created: 2026-06-25 topic: assessment status: fresh
Exam Item-Writing Rubric and Linux Foundation / CNCF MCQ Style¶
Spike date: 2026-06-25. Part A is the psychometric item-writing rubric, drawn from the Haladyna, Downing, and Rodriguez (2002) revised taxonomy of 31 guidelines and the NBME item-writing guide. Part B documents the format and tone of Linux Foundation / CNCF knowledge-based (multiple-choice) exams from the official CNCF and Linux Foundation Training curriculum pages.
The Haladyna taxonomy is the canonical, research-validated list. It is quoted verbatim from the published paper. The NBME "cover-the-options" rule and the clinical-vignette pattern are the second pillar. The two agree on every substantive point, which is why the rubric below is defensible rather than one author's opinion.
One sourcing caveat is load-bearing and called out again in Part B: the Linux Foundation does NOT publish the passing score for its associate exams on the official curriculum pages. The widely cited 75% is third-party reporting.
PART A: The item-writing rubric (psychometric)¶
A.0 The two anchoring principles¶
Two rules sit above the checklist. Get these right and most of the rest follows.
-
Single best answer, exactly one defensible key. One option is correct or clearly best. Every other option is wrong, but wrong in a way that is plausible to a test-taker who has not mastered the material. (Haladyna #19, #29.)
-
The cover-the-options test (NBME). A well-built item can be answered from the stem and lead-in ALONE, with the options covered. If a test-taker cannot form the answer before reading the choices, the stem is under-specified or the item is testing recognition rather than knowledge. This is the single most useful editing pass: write the stem, cover the options, and confirm you could answer it. (NBME, "Constructing Written Test Questions.")
A.1 The Haladyna 31-guideline taxonomy (verbatim)¶
From Haladyna, Downing, and Rodriguez (2002), Table 1, "A Revised Taxonomy of Multiple-Choice (MC) Item-Writing Guidelines." Reproduced verbatim so the rubric traces to source.
Content concerns
- Every item should reflect specific content and a single specific mental behavior, as called for in test specifications (two-way grid, test blueprint).
- Base each item on important content to learn; avoid trivial content.
- Use novel material to test higher level learning. Paraphrase textbook language or language used during instruction when used in a test item to avoid testing for simply recall.
- Keep the content of each item independent from content of other items on the test.
- Avoid over specific and over general content when writing MC items.
- Avoid opinion-based items.
- Avoid trick items.
- Keep vocabulary simple for the group of students being tested.
Formatting concerns
- Use the question, completion, and best answer versions of the conventional MC, the alternate choice, true-false, multiple true-false, matching, and the context-dependent item and item set formats, but AVOID the complex MC (Type K) format.
- Format the item vertically instead of horizontally.
Style concerns
- Edit and proof items.
- Use correct grammar, punctuation, capitalization, and spelling.
- Minimize the amount of reading in each item.
Writing the stem
- Ensure that the directions in the stem are very clear.
- Include the central idea in the stem instead of the choices.
- Avoid window dressing (excessive verbiage).
- Word the stem positively, avoid negatives such as NOT or EXCEPT. If negative words are used, use the word cautiously and always ensure that the word appears capitalized and boldface.
Writing the choices
- Develop as many effective choices as you can, but research suggests three is adequate.
- Make sure that only one of these choices is the right answer.
- Vary the location of the right answer according to the number of choices.
- Place choices in logical or numerical order.
- Keep choices independent; choices should not be overlapping.
- Keep choices homogeneous in content and grammatical structure.
- Keep the length of choices about equal.
- None-of-the-above should be used carefully.
- Avoid All-of-the-above.
- Phrase choices positively; avoid negatives such as NOT.
- Avoid giving clues to the right answer, such as
- a. Specific determiners including always, never, completely, and absolutely.
- b. Clang associations, choices identical to or resembling words in the stem.
- c. Grammatical inconsistencies that cue the test-taker to the correct choice.
- d. Conspicuous correct choice.
- e. Pairs or triplets of options that clue the test-taker to the correct choice.
- f. Blatantly absurd, ridiculous options.
- Make all distractors plausible.
- Use typical errors of students to write your distractors.
- Use humor if it is compatible with the teacher and the learning environment. (Note: this last one is for classroom assessment. For a high-stakes practice exam, drop it. Humor is a flaw in a certification-style item.)
A.2 The checkable rubric (apply to every item)¶
Each line is a yes/no gate. An item ships only when every applicable line is yes. Haladyna rule numbers and NBME are cited so a reviewer can trace each gate.
Stem
- [ ] The stem asks ONE clear thing. A single problem, a single mental behavior. (Haladyna #1.)
- [ ] The cover-the-options test passes: a competent test-taker could answer from the stem alone. (NBME.)
- [ ] The central idea is in the stem, not scattered across the options. (Haladyna #15.)
- [ ] The stem is positively worded. If negative (NOT, EXCEPT, LEAST), the negative word is capitalized and bold, and the negation is unavoidable rather than a trap. (Haladyna #17.)
- [ ] No window dressing. Every clause in the stem is load-bearing or is deliberate scenario context, not filler. (Haladyna #16.)
- [ ] Vocabulary matches the audience level; the difficulty is in the concept, not the wording. (Haladyna #8.)
Options as a set
- [ ] Exactly one defensible correct answer. A second-guessing subject-matter expert cannot argue a second option is also correct. (Haladyna #19.)
- [ ] All options are homogeneous: same category of thing (all commands, all time values, all failure causes), parallel grammar, parallel structure. (Haladyna #23.)
- [ ] Options are about equal length. The correct answer is NOT systematically the longest, the most qualified, or the most detailed. Length parity is a hard gate because length is the most common unintended cue. (Haladyna #24, #28d.)
- [ ] Options are mutually exclusive and non-overlapping. No option contains or implies another; no two can be simultaneously true. (Haladyna #22.)
- [ ] Options are in a logical or numerical order (ascending values, alphabetical, chronological) so ordering carries no signal. (Haladyna #21.)
- [ ] No "all of the above." (Haladyna #26.)
- [ ] No "none of the above" unless the key is objectively, computationally verifiable and the option is genuinely needed. Default to not using it. (Haladyna #25.)
- [ ] Three or four options. Research supports three effective options; a fourth is fine only if it is a real distractor, not filler. Do not pad to four with an implausible option. (Haladyna #18.)
Distractors
- [ ] Every distractor is plausible to someone who has not mastered the material: it represents a real misconception, a common student error, or a true-but-irrelevant fact. (Haladyna #29, #30.)
- [ ] Every distractor is factually possible and real. No invented commands, fictional flags, or nonexistent APIs. A distractor must be a thing that exists but is wrong here, not a thing that does not exist. (Haladyna #29.)
- [ ] No blatantly absurd or joke option. Absurd options reduce the effective option count and cue test-wise candidates. (Haladyna #28f.)
Test-wiseness flaws (the cue audit)
- [ ] No absolute words (always, never, all, none, completely) in any option unless they are defensibly, technically correct. (Haladyna #28a.)
- [ ] No grammatical cue: the correct option agrees with the stem's grammar no better and no worse than the distractors (a/an, singular/plural, verb tense). (Haladyna #28c.)
- [ ] No clang association: the correct option does not echo a distinctive word from the stem more than the distractors do. (Haladyna #28b.)
- [ ] No conspicuous correct choice: the key is not set apart by being more specific, more technical, more hedged, or more central than the rest. (Haladyna #28d.)
- [ ] No convergence cue: the correct answer is not the option that shares the most elements with the others (the "average of the options" trap). NBME flags this explicitly.
Cognitive level
- [ ] The item's level is deliberate, not accidental. Recall items test a single fact; application items require the test-taker to use a fact in a situation. (See A.3.)
- [ ] Application items use novel material, not verbatim text from the source. Paraphrase so recognition of familiar wording cannot substitute for understanding. (Haladyna #3.)
Independence and bank hygiene
- [ ] This item does not give away the answer to another item, and another item does not give away this one. (Haladyna #4.)
- [ ] Across the bank, the correct answer is roughly evenly distributed across positions A/B/C/D. No slot is favored. (Haladyna #20.)
Style
- [ ] Spelling, grammar, punctuation, and capitalization are correct and consistent across stem and all options. (Haladyna #12.)
- [ ] Formatted vertically: one option per line. (Haladyna #10.)
- [ ] Not a trick item. The item rewards knowledge, not the ability to spot a gotcha. (Haladyna #7.)
A.3 Recall versus application, with examples¶
Recall (definition retrieval). Tests whether the fact is stored. Lower cognitive demand. Useful for foundational terminology but easy to over-use.
In Kubernetes, which object is the smallest deployable unit that can be created and managed?
- A. Pod
- B. Container
- C. Node
- D. Deployment
This is pure recall. The answer is a definition. It is fine in moderation; a bank made only of these tests memorization, not competence.
Application (judgment in a scenario). Gives a situation and asks the test-taker to apply a principle to decide. Higher cognitive demand. This is where scenario-based items earn their place.
A platform engineer notices that a Deployment's Pods are repeatedly killed and restarted.
kubectl describe podshows the container exceeded its configured memory limit each time. Which change is the most appropriate first response?
- A. Increase the container's memory limit to match its observed working set
- B. Add a liveness probe with a longer timeout
- C. Increase the Deployment's replica count
- D. Remove the memory limit so the container can use node memory freely
Here A is best (the symptom is an OOMKill against the limit). B addresses a different failure mode (a hung process), C scales a problem that is not capacity related, and D is a real action that is wrong because removing limits trades one failure for noisy-neighbor risk. Each distractor maps to a genuine misconception. The test-taker must reason from the evidence in the stem, not recall a definition.
How to convert a recall item into an application item. Replace "what is X" with a situation where X is the right or wrong move, and write distractors that are each a plausible-but-wrong action a real practitioner might take.
A.4 Rationale quality¶
Every item carries an answer rationale. For a practice exam (the whole point of which is learning), the rationale is as important as the item.
- [ ] Explain WHY the correct answer is correct, grounded in a verifiable fact or principle, not "because it is the best."
- [ ] Explain WHY EACH distractor is wrong, individually. "B is wrong because a liveness probe addresses hangs, not memory exhaustion." A rationale that only defends the key teaches nothing about the traps.
- [ ] Every claim in the rationale is factually correct and, where possible, cites the authoritative source (official docs, the spec, the curriculum).
- [ ] The rationale does not introduce a fact that contradicts the item or reveals a second defensible answer. If writing the rationale surfaces a second correct option, fix the item.
A.5 Difficulty calibration and answer-position balance¶
- [ ] Difficulty comes from the concept and the quality of the distractors, not from obscure trivia, tricky wording, or a buried negative.
- [ ] A practice bank should span difficulty: some straightforward recall to build confidence, a majority at application level, a few hard discrimination items. Match the target exam's mix (see Part B).
- [ ] Answer-position balance is enforced at the BANK level, not the item level. After authoring, count how often the key lands in each slot and rebalance so no position is systematically favored. (Haladyna #20.)
- [ ] If you later get response data, retire or rewrite items that everyone gets right (no discrimination) or that high-scorers miss more than low-scorers (a likely keying or cue error).
PART B: Linux Foundation / CNCF knowledge-based MCQ exam style¶
B.1 What "knowledge-based" means here¶
The Linux Foundation runs two kinds of certification exam. The well-known professional exams (CKA, CKAD, CKS) are PERFORMANCE-based: a live terminal, real clusters, hands-on tasks. The associate-tier exams below are KNOWLEDGE-based: online, proctored, multiple-choice. This spike is about the second kind. Practice questions for these should look like the rubric in Part A: clean single-best-answer MCQs, scenario-leaning but not requiring a live environment.
B.2 Common format across the associate exams¶
The pattern is consistent across the CNCF associate / "Certified Associate" exams. Confirmed against the official Linux Foundation Training and CNCF pages:
- Delivery: online, proctored.
- Format: multiple-choice (some exams explicitly add multiple-SELECT items, e.g. PCA).
- Length: 60 questions.
- Duration: 90 minutes.
- Cost: USD 250, one free retake included.
- Validity: 2 years (KCNA confirmed on the LF page; the associate tier shares this).
- Scheduling window: 12 months to sit the exam.
- Passing score: widely reported as 75%, but see the flag in B.5. This is NOT printed on the official curriculum pages.
Each exam publishes a domain breakdown with percentage weightings; the percentage maps directly to how many of the 60 questions come from that domain. Build a practice bank to the same weights.
B.3 Per-exam domain weightings (from official CNCF / LF pages)¶
KCNA - Kubernetes and Cloud Native Associate. 60 questions, 90 minutes, valid 2 years. The LF curriculum page lists four headline domains; note that a second, more granular five-domain breakdown (adding Cloud Native Observability) circulates and appears to reflect a curriculum revision. Confirm the live weighting on the official page before authoring, since CNCF revises these.
LF page (four domains):
- Kubernetes Fundamentals: 44%
- Container Orchestration: 28%
- Cloud Native Application Delivery: 16%
- Cloud Native Architecture: 12%
Reported revised five-domain form (verify before relying): Kubernetes Fundamentals 46%, Container Orchestration 22%, Cloud Native Architecture 16%, Cloud Native Observability 8%, Cloud Native Application Delivery 8%.
KCSA - Kubernetes and Cloud Native Security Associate. 60 questions, 90 minutes. Six domains (2026):
- Kubernetes Cluster Component Security: 22%
- Kubernetes Security Fundamentals: 22%
- Kubernetes Threat Model: 16%
- Platform Security: 16%
- Overview of Cloud Native Security: 14%
- Compliance and Security Frameworks: 10%
PCA - Prometheus Certified Associate. 90 minutes; multiple-choice AND multiple-select; closed book (no Prometheus instance, no external sites). Five domains (2026):
- Observability Concepts: 18%
- Prometheus Fundamentals: 20%
- PromQL: 28%
- Instrumentation and Exporters: 16%
- Alerting and Dashboarding: 18%
OTCA - OpenTelemetry Certified Associate. Reported as 60 questions, 90 minutes, 75% pass. Four domains (2026):
- Fundamentals of Observability: 18%
- The OpenTelemetry API and SDK: 46%
- The OpenTelemetry Collector: 26%
- Maintaining and Debugging Observability Pipelines: 10%
CGOA - Certified GitOps Associate. Online, proctored, multiple-choice; USD 250 with one retake. Five domains (from the official CNCF page):
- GitOps Terminology: 20%
- GitOps Principles: 30%
- GitOps Patterns: 20%
- Related Practices (CaC, IaC, DevOps, DevSecOps, CI/CD): 16%
- Tooling: 14%
The CGOA is explicitly vendor-neutral: it tests GitOps principles, not Argo CD or Flux specifics, though those tools appear as study context. Practice items should test the concept, not a particular tool's flags.
CCA - Cilium Certified Associate. Online, proctored, multiple-choice; USD 250 with one retake. Eight domains (2026):
- Architecture: 20%
- Network Policy: 18%
- Service Mesh: 16%
- Network Observability: 10%
- Installation and Configuration: 10%
- Cluster Mesh: 10%
- eBPF: 10%
- BGP and External Networking: 6%
Argo. As of this spike there is no standalone "Certified Argo Associate" multiple-choice exam published on the CNCF certification catalog. Argo CD appears as study material inside the GitOps (CGOA) context. Do not assert an Argo associate exam exists without confirming on the live CNCF catalog.
B.4 Tone and style of LF / CNCF knowledge exams¶
From the official "how to ace" CNCF blog posts and the curriculum framing, the style of these knowledge-based exams is:
- Concept and terminology first, with applied judgment mixed in. They are not pure trivia. A meaningful share of items are short scenarios ("given this situation, which is true / which would you do"), not just "define X." Build practice banks with a recall-to-application mix that leans toward application, matching A.3.
- Vendor-neutral and project-accurate. Items track the upstream project's own terminology and current documented behavior. The authoritative reference is the project docs and the published curriculum, not blog folklore.
- Breadth over depth at the associate tier. The exams sample widely across the domain map rather than drilling deep into one area. Distribute practice items to the published weights.
- No live environment. Unlike CKA/CKAD/CKS, there is no terminal. Questions must be answerable by reasoning, so a good practice item never depends on running a command to see output; it can describe the output in the stem.
- Closed book. PCA is explicitly closed-book; treat the associate tier as closed-book generally. Practice items should not assume doc lookup.
B.5 What could NOT be confirmed from an authoritative source¶
Flagged honestly, per the brief:
- Passing score. The Linux Foundation does NOT publish a passing percentage on the official KCNA, KCSA, CGOA, or CCA curriculum pages (verified: the official pages omit it). The "75%" figure is consistent across many third-party study sites and one Medium write-up, but it is not first-party. The official line is that the cut score is set per exam and disclosed in the candidate handbook / score report, not advertised. Treat 75% as a strong third-party convention, not a published fact, and do not present it to learners as official.
- Exact question count and duration for some exams. KCSA, OTCA (60q / 90m) are consistently reported and align with the tier, but the official CNCF/LF pages I fetched did not always print the count and duration inline (the KCNA LF page did print 90 minutes; the CNCF KCNA page omitted count and duration). Where a figure here comes from third-party reporting rather than the official page, treat it as the tier convention (60 questions, 90 minutes) and verify on the live page before publishing a practice product.
- KCNA domain weighting drift. Two different domain breakdowns are in circulation (a four-domain and a five-domain form). CNCF revises curricula on its own schedule. Pull the weighting from the live official page at authoring time rather than trusting this snapshot.
- Argo associate exam. Not found in the official catalog as a distinct knowledge exam at spike time. Verify before assuming it exists.
Sources¶
Item-writing (Part A):
- Haladyna, T. M., Downing, S. M., and Rodriguez, M. C. (2002). A Review of Multiple-Choice Item-Writing Guidelines for Classroom Assessment. Applied Measurement in Education, 15(3), 309-334. (Table 1, the 31-guideline taxonomy, quoted verbatim.) https://www.tandfonline.com/doi/abs/10.1207/S15324818AME1503_5
- Haladyna, T. M., and Downing, S. M. (1989). A Taxonomy of Multiple-Choice Item-Writing Rules. Applied Measurement in Education, 2(1), 37-50. https://eric.ed.gov/?id=EJ391516
- Penn State Schreyer Institute, Multiple-Choice Item-Writing Rules (summary of Haladyna and Downing). https://www.schreyerinstitute.psu.edu/pdf/Multiple_Choice_Item_Writing_Rules.pdf
- NBME Item-Writing Guide, "Constructing Written Test Questions for the Health Sciences." https://www.nbme.org/educators/item-writing-guide and the PDF at https://www.nbme.org/sites/default/files/2021-02/NBME_Item%20Writing%20Guide_R_6.pdf
- NBME one-best-answer / cover-the-options principle (secondary summary): https://www.medschoolgurus.com/post/the-blueprint-behind-all-usmle-questions
Linux Foundation / CNCF exams (Part B):
- KCNA, Linux Foundation Training: https://training.linuxfoundation.org/certification/kubernetes-cloud-native-associate/
- KCNA, CNCF: https://www.cncf.io/training/certification/kcna/
- KCSA, Linux Foundation Training: https://training.linuxfoundation.org/certification/kubernetes-and-cloud-native-security-associate-kcsa/
- PCA, CNCF: https://www.cncf.io/training/certification/pca/ and the "how to ace the PCA" post: https://www.cncf.io/blog/2024/11/07/how-to-ace-the-prometheus-certified-associate-pca-exam/
- OTCA, Linux Foundation Training: https://training.linuxfoundation.org/certification/opentelemetry-certified-associate-otca/ and CNCF: https://www.cncf.io/training/certification/otca/
- CGOA, CNCF: https://www.cncf.io/training/certification/cgoa/ and Linux Foundation Training: https://training.linuxfoundation.org/certification/certified-gitops-associate-cgoa/
- CCA, CNCF: https://www.cncf.io/training/certification/cca/ and Linux Foundation Training: https://training.linuxfoundation.org/certification/cilium-certified-associate-cca/
- CNCF certification catalog (to confirm which exams exist): https://www.cncf.io/training/certification/