How to Write Multiple Choice Questions That Actually Measure Something

How to write multiple choice questions, compressed to a single sentence: write the stem as a complete, answerable question, make every wrong option something a person with partial knowledge would genuinely pick, and keep all options the same length and grammatical shape. Almost every weak quiz question breaks one of those three rules. This guide unpacks each one, with examples, the funny-question variant done properly, and the scoring math that most guides skip entirely.
I build quiz software for a living (Uplup), which means I read more multiple choice questions in a week than most teachers write in a year, and here’s what that firehose teaches you: the stem is rarely the problem. Writers polish the question and improvise the wrong answers, and the wrong answers are where measurement lives or dies.

The anatomy: stem, key, and distractors
Vocabulary first, because once you have words for the parts, the fixes get easy:
- The stem is the question itself: “Which planet has the shortest day?”
- The key is the correct answer: Jupiter.
- The distractors are the wrong options, and the name is the job description. A distractor must distract: it should attract takers whose knowledge is incomplete. “Mars, Venus, Saturn” distract; “Cheese, Tuesday, France” do not.
A multiple choice question measures exactly as well as its weakest distractor. With four options and one silly one, your question is really a three-option question; with two silly ones, a coin flip. Here’s the math laid out plainly, because it’s the fun part: a blind guesser scores 25% on a true four-option question. Each dead distractor raises that floor: one dead option pushes blind guessing to 33%, two dead options to 50%. Write ten questions with two lazy distractors each, and a taker who knows nothing averages 5/10, which looks a lot like “moderate knowledge” on your results screen. Your quiz isn’t measuring knowledge anymore; it’s measuring test-savviness.

Writing stems that don’t leak
The stem’s job is to pose one clear problem, and its most common failure is leaking the answer. The classic leaks:
- Grammar leaks. “An ___” followed by one vowel-starting option. Read every option into the stem aloud before publishing; your ear catches what your eye forgives.
- Length leaks. Writers qualify the true answer until it’s precise, so it grows. Experienced takers pick the longest option, and they’re right often enough that it’s honestly funny. Trim the key or grow the distractors.
- Echo leaks. A word from the stem reappears only in the key. “What does the tracking pixel do?” with one option containing “tracks.”
- The negative trap. “Which of these is NOT…” questions test careful reading more than knowledge. If you need one, bold the NOT, and never stack a negative stem with negative options.
- The double-barrel. “Which planet is largest and closest to the sun?” No option is cleanly right, and your sharpest takers are the ones who’ll notice.
One more, and this one’s opinionated: kill “all of the above.” A taker who recognizes two true options picks it without evaluating the third. It rewards process of elimination, not knowledge, and its cousin “none of the above” tells you what takers don’t think without revealing what they do. If several answers are genuinely correct, you want a checkbox question with partial credit, not a multiple choice question wearing a disguise.
The distractor workshop
Since the distractors carry the measurement, here’s where I go mining for good ones:
- Real mistakes. The best distractors are answers people actually give. If you teach, your students’ wrong answers from open-ended homework are a distractor goldmine. If you’re quizzing customers, the misconceptions in your support inbox are yours.
- True statements that don’t answer this question. “Which hook keeps the audience past the first sentence?” A distractor that’s a perfectly good writing tip, just not a hook, catches everyone skimming.
- The near-miss. Off by one step, one year, one decimal place. “1969” for a 1968 event distracts; “1932” doesn’t.
- The overgeneralization. The rule-of-thumb version of the precise answer. It catches takers who learned the shortcut but not the exception.
The quality test for a finished question: could a smart person who never studied this material eliminate their way to the answer? If yes, the distractors need work. And after real people take it, the data will grade your craft for you: a distractor nobody ever picks is dead weight, and a distractor outpolling the key means either a brilliant trap or a broken question. Look honestly, then decide which.
Funny multiple choice questions (done properly)
Half the search traffic around this topic wants comedy, and comedy has a real place: an absurd option relieves tension, humanizes a dry quiz, and gives the 100%-wrong taker something to enjoy. The craft rule is placement. A funny option can replace the fourth distractor; it must never replace the second. One joke among three real options keeps the question honest (the guessing floor only moves from 25% to 33%); two jokes turn your question into a coin flip with a laugh track.
A few structures that work:
- The escalating absurdity: “What’s the maximum length of a tweet? (a) 280 characters (b) 140 characters (c) 4,000 characters (d) However long your ex’s apology should have been”
- The suspiciously specific: “(d) Whatever Dave from accounting says, Dave knows things”
- The self-aware option: “(d) I’m just here to see the answer”
And one honest boundary: in graded assessments, certifications, or anything with stakes, skip the jokes entirely. Comedy in a compliance exam reads as contempt for the taker’s time, and besides, nobody has ever laughed at question 14 of a safety cert.
Application beats recall (when you can afford it)
The quiet upgrade available on almost any multiple choice question: move the stem from recall to application. “What is conditional logic?” tests whether someone memorized a definition. “A respondent answers ‘No’ to question 2; which question should appear next?” tests whether they can use the idea, and it’s nearly impossible to answer from pattern-matching alone.
The recipe is mechanical enough to apply on a rewrite pass:
- Definition stems become scenario stems: put the concept in a situation and ask what happens, what’s wrong, or what comes next.
- “Which is true about X” becomes “which would you do”: a decision shows you understanding that recognition hides.
- Single facts become short cases: two sentences of setup buys you a question that measures judgment, and judgment is usually what you actually cared about.
The cost is length, which is why the mix matters: recall questions for vocabulary and fundamentals, application questions for everything you’d fire or promote someone over. A multiple choice quiz that’s all recall flatters memorizers; a quiz that’s all cases exhausts everyone. Half and half, application weighted at more points, is the blend that measures without punishing.
Context changes the rules
The same craft bends differently depending on where the question lives, and mixing up the contexts is one of the sneakier ways good questions go bad.

In a classroom or a training course, explanations are mandatory, partial credit is kind, and recycling real student errors as distractors turns the quiz itself into a diagnostic. Also, let difficulty climb across the quiz; early failure teaches people to stop trying.
Marketing and engagement quizzes flip the priorities, because speed rules everything there. Keep stems under 20 words and options under 8, allow yourself one funny fourth option, and pitch the difficulty so most people score 60 to 80%. A marketing quiz that makes prospects feel dumb has converted its last lead.
Certification and screening sit at the strict end: no jokes, no “all of the above,” and shuffled question banks so no two takers see identical tests. I’ve noticed the stakes attract the test-savvy, so this is where the distractor-proximity dial matters most: reach for closer distractors, never obscurer facts.
And live trivia has a constraint the others don’t: read-aloud length. If the host can’t deliver the stem plus four options in 15 seconds, trim it. Options takers must hold in memory should be short enough to hold in memory.
Reading the data your questions send back
Once real people answer, every question generates two numbers worth reading, and neither requires a statistics degree:
- Difficulty (the percent who got it right). Under 30% means the question’s too hard, mis-keyed, or brilliant; over 90% means it’s a warm-up, which is fine if you placed it as one. There’s no virtue in a uniform 50%: a good quiz has an intentional difficulty shape.
- Distractor pull (which wrong options get chosen). A distractor drawing nobody is dead weight; replace it or drop to three options. A distractor out-polling the key is a red alert: either the key is wrong, the wording is ambiguous, or you’ve discovered a widespread misconception, which is the most useful finding a quiz can produce.
The workflow is simple: publish, collect 30+ responses, scan the per-question stats, fix the two worst questions, repeat. Two rounds of that loop improves a quiz more than any amount of pre-launch polishing, and I say that as someone whose job is theoretically the pre-launch polishing.
A quick word on option order
The folklore says “when in doubt, pick C,” and the folklore exists because human question-writers really do bury keys in the middle positions. The fix costs nothing: shuffle answer order per taker and the entire position game evaporates, along with any accidental patterns like alphabetized keys or “the true statement always comes last.” If your tool can’t shuffle, then at minimum assign key positions with a die roll rather than your instincts, because your instincts have a favorite letter and takers will find it.
Difficulty is a dial, not an accident
You can set a question’s difficulty deliberately, and the lever is almost always distractor proximity, not obscurity:
- Easy: distractors from a different category (“Which is a mammal?” with two birds and a fish).
- Medium: distractors from the same category (four mammals, one lays eggs).
- Hard: distractors that are true in adjacent contexts (four correct definitions, one matches this term).
Raising difficulty by asking more obscure facts is the lazy dial, and it changes what you’re measuring from understanding to exposure. Raising it through closer distractors keeps measuring the same knowledge, just at higher resolution. I genuinely think this is the least-taught idea in question writing, and it’s the one that separates quizzes that feel fair-but-hard from quizzes that feel like trick shows.
Scoring: points, partial credit, and the checkbox cousin
Multiple choice scoring is binary: right or wrong, full points or zero, and that’s correct for single-answer questions. The interesting decisions live nearby:
- Weight by difficulty. A hard question at 3 points and warm-ups at 1 make an 80% score mean something. Flat scoring across a mixed-difficulty quiz compresses everyone toward the middle.
- The multi-answer variant (“select all that apply”) deserves partial credit. Run the example: a 10-point question with 3 correct options among 5. A taker who picks 2 of the 3 with no wrong selections earns 6.7 under proportional credit and 0 under all-or-nothing. Across a quiz, that’s the difference between a score that tracks knowledge and one that punishes near-mastery like total ignorance. Proportional for learning, all-or-nothing for certification.
- Explanations after every answer. The moment after answering wrong is the single most teachable moment a quiz has; a one-line “why” turns your quiz from a filter into a teacher!
In Uplup, these are settings rather than aspirations: per-question points, both partial-credit modes on checkbox questions, and an explanation field on every question that shows after answering. Shuffle answer order while you’re in there; it neutralizes the “C looks lonely” guessing folklore and any position patterns you didn’t notice you had.
How to write multiple choice questions: a worked example
Enough theory. Let’s run the whole checklist on one real question, worst version first:
Weak: “SEO is important because… (a) It’s not (b) It helps people find you on search engines and drives organic traffic to your website over time (c) Cheese (d) All of the above”
Every leak at once: a joke where a distractor belongs, a second dead option, a length leak on the key, and “all of the above” contradicting option (a) into paradox.
Strong: “A page ranks #4 for a keyword. Which change most directly improves its click-through rate? (a) Rewriting the title tag and meta description (b) Adding more internal links to the page (c) Compressing its images (d) Buying backlinks to the domain”
One clear scenario, one key, three distractors that are all real SEO activities a partially-informed person might pick, matched lengths, no leaks. The difference between the two isn’t talent; it’s the checklist above, applied.
Frequently asked questions
The questions quiz writers send me most, short answers first.
What makes a good multiple choice question?
A clear, complete stem posing one problem, a single unambiguous key, and distractors that people with partial knowledge would plausibly choose, all in matching length and grammar. Most of how to write multiple choice questions well comes down to the distractors: that’s where the measurement actually happens, so polish them hardest.
How many options should a multiple choice question have?
Four is the standard and three is often better. Research on option counts consistently finds that a strong third distractor beats a weak fourth one, because a dead option only raises the guessing floor. Never pad to five with filler.
Should I use “all of the above” in quiz questions?
No. It rewards partial recognition (spot two true options and pick it without reading the third), breaks the one-key principle, and its data tells you almost nothing. If multiple answers are truly correct, use a checkbox question with partial credit instead.
How do you make multiple choice questions harder without being unfair?
Move the distractors closer to the key: same category, adjacent contexts, near-miss values. Difficulty through proximity keeps measuring understanding; difficulty through obscurity just measures who happened to be exposed to the fact.
Are funny multiple choice questions okay in a real quiz?
Yes, in engagement and trivia contexts, with one rule: one joke option maximum per question, replacing the fourth distractor rather than the second, so the question stays honest. Skip humor entirely in graded or certification settings.
How many points should each question be worth?
Weight by difficulty: 1 point for warm-ups, 2 to 3 for questions with close distractors, and partial credit on any select-all-that-apply variant. Flat scoring across mixed difficulty compresses scores and hides what takers actually know.
Do this next
Dig up the last quiz you wrote and audit only the distractors: for each question, ask which options a smart-but-unstudied person could eliminate from the armchair. Rewrite the two worst using the mining list above (real mistakes, near-misses, true-but-irrelevant statements), then rebuild it with shuffled options and per-question explanations and watch the answer distribution on the next 20 takers. Twenty real takers will teach you more about your questions than a hundred re-reads, and watching that distribution come in never gets old.
