How to Tell If an NCLEX Practice Question Is Any Good
A practice question is only useful if a safe, competent nurse scores meaningfully higher on it than an unsafe one. That gap is measurable, it can be measured before any student sees the item, and almost nobody publishes it. “Written by nurse educators” and “aligned to the test plan” are descriptions of a process, not evidence about an item. This page sets out the test we apply, the number it produces, and how to use it on any vendor — including us.
The failure a question bank never tells you about
A bad practice item is rarely wrong. It is usually undiscriminating — a question that a strong candidate and a dangerous one both get right, or both get wrong, for reasons unrelated to nursing judgment. Common shapes:
- The giveaway. One option is obviously absurd, so the item tests reading, not judgment.
- The coin flip. Two options are defensible and the key is arbitrary. Everyone scores 50% regardless of competence.
- The trivia item. It hinges on recalling a number rather than deciding an action, so it separates people by memory rather than by safety.
- The ambiguous stem. The right answer depends on an unstated assumption, so the strongest candidates — the ones who notice the ambiguity — are the most likely to miss it.
None of these are detectable by reading the question and nodding. They are only visible when you compare how differently two known standards of practice perform on it.
The dual-probe test
The method is simple enough to describe in four steps and is the reason we are willing to publish a number at all:
- Attempt the item as a competent, safe practitioner. Record the score.
- Attempt the same item as a plausible but unsafe practitioner — someone who reassures rather than escalates, diagnoses rather than reports, delegates what cannot be delegated. Not a random guesser: a confident wrong answer is the realistic failure mode, and a random baseline flatters the item.
- Take the gap. Competent score minus unsafe score is the discrimination index.
- Reject anything below the floor, before a student ever sees it.
Both probes must run on the same solver. If the competent attempt uses a stronger model than the unsafe one, the gap measures the difference between the two solvers rather than anything about the item, and every question passes.
What our own bank scores
| Measure | Value | Meaning |
|---|---|---|
| Scenarios in bank | 10 | Small and deliberately slow-growing. |
| Mean discrimination index | 68.9 | Average separation between a safe and an unsafe attempt. |
| Lowest in bank | 40 | The weakest item still clears the gate by ten points. |
| Rejection floor | 30 | Below this a scenario is discarded, not revised down. |
Two honest caveats. First, this is measured on our spoken clinical scenarios, where a full transcript gives far more to score than a four-option item does; a multiple-choice question has a lower ceiling on how much separation is even possible. Second, ten scenarios is a small bank. We would rather publish a real number over a small bank than an impressive claim over a large one.
The floor matters as much as the mean. It was originally set at 20, which was below the gap the tiers already implied — meaning the check could never actually bind and the pipeline had a gate that rejected nothing. It is 30 now, and the lowest scenario in the bank sits at 40.
What to ask any prep vendor
You do not need our method to use this. Four questions, and the quality of the answers tells you most of what you need:
- “How do you decide an item is good enough to publish?” A process answer (“reviewed by nurse educators”) is weaker than a measurement answer.
- “What do you measure, and what is the number?” Any stored per-item statistic counts — discrimination, point-biserial, a difficulty calibration from real attempts.
- “What proportion of authored items get rejected?” A pipeline that rejects nothing is not a filter.
- “How many items of each type do you hold?” “Covers all NGN types” is compatible with holding three bow-tie items. Ours are published and counted.
A vendor that cannot answer these is not necessarily selling something bad. It does mean neither of you can tell.
Frequently asked questions
What is a discrimination index?
The score gap between a competent attempt and an unsafe attempt at the same item. A high gap means the item separates safe practice from unsafe practice, which is the only thing a licensure practice question is for. In the NCLEXIT voice bank it averages 68.9 against a rejection floor of 30.
Why measure it before students see the item?
Because measuring it afterwards means students were the experiment. Classical item analysis needs hundreds of live attempts to stabilise, so a new item is unvalidated for months. Probing it with a known-safe and known-unsafe attempt gives you the same signal on day one.
Can this be applied to multiple-choice questions?
Yes, though the ceiling is lower — a single-best item has far less to separate on than a full spoken transcript. The principle is unchanged: if a knowingly unsafe approach scores about the same as a competent one, the item is not measuring judgment.
Is a higher discrimination index always better?
Not without limit. An extremely high gap can mean the item has one obvious trap rather than a genuine judgment. It is a floor to clear and a distribution to watch, not a score to maximise.
Do other NCLEX prep companies publish this?
We have not found one that publishes a per-item quality measurement. Most publish bank size and author credentials. That is why the four questions above are worth asking directly rather than looking for the answer on a marketing page.
See a scenario that failed the gate
Drop your email and we’ll send a rejected scenario alongside a passing one, with both probe transcripts and the scores.