
A cut score is the point on your exam scale that separates candidates who pass from those who fail, and for a credentialing program it is a documented policy decision rather than a number someone picked because it felt about right. Nearly every certification program I look at can tell me its passing score. Far fewer can tell me how that number was arrived at, who was in the room, and what evidence would hold up if a failed candidate challenged it. That gap is the whole problem, because the score itself is never what gets defended. The process behind it is.
It is the operational definition of "minimally competent" for your credential, expressed as a number.
That framing matters more than it sounds. You are not measuring who is best. You are drawing a line at the level of knowledge and skill below which someone should not be practicing, certifying, or holding the credential. Everything above the line is a pass, whether it clears by one point or thirty.
Two interpretations sit underneath this, and they lead somewhere very different.
Criterion-referenced means the standard is absolute. Candidates are judged against a defined level of competence, and in principle everyone can pass or everyone can fail. This is what credentialing programs use, because the public interest question is whether this person is competent, not whether they beat their peers.
Norm-referenced means the standard is relative. The top thirty percent pass, or the cut sits one standard deviation below the mean. This suits selection into a fixed number of places. It suits certification badly, because it makes competence depend on who happened to sit the exam that cycle.
If you have a criterion-referenced credential and a norm-referenced cut score, you have a contradiction sitting in your program documentation.
Because nothing produced it except habit. As Assessment Systems puts it, for a criterion-referenced interpretation it is not legally defensible to conveniently pick a round number. You need a formal process.
Think about what a round number actually claims. It says a candidate who knows 70 percent of this content is competent to practice and one who knows 69 percent is not, and that this holds true across every form of the exam regardless of whether one form happened to be harder. Nobody believes that when it is stated plainly, yet it is the implicit position of every program running on an inherited percentage.
The exposure shows up in three places. A candidate appeals and asks how the standard was established. An accreditation review asks for your documentation. Or you revise the exam, the new form is harder, and pass rates fall with no explanation that does not sound like an accident.
Methods split into two families. Item-centered approaches ask experts to judge the questions. Examinee-centered approaches ask them to judge the candidates. Your choice depends mostly on one thing: whether you already have performance data.
| Method | What panelists do | Best when |
|---|---|---|
| Modified Angoff | Estimate the chance a minimally competent candidate answers each item correctly | New or continuously delivered exams with no prior candidate data. The default for credentialing |
| Bookmark | Place a bookmark in a booklet of items ordered by difficulty | You already have item difficulty statistics. Panelists often find it more intuitive |
| Contrasting groups | Classify real candidates as competent or not, then compare score distributions | You have a cohort experts genuinely know, common in clinical settings |
| Borderline group | Identify candidates who are genuinely on the fence and take their median score | Performance and skills assessment where judges observe candidates directly |
For most certification bodies the answer is a modified Angoff study, and the reason is practical rather than theoretical. A survey of credentialing organizations with NCCA-accredited programs found roughly three quarters used a modified Angoff procedure. Bookmark is elegant, but it needs item difficulty data you will not have before launch, which is exactly when you need a cut score. Our Angoff score guide walks through running that panel step by step.
Defensibility lives in documentation, not in the number. Six things need to be on paper.
None of this is exotic. It is simply the difference between a program that can answer "how did you get that number" in one email and one that spends three weeks reconstructing a meeting nobody has taken minutes at.
Here is the uncomfortable part that mature programs handle openly and immature ones ignore.
Every score carries measurement error. A candidate's observed score is their true ability plus noise from item sampling, fatigue, and chance. That error is not evenly distributed either, and the standard error of measurement conditional on score level is what tells you how precise your exam is at the point that matters, which is the cut score itself.
The consequence is that candidates sitting within a point or two of the line are the ones most likely to be misclassified in either direction. You cannot eliminate that. You can do three things about it. Build the exam so it has the most measurement precision near the cut score rather than spread evenly across the range, which is a test construction decision that traces back to how you tag and select items in your item bank. Decide and document a policy on borderline cases before you have one in front of you. And resist the temptation to add a buffer to the cut score to feel safer, because an unexplained adjustment is exactly the thing that fails review.
On a schedule, and on specific triggers. The schedule is usually tied to your practice analysis cycle, since a new content outline means the old standard no longer refers to the same construct.
The triggers matter more. Revisit when the content outline changes materially, when the exam form changes substantially, when scope of practice or regulation shifts, or when pass rates move sharply without any change in candidate preparation. That last one is a signal to investigate, not a signal to adjust the cut score. If your pass rate falls and you lower the standard to bring it back, you have replaced a criterion-referenced credential with a norm-referenced one and told nobody.
Between studies, equating across forms is what keeps the standard stable. The cut score belongs to the construct, not to a particular set of questions, which is the same discipline covered in our guide to psychometric test development.
Is a cut score the same as a pass rate?
No, and conflating them causes real damage. The cut score is the standard you set. The pass rate is what happens when candidates meet it or do not. A stable standard with a moving pass rate is normal and usually tells you something about candidate preparation.
How many panelists do we need?
Enough to be credibly representative of the practice, which usually means eight to twelve rather than three or four. Spread across setting and geography matters as much as headcount.
Can we just use the same cut score on a new exam form?
Only if the forms are equated. Two forms built to the same outline are rarely identical in difficulty, and applying one raw score to both silently changes the standard.
Do we need a new study every year?
Usually not. Tie the cycle to your practice analysis and to material changes in the exam or the profession, and document the reasoning either way.
What if the board wants to override the recommended score?
That is legitimate, because adopting a standard is a policy act. What is not legitimate is doing it without recording who decided, on what grounds, and what the study had recommended.
A cut score is a claim about competence, and claims need evidence. The programs that handle challenges calmly are not the ones with a cleverer number. They are the ones that can produce the minimally competent candidate definition, the panel roster, the ratings, and the board minutes without going hunting.
Most of that evidence is generated by your assessment platform as a side effect of running the exam properly, provided it stores item statistics, form composition, and score history in a way you can retrieve years later rather than at the moment of panic. If you want to see how that record looks against your own certification program, take a look at our online assessment platform or book a demo and we will walk your exam through it.
See how 200+ associations and healthcare organizations deliver continuing education, certification, and non-dues revenue with Oasis.
Whether you're running continuing education, certification, or member growth programs, Oasis LMS helps deliver high-impact education efficiently and at scale.
