Cut Score: How to Set and Defend a Passing Score

Cut Score: How to Set and Defend a Passing Score

Cut Score: How to Set and Defend a Passing Score

A cut score is the point on your exam scale that separates candidates who pass from those who fail, and for a credentialing program it is a documented policy decision rather than a number someone picked because it felt about right. Nearly every certification program I look at can tell me its passing score. Far fewer can tell me how that number was arrived at, who was in the room, and what evidence would hold up if a failed candidate challenged it. That gap is the whole problem, because the score itself is never what gets defended. The process behind it is.

Key takeaways

  • A round number is a red flag. If your passing score is 70 percent because 70 sounded reasonable, you have a policy with no evidence behind it.
  • Standard setting is an accreditation requirement. Programs seeking NCCA or ANSI accreditation are expected to show a formal, documented study.
  • The method has to match your data situation. Some approaches need live candidate performance data you will not have before launch.
  • Cut score and pass rate are different things. One is a standard, the other is an outcome. Managing the second by moving the first is how programs lose defensibility.
  • Measurement error does not disappear at the boundary. Candidates near the cut score are the ones most likely to be misclassified, and you should know by how much.

What is a cut score, really?

It is the operational definition of "minimally competent" for your credential, expressed as a number.

That framing matters more than it sounds. You are not measuring who is best. You are drawing a line at the level of knowledge and skill below which someone should not be practicing, certifying, or holding the credential. Everything above the line is a pass, whether it clears by one point or thirty.

Two interpretations sit underneath this, and they lead somewhere very different.

Criterion-referenced means the standard is absolute. Candidates are judged against a defined level of competence, and in principle everyone can pass or everyone can fail. This is what credentialing programs use, because the public interest question is whether this person is competent, not whether they beat their peers.

Norm-referenced means the standard is relative. The top thirty percent pass, or the cut sits one standard deviation below the mean. This suits selection into a fixed number of places. It suits certification badly, because it makes competence depend on who happened to sit the exam that cycle.

If you have a criterion-referenced credential and a norm-referenced cut score, you have a contradiction sitting in your program documentation.

Why is 70 percent not a defensible passing score?

Because nothing produced it except habit. As Assessment Systems puts it, for a criterion-referenced interpretation it is not legally defensible to conveniently pick a round number. You need a formal process.

Think about what a round number actually claims. It says a candidate who knows 70 percent of this content is competent to practice and one who knows 69 percent is not, and that this holds true across every form of the exam regardless of whether one form happened to be harder. Nobody believes that when it is stated plainly, yet it is the implicit position of every program running on an inherited percentage.

The exposure shows up in three places. A candidate appeals and asks how the standard was established. An accreditation review asks for your documentation. Or you revise the exam, the new form is harder, and pass rates fall with no explanation that does not sound like an accident.

Which standard setting method should you use?

Methods split into two families. Item-centered approaches ask experts to judge the questions. Examinee-centered approaches ask them to judge the candidates. Your choice depends mostly on one thing: whether you already have performance data.

MethodWhat panelists doBest when
Modified AngoffEstimate the chance a minimally competent candidate answers each item correctlyNew or continuously delivered exams with no prior candidate data. The default for credentialing
BookmarkPlace a bookmark in a booklet of items ordered by difficultyYou already have item difficulty statistics. Panelists often find it more intuitive
Contrasting groupsClassify real candidates as competent or not, then compare score distributionsYou have a cohort experts genuinely know, common in clinical settings
Borderline groupIdentify candidates who are genuinely on the fence and take their median scorePerformance and skills assessment where judges observe candidates directly

For most certification bodies the answer is a modified Angoff study, and the reason is practical rather than theoretical. A survey of credentialing organizations with NCCA-accredited programs found roughly three quarters used a modified Angoff procedure. Bookmark is elegant, but it needs item difficulty data you will not have before launch, which is exactly when you need a cut score. Our Angoff score guide walks through running that panel step by step.

What actually makes a cut score defensible?

Defensibility lives in documentation, not in the number. Six things need to be on paper.

  1. A written definition of the minimally competent candidate. Panelists cannot rate items against a standard nobody has articulated. This is the single most skipped step, and it quietly invalidates everything downstream.
  2. A qualified, representative panel. Names, credentials, practice settings, geography, years of experience. A panel of six people from one hospital is not representative of a national credential.
  3. A link back to the job analysis. Your content outline should trace to a practice analysis, and your standard should sit on top of that outline.
  4. The raw ratings and how they were aggregated. Including any second round after discussion, and what changed between rounds.
  5. The policy decision on top of the recommendation. A standard setting study produces a recommended cut score. Your board adopts one. If the board adjusted it, record the rationale.
  6. A date and a review schedule. A cut score set against a 2019 content outline is not evidence for a 2026 exam.

None of this is exotic. It is simply the difference between a program that can answer "how did you get that number" in one email and one that spends three weeks reconstructing a meeting nobody has taken minutes at.

What about measurement error at the boundary?

Here is the uncomfortable part that mature programs handle openly and immature ones ignore.

Every score carries measurement error. A candidate's observed score is their true ability plus noise from item sampling, fatigue, and chance. That error is not evenly distributed either, and the standard error of measurement conditional on score level is what tells you how precise your exam is at the point that matters, which is the cut score itself.

The consequence is that candidates sitting within a point or two of the line are the ones most likely to be misclassified in either direction. You cannot eliminate that. You can do three things about it. Build the exam so it has the most measurement precision near the cut score rather than spread evenly across the range, which is a test construction decision that traces back to how you tag and select items in your item bank. Decide and document a policy on borderline cases before you have one in front of you. And resist the temptation to add a buffer to the cut score to feel safer, because an unexplained adjustment is exactly the thing that fails review.

When should you revisit a cut score?

On a schedule, and on specific triggers. The schedule is usually tied to your practice analysis cycle, since a new content outline means the old standard no longer refers to the same construct.

The triggers matter more. Revisit when the content outline changes materially, when the exam form changes substantially, when scope of practice or regulation shifts, or when pass rates move sharply without any change in candidate preparation. That last one is a signal to investigate, not a signal to adjust the cut score. If your pass rate falls and you lower the standard to bring it back, you have replaced a criterion-referenced credential with a norm-referenced one and told nobody.

Between studies, equating across forms is what keeps the standard stable. The cut score belongs to the construct, not to a particular set of questions, which is the same discipline covered in our guide to psychometric test development.

Frequently asked questions

Is a cut score the same as a pass rate?
No, and conflating them causes real damage. The cut score is the standard you set. The pass rate is what happens when candidates meet it or do not. A stable standard with a moving pass rate is normal and usually tells you something about candidate preparation.

How many panelists do we need?
Enough to be credibly representative of the practice, which usually means eight to twelve rather than three or four. Spread across setting and geography matters as much as headcount.

Can we just use the same cut score on a new exam form?
Only if the forms are equated. Two forms built to the same outline are rarely identical in difficulty, and applying one raw score to both silently changes the standard.

Do we need a new study every year?
Usually not. Tie the cycle to your practice analysis and to material changes in the exam or the profession, and document the reasoning either way.

What if the board wants to override the recommended score?
That is legitimate, because adopting a standard is a policy act. What is not legitimate is doing it without recording who decided, on what grounds, and what the study had recommended.

The bottom line

A cut score is a claim about competence, and claims need evidence. The programs that handle challenges calmly are not the ones with a cleverer number. They are the ones that can produce the minimally competent candidate definition, the panel roster, the ratings, and the board minutes without going hunting.

Most of that evidence is generated by your assessment platform as a side effect of running the exam properly, provided it stores item statistics, form composition, and score history in a way you can retrieve years later rather than at the moment of panic. If you want to see how that record looks against your own certification program, take a look at our online assessment platform or book a demo and we will walk your exam through it.

Evaluating LMS options for your organization?

See how 200+ associations and healthcare organizations deliver continuing education, certification, and non-dues revenue with Oasis.

Book a demoNot ready? Read our case studies
Sam Hirsch

Sam Hirsch

Vice President, Sales and Marketing

Sam Hirsch is the Vice President of sales and marketing at 360 Factor. He has helped over 250 associations find the right LMS for their organization.

Share on socials:
oasis lms

Deliver learning that drives impact

Whether you're running continuing education, certification, or member growth programs, Oasis LMS helps deliver high-impact education efficiently and at scale.

Book a demo
Walk through use cases with us
Deliver learning that drives impact with Oasis LMS