What Is a BARS (Behaviorally Anchored Rating Scale)?

The Fabric Team
July 22, 2026
9 min read

A behaviorally anchored rating scale AKA BARS is a performance rating method where every point on the scale is tied to a specific, observable behavior. Instead of "rate this person 1 to 5 on communication," a BARS reads "5 = summarises a technical problem in one sentence a non-engineer understands; 3 = explains the problem with jargon, requires follow-up questions to be understood; 1 = describes the symptom, not the problem, and gets stuck when pressed." That anchoring is the whole point. It replaces subjective impressions with observations two raters can actually agree on. This guide explains what BARS is, walks through a worked five-point scale, shows how to build one, and covers where it fits (and where it doesn't) in modern performance and hiring processes.

What is a behaviorally anchored rating scale?

A behaviorally anchored rating scale is a performance appraisal instrument in which each rating point on a numeric scale is defined by a description of a specific, observable behavior representative of that level of performance. The method was introduced in the 1963 industrial-organizational psychology literature by Smith and Kendall, who were trying to solve a rating-scale problem that hasn't gone away in six decades: two managers rating the same employee frequently disagree, not because the employee's performance is ambiguous, but because "rate their communication 1 to 5" means whatever each rater privately thinks it means. BARS constrains what a rating means. A "4" for a customer-support agent isn't "pretty good communication" — it is "acknowledged the issue, restated it, walked the customer through resolution, confirmed the fix at close." Behaviorally anchored means the anchors are behaviors, not adjectives. A BARS is one of the more rigorous rating instruments available for judgment-heavy roles, which is why it is still routinely used in performance management, promotion reviews and structured interview scoring.

BARS example: a worked 5-point scale

The clearest way to see BARS is a full example. Here is a five-point BARS for one dimension handling customer objections on a customer-success role:

RatingAnchor behavior
5 — ExceptionalReframes the objection back to the underlying need, addresses that need with a concrete option, and closes with a mutually agreed next step.
4 — StrongAcknowledges the objection, addresses the customer's stated concern, and proposes a next step, though may miss the underlying need behind the objection.
3 — Meets barAnswers the objection factually and accurately, but does not actively steer the conversation forward or propose a next step.
2 — Below barAddresses part of the objection or defers it without a clear plan to return; leaves the customer to drive the next step.
1 — UnacceptableBecomes defensive, changes the subject, or ignores the objection outright; customer disengages or escalates.

Notice what the anchors do not say. They don't say "great communication skills" or "positive attitude." They describe what the person did reframed, acknowledged, answered, deferred, ignored and let the rater match observed behavior to the closest anchor. If two reviewers watch the same call and one picks 4 and the other picks 3, the disagreement is now over an interpretable behavioral difference ("did they propose a next step or not?") rather than an untraceable difference in taste.

BARS vs. other rating scales

BARS is one of several performance-rating instruments. The comparison that matters most in practice is against the graphic rating scale (the default in most performance management software) and the behavioral observation scale (BOS), which is BARS's closest cousin.

MethodWhat raters doStrengthsWeaknesses
Graphic rating scaleRate a trait (e.g. "communication") on a numeric or labelled scaleFast, cheap to build, familiarRater subjectivity is very high; two ratings of "4" often mean different things
BARSMatch observed behavior to the closest anchor on a scaleHigh rater consistency; defensible; forces observation of behaviorExpensive to build; scale must be rebuilt when the role changes materially
BOS (Behavioral Observation Scale)Rate the frequency with which specific behaviors were observed (e.g. 1 = never, 5 = always)Good rater consistency; captures behavior frequencyLoses the qualitative "how well" dimension that BARS captures
MBO (Management by Objectives)Rate against pre-agreed outcome targetsAligns rating to business outcomesDoesn't measure how results were achieved, which matters for coaching

The choice usually comes down to what you're trying to defend. If you need to explain to a candidate, an internal appeals process, or an employment tribunal why a rating was what it was, BARS gives you the strongest audit trail because the anchor itself is the reasoning.

How to develop a BARS in 6 steps

The Smith–Kendall procedure has been refined over the years but the shape is unchanged. Building a defensible BARS for a role takes roughly the following six steps:

  1. Identify performance dimensions. With subject-matter experts (SMEs) for the role, list 4–8 dimensions on which the role's performance genuinely varies (e.g. for a support agent: handling objections, escalation judgment, technical accuracy, tone, throughput). If a dimension doesn't vary between good and bad performers, drop it.
  2. Collect critical incidents. Ask SMEs to describe short, specific examples of behavior they have observed representing exceptional, average, and poor performance on each dimension. Real incidents, not hypotheticals.
  3. Retranslate the incidents. Give the incident pool to a second, independent panel of SMEs and ask them to sort each incident into a dimension. Discard any incident that the second panel doesn't reliably place in the same dimension the first panel did. This retranslation step is what separates BARS from ordinary "let's write some anchors" scales — it is what buys the rater consistency.
  4. Scale the retained incidents. For each surviving incident, ask the second panel to rate the performance level it represents on the numeric scale (usually 1–5 or 1–7). Retain only incidents with tight agreement on the level (low standard deviation across raters).
  5. Assemble the scale. For each dimension, pick one clean anchor per scale point. Write it as a behavior, not an adjective.
  6. Pilot and calibrate. Have a small group of managers rate a handful of real employees or interview recordings using the draft BARS, then debrief. Fix anchors that everyone interprets differently.

Done properly, this is weeks of work per role family. Done improperly (skipping retranslation, or having a single person write the anchors), you end up with a graphic rating scale in a costume.

Advantages and disadvantages of BARS

Advantages:

  • Higher inter-rater reliability than graphic scales: Two managers rating the same person on the same behavior usually agree closely.
  • Anchors serve as coaching language. "You were a 3 on handling objections, here's what a 4 looks like" is a much more useful conversation than "you got a 3 on communication."
  • Legally and procedurally defensible. When a rating is challenged, the anchor is the reasoning.
  • Focuses attention on behavior, which is observable and coachable, rather than personality traits, which are not.

Disadvantages:

  • Expensive to build well. Multiple SME panels, retranslation, calibration. A real BARS is not a template exercise.
  • Rigid. If the role changes materially (new tooling, new customer segment), anchors written a year ago may no longer reflect the actual work.
  • Can encourage "checkbox behavior" if raters over-rely on matching to a specific anchor without weighing the wider context.
  • Doesn't scale well across radically different roles, each role family typically needs its own scale.

Where BARS actually fits (and where it doesn't)

BARS is worth the effort when three things are true at once: performance depends on how work is done and not just outputs; two raters would plausibly disagree without structure; and the rating needs to hold up under scrutiny. Structured interview scoring, promotion decisions, calibration reviews across managers, and coaching conversations all fit. Where BARS is usually overkill: narrowly transactional work where output can be counted directly (packages per hour, tickets closed), roles that change too fast for the anchors to stay current, and small teams where the cost of building a defensible scale outweighs the marginal gain in rating quality.

BARS applied to interview scoring

The strongest use of behaviorally anchored ratings in modern hiring is inside structured interview scoring. Rather than asking a panel to rate a candidate on "problem-solving" (which reduces to a vibe), an interview rubric can define what a 5, 4, 3, 2, 1 look like for a specific problem-solving artefact, how they framed the ambiguity, whether they surfaced trade-offs, how they responded to a counterexample. Fabric's Interview Engine scores candidates against structured, role-specific criteria that behave like a BARS: each score band is defined by observable behavior in the interview transcript and recording, not by a rater's summary impression. The engine screens, scores, and shortlists — the recruiter or panel then decides on the shortlist. It's a signal for your team to weigh, not an automatic reject. That framing is what keeps structured scoring including BARS-style scoring into a defensible input to hiring decisions rather than an unaccountable one.

This article is for informational purposes only. Fabric's Interview Engine screens, scores, and records Round 1 interviews; it does not make the final hiring decision. The recruiter or hiring panel using Fabric remains responsible for all hiring decisions.

Related Posts

FAQ

What is an example of a behavior rating scale?

A behaviorally anchored rating scale for the dimension "handles customer objections" might read: 5 – reframes the objection, addresses the concern, and closes with a next step. 3 – answers the objection factually but does not steer the conversation forward. 1 – becomes defensive, changes the subject, or ignores the objection. Each rating point is tied to a specific observable behavior.

What are the 5 levels of performance rating?

A common five-level performance rating scale is: 5 – consistently exceeds expectations, 4 – exceeds expectations, 3 – meets expectations, 2 – partially meets expectations, 1 – does not meet expectations. On a BARS, each of those five levels is anchored to a described behavior for a specific job dimension, so reviewers rate what the person did rather than what they seemed like.

What is the difference between a behaviorally anchored rating scale and a graphic rating scale?

A graphic rating scale uses general labels or numbers to rate traits or behaviors, often without detailed descriptions, so two raters can pick the same number for very different reasons. A BARS ties each point on the scale to specific observable behaviors, which makes evaluations more consistent between raters and easier to defend when someone questions a rating.

How do you develop a behaviorally anchored rating scale?

The standard process is: identify the performance dimensions that matter for the role, collect critical-incident examples of high and low performance from people who know the role, sort those incidents into rating bands, have a second panel independently re-sort them to validate, and then write the final scale with retained incidents as anchors.

When should an organization use BARS?

BARS is most useful when performance is measured on judgment-heavy behaviors that vary in quality (customer interactions, engineering craft, coaching, negotiation) and when rater consistency matters for interview scoring, promotion decisions, calibration reviews, and any process that has to defend a rating externally.

Can BARS be used for all types of jobs?

In principle yes, but BARS is expensive to build well and pays off most on jobs where behavioral quality varies and matters for client-facing roles, technical roles with judgment components, and manager roles. For narrowly defined production or transactional work, simpler output-based measures often do the job for less effort.

Try Fabric for one of your job posts