How Fabric scores communication
Every Fabric interview gives candidates a spoken English score out of 10. This page will help you understand how it works.
What is Fabric's communication score?
After every Fabric AI interview, an AI agent listens to the full recording and rates the candidate's spoken English out of 10.
- It answers one question: how easy is it to have a spoken conversation with this person?
- It's a score on our own scale. It is not an IELTS or CEFR score, though we reference both later in this guide so you have familiar anchors.
- The number is the average of five dimensions we listen for (clarity, flow, vocabulary, grammar, and structure), each covered in detail further down.
- A 6 means functional professional communication. That's the middle of our scale and a perfectly good score, not a warning sign.
Reading the score for your role
There's no single passing score. A number that's fine for a backend engineer can be a problem for an account executive, so here's how we'd read it for three broad families of roles.
One thing before the tables: these bars already bake in how strict our scale is. Someone with an IELTS 7 certificate will usually score in the high 6s to low 7s with us, which is exactly where the top rows below sit. So read the bars as written, without adding your own safety margin.
Engineering, technical, and internal roles
Engineers, analysts, data, ops, and finance. This covers everything from internal work where most collaboration happens in writing to international teams that talk all day.
Clear pass. Clears the bar for any engineering role, including client-facing technical work. Explanations just land.
Typical hire. Your typical competent hire. Perfectly fine for standups, reviews, and demos on collaborative or international teams. You'll hear some hesitation, but it never gets in the way.
Depends on the team. Fine for internal, code-heavy roles where most collaboration happens in writing. If the team talks all day, listen to the recording before moving forward.
Listen first. At this level, listening takes real effort. Pull up the recording and judge against the actual role rather than going off the number alone.
Worth keeping in mind
- A gap of 6.4 versus 6.8 is real but small. Once the gap grows past a full point, you can rely on it.
- This score only covers spoken communication. It says nothing about writing, and nothing about whether the person can do the job, so pair it with the interview's content evaluation.
- When you're on the fence, open the per-dimension feedback. The quoted evidence shows you exactly what drove the number.
What each score sounds like
Numbers only get you so far, so here's what the bands actually sound like. Each one comes with a short excerpt from a real interview, with anything identifying removed.
You would notice this person in any room. The vocabulary is wide and natural, the storytelling is effortless, and the only flaws are rare slips under pressure.
Precise and easy to listen to across the board, with crisp articulation, complex sentences that stay under control, and clear signposting. What separates this from 9+ is the occasional visible seam, like a restart or a small slip.
Comfortably professional. Everything gets across, but one or two habits are noticeable, maybe filler chains, small recurring grammar slips, or uneven pacing. You register them without ever being slowed down.
Still solidly professional, just with the rough edges spread a bit wider. Both the delivery and the structure show some wear, but you never have to work to follow the meaning.
This is the typical competent candidate. They get everything across, but you can hear the effort throughout: restarts mid-sentence, plenty of fillers, simple connectors. None of it ever blocks the conversation.
Workable, but you feel the effort. Simple topics go fine, while longer explanations start to strain, with persistent word searches, repeated fragments, and grammar that sometimes makes you reconstruct the sentence.
You have to work to follow along. Sentences break midway, grammar errors force you to fill in gaps, and answers mostly hang together with “and... so...”. If a candidate scores here, listen to the recording before deciding anything.
No sample for this band.
What the score listens for
The score is an average of five things, each rated out of 10. Click through them to see what we listen for, and what strong and weak actually sound like.
What we listen for
How much work it takes to catch the words. We pay attention to articulation, word stress, and whether the ends of words come through.
Strong sounds like
You catch every word the first time.
Weak sounds like
Words get swallowed or blurred, and you find yourself wanting to rewind.
How the score works
The scale
Each of the five dimensions gets a score out of 10, and the overall score is simply their average. A 6 means functional professional communication. That's the middle of our scale, and it's a perfectly good score, not a warning sign.
The evidence rule
The agent has to back up every claim it makes, good or bad, by quoting the candidate's actual words with timestamps. If it can't point to a moment in the recording, the claim doesn't go in.
Repeatability
Score the same recording twice and you'll usually land within about a third of a point, which is roughly how much two careful human raters drift. A gap of more than a full point between candidates is a real difference.
How the score maps to CEFR and IELTS
The 0 to 10 score is what we generate. The CEFR label you see next to it is simply read off this table. Keep in mind that these correlations are approximate, especially the IELTS column. No official equivalence exists between Fabric, CEFR, and IELTS scores.
The CEFR and IELTS columns are approximate reference points, not official conversions.
How we calibrated it
We set the band boundaries using the official CEFR descriptors from the Council of Europe, then sanity-checked them against the EF English Proficiency Index (2025). The median candidate on our scale lands exactly where EF puts the median Indian professional, which is upper B1. Across production interviews that shakes out to 51% B1, 30% B2, and 3% C1, and rescoring the same recording stays within about a third of a point.
How it compares to human examiners
We also ran our scorer over Speak & Improve, a Cambridge University dataset of speaking tests marked by human examiners. Our ranking of speakers matched theirs with a rank correlation of 0.77, which is roughly how well two trained human raters agree with each other, and 86% of speakers landed within one CEFR level of their examiner's mark. Where we differ from Cambridge, we differ in one consistent direction: about 0.7 points stricter on average, a little more so at the top of the scale.
Why we score stricter than test certificates
The Cambridge benchmark confirms that our scores sit below traditional test marks. That's a choice rather than a flaw, and it comes down to three things.
A harder setting
A live, unscripted interview asks more of a speaker than a prepared language test. Candidates have to sustain long answers, reason on the spot, and use professional vocabulary under pressure.
Language only
Speaking tests partly reward completing the task. We ignore content and task success entirely, so there's no credit to soften the language judgment.
Reserved top bands
We save 8+ for speech that stays exceptional through an entire real interview. We'd rather a high score under-promise than over-promise, since a hiring decision rides on it.
- Treat the Fabric score as the primary output. The CEFR label is just a familiar anchor on top of it, and because the mapping runs strict at the top end, someone holding an IELTS 7 certificate will often land in our 6.5 to 8 band.
- No paired Fabric-versus-IELTS dataset exists, and the two tests measure overlapping but different things, so don't treat the IELTS column as an official conversion.
- Position within a band matters too. A 6.8 and a 7.9 share a label but read very differently in the per-dimension feedback.
What never affects the score
A lot of things that sway human interviewers have no effect on this score. Here's the full list.
Answer correctness or depth
The interview itself evaluates what candidates said. This score only cares how they said it, so a wrong answer in excellent English scores the same as a right one with the same delivery.
Skipping or admitting a gap
Saying “I'm not familiar with that, can we move on?” gets scored as a sentence like any other. There's no penalty for admitting a gap, and no bonus for being polite about it.
Thinking time
A pause spent working out a problem is thinking, and thinking is free. Only pauses spent hunting for a word or a grammar form count against fluency.
Accent
An accent is never an error, and never labeled good or bad. What matters is whether the listener has to work, and most accents cost nothing. Regional grammar patterns from established varieties of English count as minor at most.
Answer length
A short, well-formed answer can earn any score on the scale. Nobody earns points just for talking longer.
Technical knowledge
Using jargon well, or misusing it, tells you about domain knowledge rather than language skill. It moves nothing.
Audio and circumstance
Background noise, choppy calls, glitches, and asking to repeat something unclear are never held against the candidate.
Enthusiasm and attitude
Confidence, energy, and effort are great to see, but they aren't language skills, so they aren't allowed to appear anywhere in this assessment.
A separate read on mother tongue influence
Alongside the score, every interview also comes with an MTI analysis. It describes how a candidate's first language shows up in their English, and it stays completely separate from the score.
What it covers
- The likely accent family, which sounds tend to get swapped, and the rhythm and stress patterns carried over from the first language.
- How strong that influence is. What it never does is call the accent good or bad.
- There's no score attached. It's a description of how someone sounds, not a judgment of their skill.
Why it never touches the score
If accent fed into the score, candidates would get marked down for how they sound rather than for how well they communicate, and the score would stop being trustworthy. And if the MTI analysis came with a score, a description of someone's accent would quietly turn into a judgment of it. Professional language assessors have drawn this line for a long time, and we follow it.
In practice, the score tells you whether spoken communication will work in the role, and MTI gives you context about the speech itself. For voice-heavy roles that context is genuinely useful, and nobody gets penalized for their accent along the way.