ChivoxAI
/english-speech-assessment

English speech assessment and pronunciation scoring for EdTech

Turn learner speech into structured feedback that a tutor, coach or voice agent can use immediately. Chivox evaluates the whole utterance and preserves the phoneme-level evidence needed for precise, helpful correction — without forcing your product to show every score.

Learner practising English pronunciation with phoneme-level feedback on a laptop
Phoneme-level accuracy
Score pronunciation, fluency and stress down to individual phonemes.
Phoneme
diagnostic detail
Word → passage
assessment range
Closed + open
read-aloud and free speaking
Adult English learner reviewing phoneme-level pronunciation highlights on a laptop
/What you can score

From phonemes to open speaking — not just a percentage

Closed reading tasks return pronunciation, fluency, integrity, intonation and stress. Diagnosis kernels flag omitted, inserted, repeated or misread words. Semi-open and open items can score content, grammar, fluency and pronunciation independently.

  • Word scoring with phoneme-level pronunciation accuracy
  • Sentence and paragraph scoring: pronunciation, integrity, fluency, intonation and stress
  • Diagnosis & correction: detect omissions, insertions, misreads and repetitions
  • Semi-open / open kernels: independent scores for content, grammar, fluency and pronunciation
  • Free-speech recognition scoring that maps spoken practice to word-level pronunciation and fluency marks
Language product designer comparing drill, tutor and reading practice flows on a large monitor
/Where it ships

One engine across practice, tutoring and exam prep

The same English scoring layer covers classroom drills, homework, live tutoring, speaking contests and exam preparation — so product teams do not rebuild assessment for every activity type.

  • AI language tutors and pronunciation coaches
  • Pre-class, in-class and homework speaking practice
  • Exam prep for high-stakes English listening & speaking formats
  • Speaking challenges, level checks and contest-style tasks
  • Voice-agent quality checks inside conversational loops
Product engineer mapping assessment scores to learner-facing coaching cues on a whiteboard and laptop
/Integration

Keep pedagogy and UX in your product — Chivox returns evidence

Call through MCP or your existing service workflow. Map scores to your rubrics and lesson logic, then let an LLM explain only the correction the learner needs on this turn.

  • MCP tools or classic SDK / API — same scoring family
  • Stable structured fields for progress views and analytics
  • Hide raw complexity; surface one age-appropriate next step
  • Keep assessment stable while prompts and UI copy evolve
  • Separate low-quality audio retries from pronunciation coaching
//what the english engine returns

Dimensions you can map — not a single percentage.

Closed reading tasks expose pronunciation, fluency, integrity, intonation and stress. Open items add content and grammar so the product can choose which mismatch to coach first.

Phoneme accuracy

Per-sound scores inside each word, so a tutor can isolate /θ/ without rejecting the rest of the sentence.

Fluency

Pace, pauses and overall flow — useful when the words are right but the delivery still needs work.

Integrity

Whether the attempt is complete enough to coach, or should be retried for clipping, silence or missing audio.

Stress & intonation

Word stress and sentence melody for items where a flattened contour is the real problem.

Open speaking

Independent content, grammar, fluency and pronunciation marks when there is no single reference text.

Diagnosis

Omissions, insertions, misreads and repetitions — the error type, not only a lower score.

//item types

One engine from a single word to unconstrained speech.

  • Word & phonics
  • Sentence read-aloud
  • Paragraph reading
  • Diagnosis & correction
  • Semi-open dialogue
  • Open speaking
  • Real-time reading
  • Oral multiple choice
//three product moments

The same evidence, three different English activities.

Keep the learner in the activity. Use the payload to decide whether this turn needs a phoneme cue, a fluency hint, or a content follow-up.

Adult learner practising a minimal pair with phoneme-level feedback01

Word and phoneme drills

Learner experience
One contrast, one retry, while the sound is still in the mouth.
Engine signal
Per-phoneme scores and error type for the target word.
Adult learner reading an English sentence aloud with fluency feedback02

Sentence and paragraph reading

Learner experience
Finish the line, then hear what to fix — not a wall of marks.
Engine signal
Pronunciation, fluency, integrity, stress and word-level rows.
Adult learner describing a picture as an open English speaking task03

Open speaking

Learner experience
Describe, retell or answer — without a script to read.
Engine signal
Content, grammar, fluency and pronunciation scored independently.
/A real product moment
Illustrative walkthrough

From “think” to one correction the learner can act on

Imagine a B1 learner reading “I think the train leaves at three.” The sentence is understandable, but the first sound in “think” is consistently replaced with /s/. The product does not need to interrupt the whole sentence or show six scores.

Learner says
“I sink the train leaves at three.”
Attempt → evidence → next action
01
Target
think · /θɪŋk/

Expected voiceless dental fricative at the start.

02
Detected
/θ/ scored 54

The remaining sounds and sentence rhythm are usable.

03
Priority
Coach one sound

Pronunciation is more useful to address than pace on this turn.

Tutor can respond
“Almost there. For think, let a little air pass over the tip of your tongue: th-ink. Try just the word once, then we’ll put it back in the sentence.”

Implementation note. The score identifies the location; your lesson rules and tutor voice decide the explanation, retry length and success threshold.

/show-not-tell

See the phoneme, not only the score

Word and phoneme rows preserve the exact evidence a tutor needs to explain an error, while fluency fields help the product decide whether to coach pace, pauses or pronunciation first.

result.json · English
{
  "overall": 85,
  "pron": 88,
  "fluency": { "overall": 78, "speed": 65, "pause": 2 },
  "details": [{
    "char": "think",
    "score": 72,
    "phone": [
      { "phoneme": "θ", "score": 54, "dp_type": "mispronounced" },
      { "phoneme": "ɪ", "score": 86, "dp_type": "normal" },
      { "phoneme": "ŋk", "score": 83, "dp_type": "normal" }
    ]
  }]
}
/workflow

From spoken attempt to one useful correction

Keep exam-grade detail behind the scenes while the learner receives a clear, product-specific next step.

  1. 01
    Step 1
    Capture learner speech

    Live stream or a complete recording, with audio-quality checks first.

  2. 02
    Step 2
    Score closed or open item type

    Word, sentence, paragraph, diagnosis, or open speaking — pick the kernel that matches the activity.

  3. 03
    Step 3
    Return phoneme & dimension evidence

    Overall, fluency, integrity, stress and per-phoneme rows stay in the payload.

  4. 04
    Step 4
    Coach one clear correction

    Your product chooses the wording, retry and what the learner actually sees.

/questions

Common implementation questions

What English item types does the engine cover?

Word, sentence and paragraph reading; diagnosis and correction on follow-read tasks; semi-open and open speaking (for example picture talk, oral composition, story retell); and free-speech recognition scoring for less constrained practice.

Is this the same technology used in exams?

Yes. The English scoring engine belongs to the same exam-grade family Chivox has used for high-stakes English listening and speaking assessment for more than a decade — adapted here for product, tutor and agent integrations.

Does the learner have to see every score?

No. Most products translate the detailed response into a smaller set of coaching cues — one sound, one phrase, or one content tip — while keeping the full payload for analytics and LLM reasoning.

How is open speaking different from read-aloud scoring?

Read-aloud tasks compare speech to a reference text. Semi-open and open items can score content, grammar, fluency and pronunciation independently, so picture talk, retelling and conversation practice are not forced into a reading rubric.

/contact

Let’s build a better speaking experience together.

Tell us what you’re building and whether you need SDK/API, MCP, or function calling. We’ll reply within one business day with pilot credits, pricing, or a deployment plan.