Real Voices vs Synthetic Tones: What Jitter, Shimmer and CPPS Actually Show
A synthetic test tone scores almost perfectly on the classic voice measures. Real people never do. Here is the comparison, measured, and the calibration mistake it led us into.
Jitter and shimmer are the two voice measures you will meet first in almost any article about voice quality. Jitter is the small variation in length from one vocal fold cycle to the next. Shimmer is the matching variation in loudness. Both are standard in speech science, and both are easy to misread.
The problem is that the published "normal" ranges were defined on sustained vowels, where a speaker holds one steady "ahh" for several seconds. Online voice tests, including ours, record ordinary speech. This guide shows what happens when you apply vowel-based expectations to speech, using three real readings and one synthetic tone.
The recordings
Reader A and Reader B are healthy adult voices from public-domain LibriVox audiobooks, each 20 seconds of continuous reading. Reader C is a rougher voice on a lower-quality recording. The synthetic tone is a 20-second computer-generated voice-like signal we use for automated testing, with perfectly regular cycles.
All four were analysed with the same code the site runs in your browser. The table below lists the raw measurements for the full 20 seconds of each.
| Recording | Jitter | Shimmer | CPPS | Score |
|---|---|---|---|---|
| Reader A | 4.01% | 0.99 dB | 6.66 dB | 7.9 |
| Reader B | 2.52% | 0.99 dB | 6.73 dB | 8.0 |
| Reader C (rough) | 3.20% | 1.01 dB | 2.91 dB | 6.7 |
| Synthetic tone | 0.00% | 0.00 dB | 7.35 dB | 7.6 |
Finding 1: real speech has far more jitter than the textbook says
Clinical guidance for sustained vowels usually treats jitter under about 1% as normal. Both healthy readers here are well above that, at 2.52% and 4.01%. They are not unhealthy voices. In connected speech, the pitch rises and falls with every phrase, and cycle-by-cycle comparison counts those intended changes as if they were instability.
Our earlier calibration round on eleven public-domain readings showed the same pattern. Jitter ranged from 2.5% to 8.2% and shimmer from 1.0 to 2.4 dB, in voices that sounded perfectly ordinary. A tool that grades speech against vowel thresholds will mark almost every real person down.
Finding 2: the rough voice is not where jitter says it is
Reader C is clearly the roughest recording to the ear, yet its jitter of 3.20% sits between the two healthy readers. Jitter simply does not separate these voices in running speech. CPPS does: Reader C measures 2.91 dB against 6.66 and 6.73 dB for the healthy readers, less than half.
CPPS, the smoothed cepstral peak prominence, measures how clearly the harmonic structure of the voice stands out from noise across the whole spectrum. It is widely recommended for connected speech, which is why it holds up where jitter does not. It is now the main clarity measure in our score, and jitter carries less weight than it once did.
Finding 3: a perfect score on jitter and shimmer means a fake voice
The synthetic tone reads 0.00% jitter and 0.00 dB shimmer. No human voice does that, because vocal folds are living tissue and every cycle is slightly different. If a tool shows near-zero values for these two measures on your recording, either the recording was heavily processed or the tool is not measuring cycle-level detail at all.
The synthetic tone also scores a high CPPS of 7.35 dB, because its harmonics are perfectly regular. That combination, very clean and perfectly steady, is a fingerprint of generated audio rather than a sign of a great voice.
The mistake this caused us
The first versions of Voice Rater were tested mostly on synthetic signals, because they are convenient and repeatable. Every check passed. Then testing with real people showed results coming out about 14 years older than the speaker's actual age, which is the kind of error that makes a tool impossible to trust.
Running the eleven real readings through the old version explained it. Healthy readers scored between 4.1 and 5.8 out of 10, with a median estimated age of 53, while the synthetic tone scored 7.2 and was judged 26. The thresholds had been tuned on a signal that has no natural variation, so real voices looked damaged by comparison. After switching the main measure to CPPS and recalibrating on real speech, the same healthy readers score 7.1 to 8.0, with a median age of 35.
The lesson is now a fixed rule in our testing. Real recordings are the baseline, and the synthetic tone is kept only as a check that it never outscores a healthy human. At 7.6 it currently sits below Reader B's 8.0 and just under Reader A's 7.9.
How to read your own numbers
- Jitter of 2% to 8% in speech is ordinary. On its own it says little about voice health.
- Look at CPPS for clarity. In this test the two healthy readers measured about 6.7 dB and the rough recording 2.9 dB. A gap of several decibels is what separates a clear voice from a strained one.
- Zero jitter and zero shimmer are a warning. They point to generated or heavily processed audio, not to a flawless voice.
Every result page on the site has a raw measurements panel, so you can compare your own figures with the table above. For any lasting hoarseness or change in your voice, a speech-language pathologist can run proper clinical measurements on sustained vowels.
See your own jitter, shimmer and CPPS. Open the raw measurements panel after any recording.
Rate your voice