EN
Rate my voice

How Voice Rater Works

The complete signal chain, the thresholds behind every sub-score, and the one measurement I had wrong for weeks without noticing.

Voice Rater runs every calculation inside your browser tab. No audio is uploaded, and there is no backend that could receive it if there were. This page documents what happens between pressing record and seeing a number, because a score you cannot audit is worth about as much as a random one.

None of this is novel signal processing. Every technique here is standard in speech science, and the reference implementation everyone uses is Praat. What is unusual is doing it client-side: the normal way to build this is to send a WAV file to a server running Praat, which means hosting costs, a queue, and the user handing over a recording of their own voice. The whole exercise here was seeing how far the Web Audio API gets on its own.

1. Capturing the audio

Three getUserMedia constraints have to be switched off, and this is the single most common way to get meaningless results:

Diagram of the seven analysis steps: capture, framing and gates, pitch, jitter and shimmer, clarity, formants, and the weighted 1 to 10 score
The seven analysis steps described below. Every one of them runs inside your browser tab.

PCM is collected through a ScriptProcessorNode. That node is deprecated and AudioWorklet is the modern replacement, but it proved unreliable on several iOS versions, and ScriptProcessorNode remains the only method returning continuous non-overlapping frames everywhere. One catch: in some browsers the callback is never scheduled unless the node is connected to the destination, so it is routed through a gain node set to zero. Connected, but silent.

The sample rate is whatever the device reports. Nothing in the chain assumes 44.1 kHz.

2. Framing and the two gates

The recording is cut into 40 ms frames with a 20 ms hop. Forty milliseconds is a compromise: it holds at least two complete periods even for a 60 Hz bass voice, which autocorrelation needs, while staying short enough that the vocal tract has not visibly changed shape within the frame.

Two gates then discard frames that would poison the statistics:

GateThresholdWhy
Loudness−45 dBFSRoom tone and breath between words would otherwise be analysed as if they were voice
PeriodicityNSDF peak < 0.55Unvoiced consonants (s, f, sh) have no fundamental frequency to measure

The proportion of frames that survive the second gate becomes the delivery flow sub-score. A take that is mostly breath and pauses genuinely scores lower there, which is the intended behaviour. Frames rejected by the loudness gate are not thrown away: a pause holds your microphone and your room and nothing else, which is what section 5 measures.

3. Fundamental frequency: the McLeod method

F0 comes from the normalised square difference function, the core of the McLeod pitch method, searched between 60 and 500 Hz. Plain autocorrelation was the first thing tried and it fails the same way it always fails: it happily locks onto the octave below, because a signal correlates strongly with itself at twice its period. Normalising by the frame energy at each lag, then taking the first sufficiently strong local maximum rather than the global one, mostly removes that failure mode.

Two details matter more than they look:

4. Jitter and shimmer: two mistakes worth reading

This is the part I got wrong twice. The first error is, I suspect, why other browser-based voice tools feel like they return random numbers.

The obvious implementation, once you already have a per-frame F0 track, is to measure how much F0 varies from frame to frame and call that jitter. It is wrong by an order of magnitude. A 40 ms frame contains five or six glottal cycles, and the autocorrelation has already averaged across all of them. The perturbation you are trying to measure is the thing the estimator smoothed away before you saw it.

Concretely: that version reported around 0.09% where the true figure was closer to 1%. Every voice scored full marks on stability, and a deliberately hoarse recording came out ahead of a clean one.

The correct approach is the textbook clinical definition, in the time domain, cycle by cycle:

At least five detected cycles are required, and only on locally steady ground: an 80 ms window is placed where consecutive frames agree on F0 to within 3%, hunting inside a sentence for stretches that behave like a sustained vowel. Across eleven people reading aloud, jitter comes out 2.5% to 8.2% and shimmer 1.0 to 2.4 dB, far larger than a clinic would quote. Which is the second mistake. Clinical thresholds assume a sustained vowel, but intonation, stress and word transitions move the fundamental on purpose, so a cycle-by-cycle comparison counts every deliberate movement as perturbation. With 1% as the healthy boundary, every one of them scored in the fours and fives on a page promising the sevens, while a synthetic tone scored 7.2.

5. Voice quality: cepstral peak prominence

What replaced those thresholds is smoothed cepstral peak prominence, which speech pathology adopted for this situation: it asks how closely the spectrum resembles a clean harmonic comb, a question intonation does not disturb. Each frame gets a log power spectrum, then a second FFT turning it into a cepstrum, then a peak search across the quefrency band matching the 60 to 500 Hz pitch range. Peak height above a least-squares regression line through the same band is the prominence, smoothed across five neighbouring frames and eleven quefrency bins; the median over the take is reported. It runs on a 16 kHz downsampled copy in 40 ms frames with a 25 ms hop, which costs 0.05 dB and takes under half a second instead of nearly three.

Absolute values are not portable between implementations, so the literature's boundary is the wrong number to copy. Calibration uses this implementation's own readings, listed in section 7: full marks at 9.5 dB, zero at 3.2 dB. Noise is the complication. A raised floor buries harmonic structure and pulls the prominence down about 3.2 dB between a clean recording and one at 10 dB signal-to-noise, reading as a rougher voice, so the score compensates. That means knowing how clean the recording is, and the old harmonic-to-noise ratio cannot say: on connected speech it barely moves between speakers, pinning the cleanliness figure near 10% for everyone and switching the compensation off in practice. What replaced it is a true signal-to-noise ratio, the median level of the voiced frames over the median level of the frames below the loudness gate, where twelve decibels counts as dirty and thirty-four as clean. With no pauses at all, as happened twice, the quietest tenth of all frames stands in.

A side-by-side measurement of three human readings and one synthetic test tone, and the calibration mistake that synthetic signals led to, is in real voices vs synthetic tones.

6. Formants via LPC

F1 and F2 come from a 16th-order linear prediction filter solved with Levinson-Durbin recursion, on a copy of the signal downsampled to 11,025 Hz. The downsampling is the important choice: formants of interest all sit below about 5 kHz, and running a 16-pole model against a 24 kHz band spends most of its poles describing high-frequency content that carries no formant information at all.

The classic way to extract formants from LPC coefficients is to find the roots of the prediction polynomial. That is mathematically cleaner and numerically fragile: root finding on a 16th-order polynomial in JavaScript, on arbitrary consumer microphone input, produces occasional nonsense. Instead the LPC spectral envelope is evaluated on a frequency grid and scanned for local maxima. Less elegant, considerably more stable, and the resolution is more than sufficient when the goal is telling F1 around 500 Hz from F1 around 700 Hz.

Formants are the second line of evidence for the masculine-feminine read-out. Pitch alone gets it wrong in the cases people care about most: a man in falsetto and a woman with a low speaking voice can share an F0, but their vocal tract resonances usually differ, because resonance depends on tract length, which vocal fold frequency does not change.

7. From measurements to a 1–10 score

Seven sub-scores, each normalised to 0–1 and weighted, the weights differing between speaking and singing. Range counts for 6% when you talk and 20% when you sing, because a monotone speaking voice is normal and a monotone singing voice is not. Clarity carries the most weight in both, 32% speaking and 26% singing, being the sub-score that tracks the voice and not the performance.

Bar chart comparing sub-score weights for speaking and singing: clarity 32 and 26 percent, range 6 and 20 percent
Sub-score weights for speaking and singing, exactly as set in the scoring code.
Sub-scoreMeasured fromFull marks atZero atSpeaking weight
ClarityCepstral peak prominence9.5 dB3.2 dB32%
Pitch stabilityJitter2.0%7.0%16%
Volume controlShimmer0.8 dB2.8 dB12%
ExpressivenessPitch spread3.2 semitonesfar from target14%
Tone brightnessSpectral centroid1900 Hzfar from target12%
Delivery flowVoiced frame share80%15%8%
RangeSemitone span10 semitones1 semitone6%

Two of these are band scores with a sweet spot: expressiveness and tone brightness both peak at a target and fall away in either direction. A completely flat pitch contour reads as reciting; an enormous one reads as unsteady, and earns no expressiveness credit for its size.

Why the full-marks end is set at professional level

The first calibration set each threshold at a healthy-voice pass mark and lifted the curve with a 0.62 exponent. That crowded trained and ordinary voices alike into the top of the scale: an average person scored 9.0, a disordered voice still managed 5.0, and the result is indistinguishable from not measuring.

The fix was to move the full-marks end up to professional standard and raise the exponent to 0.9. A purely linear mapping was tried too, and it pushes healthy voices into the low sixes, which reads as an insult.

The exponent survived the September rewrite. The clinical thresholds it was paired with did not:

Measured on connected speechReading
CPPS, healthy adults reading aloud (our tests)6.7 – 7.0 dB
CPPS, low-quality recording3.1 dB
CPPS, synthetic voice6.8 dB
Jitter, eleven readers2.5 – 8.2%
Shimmer, eleven readers1.0 – 2.4 dB
Harmonic-to-noise ratio, eleven readers4.6 – 9.5 dB
Old total score, eleven readers4.1 – 5.8

That table forced the rebuild: clarity moved onto cepstral peak prominence, jitter and shimmer were rescaled, both losing weight. The lesson for anyone calibrating something similar: constructed metric objects fed to the scoring function only prove the curve behaves. Synthetic audio proves less: perfectly periodic, it outscored every human in the set.

8. Pitch accuracy: a separate measurement

Everything above scores how a voice behaves. It cannot say whether you sang the right note, because it has no idea what note you were aiming for. The pitch accuracy test on the Rate My Singing page closes that gap by supplying the target itself: five reference notes play, you sing them back, and each one is compared against a known frequency.

The detector is the same McLeod implementation from section 3. What differs is everything around it.

The reference plays before recording, never during. Sung-along reference tones leak into the microphone, and the detector has no way to tell which of the two fundamentals belongs to the singer. Playing first and recording after costs six extra seconds and removes the ambiguity entirely.

Only the middle half of each note is measured. A note is held for 1.2 seconds, and the window from 25% to 75% is what gets analysed. Sliding up to pitch at the start and tapering off at the end are normal singing behaviours. Including them would penalise the technique that makes singing sound musical. The frames inside the window are reduced to their median, so a single mistracked frame cannot drag the result.

Deviation is reported in cents, and octaves are normalised away. The error is 1200 × log₂(sung / target), wrapped into ±600 cents. Without that wrapping, a bass following the higher reference range would be marked 1200 cents wrong on every note, when singing the pattern an octave down is musically correct and deserves full marks. The test compares pitch class, and the octave a note is sung in has no effect on the result.

The score curve has its inflection points set to the same thresholds used for the per-note verdicts, so the total and the individual labels cannot contradict each other:

Average deviationVerdictScore
≤ 15 centsSpot on9.2 – 10
15 – 30 centsClose7.5 – 9.2
30 – 50 centsSlightly off5 – 7.5
> 50 centsOffbelow 5

Fifteen cents is roughly where the error stops being audible to most listeners, and fifty is half a semitone. A note that fails the three-frame minimum is reported as not heard and scores nothing. Quietly dropping it would let someone sing one note correctly, hum through the rest, and get a perfect result.

One systematic effect is separated out from the average: if all five deviations lean the same way, the test says so. A consistent lean is usually mechanical, and fixable. Flat throughout often traces to breath support, sharp to pushing. Errors scattered either side of the target point at ear training instead. Keeping the sign on each deviation before averaging is what makes that distinction visible.

9. Reading the results in context

Being specific about the limits is part of the method, and belongs here as much as the thresholds do.

10. Checking any of this yourself

Every result page has a raw measurements panel showing your F0, cepstral peak prominence, jitter, shimmer, harmonic-to-noise ratio, spectral centroid, F1, F2 and voiced-frame share behind your score, so you can see which measurement drove the result.

The analysis code is a single unminified JavaScript file served from this site. Open the developer tools network tab and read voice.js. The functions above appear in the same order as this page. The privacy claim is checkable there too: watch the network tab while recording and analysing. Nothing is sent, because there is nowhere to send it.

Try it against your own voice. Five seconds, no account, and every number above is shown to you.

Rate your voice free

Where to go next

Rate My VoiceAll seven sub-scores on your speaking voice, explained one by one. Voice Gender DetectorThe F0 and formant evidence behind the masculine-feminine read-out. AboutWhy this exists, and what it deliberately does not do.