Pitch vs Resonance: What Makes a Voice Sound Feminine or Masculine
Raise the pitch of a male voice to about 200 Hz and a pitch-based test will call it feminine, while a speech model trained on real voices mostly will not. We measured how far pitch alone gets you on 40 real voices, and what changes when resonance is taken into account.
What is resonance in voice?
Resonance is how the vocal tract, the throat and mouth above the vocal folds, shapes the sound the folds make. It shows up as formant frequencies, the bands of sound the vocal tract strengthens in each vowel. It is separate from pitch, which is how fast the vocal folds vibrate, measured in hertz. A shorter vocal tract pushes the formants up and makes a voice sound brighter, and a speaker can change its effective length and shape with larynx height, tongue position and lip shape, which is why two people can speak at the same pitch and still sound clearly different.
Voice training communities have long warned that pitch on its own is not enough. Research supports that. In a 2009 listening study (Hillenbrand and Clark, Attention, Perception & Psychophysics 71:1150-1166), shifting both pitch and formants changed the perceived sex of a speaker about 82% of the time, and shifting only one was less effective. We wanted to see what that means for online voice tests, including our own, so we ran the experiment on 40 real voices.
The test: change the pitch, keep the resonance
We took 12 seconds of reading from each of 20 women and 20 men in LibriSpeech, a public speech corpus of audiobook recordings released under a CC BY 4.0 licence, with each speaker's sex listed in its metadata. Then we used Praat, the standard phonetics program, to change only the pitch of every recording. Each man was moved to a median of about 200 Hz and each woman down to about 120 Hz. We used Praat's "Change gender" function with the formant shift ratio left at 1.0, so the formants were not deliberately moved.
Every recording, original and shifted, then went through two analyses. The first is the standard reading on our voice gender detector, which combines pitch and formants with pitch weighted at 68%. The second is an open-source speech model, Common Voice Gender Detection, which was trained to classify voices as female or male and hears the whole signal rather than a few measurements. The model runs in the browser as the optional deep resonance check on that page.
What happened
| Recordings | Pitch-based reading | Speech model |
|---|---|---|
| 40 original voices, labelled correctly | 38 of 40 | 37 of 40 |
| 20 men, pitch raised to about 200 Hz, called feminine | 20 of 20 | 3 of 20 |
| 20 women, pitch lowered to about 120 Hz, called masculine | 20 of 20 | 2 of 20 |
One difference in setup: the pitch-based reading was run on the full 12-second clips, while the speech model was run on 5-second takes, the length the site records. The model had the shorter clips, so clip length does not explain the gap.
On untouched recordings, both methods did well. The difference appears as soon as pitch and resonance disagree. The pitch-based reading followed the pitch every single time. The speech model kept hearing most of the shifted men as men and most of the shifted women as women, which is consistent with the research finding that moving only one cue is less effective. We did not test human listeners ourselves.
We should be plain about what this says about our own tool. The standard reading on this site leans so heavily on pitch that, in this test, a raised pitch alone pushed it into the feminine range every time. That is why we added the deep resonance check, and why we now say so on the detector page.
Why the formant measurement struggles
If formants matter, why does a measurement that includes them still miss? Because which vowel is being spoken moves the second formant much further than a speaker's sex does, as vowel studies such as Hillenbrand, Getty, Clark and Wheeler (1995, Journal of the Acoustical Society of America 97:3099-3111) show, so the average gap between men and women is small by comparison. Measured with Praat across all 40 readers, the women's median formants averaged 509, 1748 and 2874 Hz for the first three formants, against 441, 1554 and 2620 Hz for the men: roughly 10 to 15% apart, with plenty of overlap. The formants our own detector code estimated separated the groups even less: the second formant averaged 1571 Hz for the women and 1464 Hz for the men, a gap of about 7%.
A few seconds of speech mixes many vowels, so the averages blur. Research studies usually measure formants vowel by vowel, which needs a transcript or a speech recogniser. A trained speech model works from the whole signal instead of a few averaged measurements, and in our test it stayed much closer to the original labels when only pitch was changed. We did not test why.
Shifting resonance alone moves a speech model
We also ran the opposite case on a single 20-second take from one male reader on LibriVox, whose speaking pitch sat around 123 Hz. Raising his formants by 10% while leaving his pitch alone moved the speech model's female probability from 0% to 100%. The pitch-based reading barely moved, from 23% to 25% feminine. Raising only his pitch to about 200 Hz did the reverse: the pitch-based reading jumped to 80% feminine, while the model's female probability stayed at 7%.
One voice is an illustration, not a study, and Praat's formant shift also changes other qualities of the sound. But it points the same way as the 40-voice test: for this model, resonance is not a minor extra. The research above suggests people also rely on both cues together, but we did not test listeners.
What this means if you are training your voice
- Treat pitch as one cue, not the whole picture. If a pitch reading sits in the feminine range but the model's result does not, resonance may be the cue that has not moved yet. Pitch and resonance can change at different speeds, and neither tool can tell you what you should aim for.
- Use two measures that can disagree. On the gender detector, run the deep resonance check after a take. If your pitch cue says feminine and the model says masculine, that gap shows the two cues are not matching yet.
- Keep everything else the same. Same sentence, same room, same distance from the microphone. The next guide on tracking voice training progress shows how much a reading moves between takes on its own.
- Do not force it. Pushing pitch or squeezing the throat to brighten resonance can strain the voice. A speech-language pathologist or experienced voice coach can guide resonance training; a test can only show you the numbers.
Everything above works in both directions. In our test, lowering the pitch of 20 women to about 120 Hz fooled a pitch-based reading every time and the speech model only twice. These tools describe how a model hears a recording, not who anyone is, and what you aim for, feminine, masculine or somewhere in between, is your own decision.
Limits of this test
Praat's pitch shift is a well-tested technique, but a shifted recording is not the same as a person who has trained their voice. Real training changes pitch, resonance, intonation and voice quality together, and often unevenly. The speech model was trained on recordings labelled female or male, so it describes how a voice tends to be heard, not anyone's identity, and it was wrong on 3 of the 40 original voices.
LibriSpeech readers are volunteers reading audiobooks aloud, and we did not select for accent, so casual speech and many accents are probably underrepresented. We used 20 voices per group, one pitch target per group and no human listeners. We also build the detector we tested, so treat this as a self-check, not independent validation.
Check your pitch and your resonance separately. Record once, then run the deep resonance check. Nothing is uploaded.
Open the voice gender detector