EN
Rate my voice

Pitch vs Resonance: What Makes a Voice Sound Feminine or Masculine

Raise the pitch of a male voice to about 200 Hz and a pitch-based test will call it feminine, while a speech model trained on real voices mostly will not. We measured how far pitch alone gets you on 40 real voices, and what changes when resonance is taken into account.

By Daniel Hart, who builds and tests Voice Rater · Published October 2, 2026

What is resonance in voice?

Resonance is how the vocal tract, the throat and mouth above the vocal folds, shapes the sound the folds make. It shows up as formant frequencies, the bands of sound the vocal tract strengthens in each vowel. It is separate from pitch, which is how fast the vocal folds vibrate, measured in hertz. A shorter vocal tract pushes the formants up and makes a voice sound brighter, and a speaker can change its effective length and shape with larynx height, tongue position and lip shape, which is why two people can speak at the same pitch and still sound clearly different.

Voice training communities have long warned that pitch on its own is not enough. Research supports that. In a 2009 listening study (Hillenbrand and Clark, Attention, Perception & Psychophysics 71:1150-1166), shifting both pitch and formants changed the perceived sex of a speaker about 82% of the time, and shifting only one was less effective. We wanted to see what that means for online voice tests, including our own, so we ran the experiment on 40 real voices.

The test: change the pitch, keep the resonance

We took 12 seconds of reading from each of 20 women and 20 men in LibriSpeech, a public speech corpus of audiobook recordings released under a CC BY 4.0 licence, with each speaker's sex listed in its metadata. Then we used Praat, the standard phonetics program, to change only the pitch of every recording. Each man was moved to a median of about 200 Hz and each woman down to about 120 Hz. We used Praat's "Change gender" function with the formant shift ratio left at 1.0, so the formants were not deliberately moved.

Every recording, original and shifted, then went through two analyses. The first is the standard reading on our voice gender detector, which combines pitch and formants with pitch weighted at 68%. The second is an open-source speech model, Common Voice Gender Detection, which was trained to classify voices as female or male and hears the whole signal rather than a few measurements. The model runs in the browser as the optional deep resonance check on that page.

Bar chart: the pitch-based reading was fooled by all 20 pitch-raised men and all 20 pitch-lowered women. The speech model was fooled by 3 men and 2 women
How many shifted voices each method got wrong. Pitch was changed, formants were left as they were. The model was tested on 5-second takes, the length the site records.

What happened

RecordingsPitch-based readingSpeech model
40 original voices, labelled correctly38 of 4037 of 40
20 men, pitch raised to about 200 Hz, called feminine20 of 203 of 20
20 women, pitch lowered to about 120 Hz, called masculine20 of 202 of 20

One difference in setup: the pitch-based reading was run on the full 12-second clips, while the speech model was run on 5-second takes, the length the site records. The model had the shorter clips, so clip length does not explain the gap.

On untouched recordings, both methods did well. The difference appears as soon as pitch and resonance disagree. The pitch-based reading followed the pitch every single time. The speech model kept hearing most of the shifted men as men and most of the shifted women as women, which is consistent with the research finding that moving only one cue is less effective. We did not test human listeners ourselves.

We should be plain about what this says about our own tool. The standard reading on this site leans so heavily on pitch that, in this test, a raised pitch alone pushed it into the feminine range every time. That is why we added the deep resonance check, and why we now say so on the detector page.

Why the formant measurement struggles

If formants matter, why does a measurement that includes them still miss? Because which vowel is being spoken moves the second formant much further than a speaker's sex does, as vowel studies such as Hillenbrand, Getty, Clark and Wheeler (1995, Journal of the Acoustical Society of America 97:3099-3111) show, so the average gap between men and women is small by comparison. Measured with Praat across all 40 readers, the women's median formants averaged 509, 1748 and 2874 Hz for the first three formants, against 441, 1554 and 2620 Hz for the men: roughly 10 to 15% apart, with plenty of overlap. The formants our own detector code estimated separated the groups even less: the second formant averaged 1571 Hz for the women and 1464 Hz for the men, a gap of about 7%.

A few seconds of speech mixes many vowels, so the averages blur. Research studies usually measure formants vowel by vowel, which needs a transcript or a speech recogniser. A trained speech model works from the whole signal instead of a few averaged measurements, and in our test it stayed much closer to the original labels when only pitch was changed. We did not test why.

Shifting resonance alone moves a speech model

We also ran the opposite case on a single 20-second take from one male reader on LibriVox, whose speaking pitch sat around 123 Hz. Raising his formants by 10% while leaving his pitch alone moved the speech model's female probability from 0% to 100%. The pitch-based reading barely moved, from 23% to 25% feminine. Raising only his pitch to about 200 Hz did the reverse: the pitch-based reading jumped to 80% feminine, while the model's female probability stayed at 7%.

One voice is an illustration, not a study, and Praat's formant shift also changes other qualities of the sound. But it points the same way as the 40-voice test: for this model, resonance is not a minor extra. The research above suggests people also rely on both cues together, but we did not test listeners.

What this means if you are training your voice

Everything above works in both directions. In our test, lowering the pitch of 20 women to about 120 Hz fooled a pitch-based reading every time and the speech model only twice. These tools describe how a model hears a recording, not who anyone is, and what you aim for, feminine, masculine or somewhere in between, is your own decision.

Limits of this test

Praat's pitch shift is a well-tested technique, but a shifted recording is not the same as a person who has trained their voice. Real training changes pitch, resonance, intonation and voice quality together, and often unevenly. The speech model was trained on recordings labelled female or male, so it describes how a voice tends to be heard, not anyone's identity, and it was wrong on 3 of the 40 original voices.

LibriSpeech readers are volunteers reading audiobooks aloud, and we did not select for accent, so casual speech and many accents are probably underrepresented. We used 20 voices per group, one pitch target per group and no human listeners. We also build the detector we tested, so treat this as a self-check, not independent validation.

Check your pitch and your resonance separately. Record once, then run the deep resonance check. Nothing is uploaded.

Open the voice gender detector

Related

Male vs female voice frequencyThe real Hz ranges of 40 readers, and where they overlap. Track voice training progressHow big a change has to be before it is more than noise. Voice Hz TestYour median, lowest and highest pitch from one take.