Which whisper.cpp model do you actually need on a Mac?
We ran tiny.en, base.en, small.en and medium.en over the same four-minute interview, clean and noisy. small.en was within three words of medium.en at a third of the time, and a jargon prompt mattered more than model size.

small.en is the one to run: on an Apple M3 it transcribed a four-minute
technical interview at 1.37% word error rate, 17.6 times faster than real
time, and medium.en beat it by three words while taking three times as
long. The bigger lever was not model size at all. A one-line --prompt
listing the jargon took base.en from 2.74% to 1.22%, better than small.en
without one.
Our team runs whisper.cpp for live transcription in an internal tool that
listens to technical interviews on an M3 laptop. It runs base.en. The
question was whether that was a false economy. Forum advice splits into "just use large-v3" and "tiny is useless",
and neither side posts numbers. So we measured whisper.cpp model size against
accuracy and speed on the same recording.
- On clean speech every model was usable. Word error rate went 3.66%,
2.74%, 1.37% and 0.91% from
tiny.entomedium.en. That is 24, 18, 9 and 6 wrong words out of 656. - Steady noise barely separated
small.enfrommedium.en. With pink noise at 5 dB SNR they were 2.59% and 2.44%: one word apart. - Background voices did. With three other talkers at 5 dB,
small.enhit 9.45% andmedium.en4.88%. At 0 dBsmall.entranscribed the background talkers too and scored 104.27%. - Every model misheard Redis and idempotency without a prompt. With one,
base.enupward got all of them on clean audio.tiny.enignored it. - Metal matters more than the model.
base.entook 7.09 s on the GPU and 39.75 s on four CPU threads for the same file.
What audio did we test on?
A transcript you did not write cannot be scored. So we wrote one: a 646-word mock interview about a payments platform, with the jargon a technical interview is full of: Kubernetes, PostgreSQL, Redis, idempotency keys, Prometheus, Grafana, Kafka, "p99 latency", "99.9 percent", "200 thousand orders". Twenty turns, ten questions and ten answers.
We voiced it with the Kokoro TTS model locally, interviewer as af_heart and
candidate as am_michael, with 0.6 s between turns. The result is 256.1
seconds of 16 kHz mono audio.
Synthetic speech is cleaner than a person. No ums, no restarts, no crosstalk, no laptop fan, even diction. The absolute error rates here are optimistic. The comparison between models on identical audio is the point.
For noise we mixed two kinds into the clean file with a small script. Pink
noise from ffmpeg's anoisesrc stands in for a fan or air conditioning.
"Babble" is the interview itself, rotated by 61, 127 and 193 seconds and
summed: three overlapping copies of the same two voices, standing in for other
people talking nearby. Because the babble talkers say the same script, a model
that follows the wrong voice can still score some words right by accident.
We scaled the noise so that 10 * log10(P_speech / P_noise) hit the target exactly, with both powers taken
as the mean square over the whole clip. The script printed the achieved SNR
back: 20.00, 10.00, 5.00 and 0.00 dB, nothing clipped. Because the speech power
includes the pauses between turns, the SNR during actual speech is a little
higher than the label.
Hardware and versions: Apple M3, 16 GB, macOS 26.4.1, whisper.cpp 1.9.2 from
Homebrew (whisper-cli, ggml 0.18.1), ffmpeg 7.1.1. The models are the ggml
English-only files from ggerganov/whisper.cpp on Hugging Face. whisper-cli
ran with its defaults: Metal on, 4 threads of 8, beam size 5, best of 5,
language forced to English. We skipped the large models. large-v3 is not
English-only, its file is 3.10 GB (twice medium.en's 1.53 GB), and medium
was already the slow end of what live transcription can use.
We checked the headline figures a second time, from scratch: the same script
voiced again by Kokoro, the same models, a separate run. small.en scored
1.37% clean and 0.76% with the prompt, and base.en with the prompt 1.22%,
all identical. base.en without a prompt came back at 20 errors instead of
18. Kokoro does not produce bit-identical audio twice (the second file was
0.4 ms longer), and two words is the size of that wobble. The small.en run
also took 11.2 s rather than 14.6 s, on a less busy machine, so treat the
speeds below as conservative.
How accurate is each whisper.cpp model?
Word error rate is substitutions plus deletions plus insertions, divided by the number of reference words. Both sides go through the same normaliser first: lowercase, punctuation dropped, digits spelled out ("200,000" and "200 thousand" both become "two hundred thousand"), so a formatting choice is not counted as a mistake. The reference normalises to 656 words. One wrong word is 0.15 percentage points.
| Model | File | Clean | Pink 20 dB | Pink 10 dB | Pink 5 dB | Pink 0 dB | Babble 10 dB | Babble 5 dB | Babble 0 dB |
|---|---|---|---|---|---|---|---|---|---|
| tiny.en | 78 MB | 3.66% | 4.12% | 4.73% | 5.64% | 19.66% | 7.47% | 23.02% | 81.71% |
| base.en | 148 MB | 2.74% | 3.35% | 3.96% | 5.34% | 11.74% | 5.03% | 21.04% | 80.79% |
| small.en | 488 MB | 1.37% | 1.37% | 1.98% | 2.59% | 6.25% | 2.74% | 9.45% | 104.27% |
| medium.en | 1.53 GB | 0.91% | 1.37% | 1.68% | 2.44% | 3.51% | 1.68% | 4.88% | 62.96% |
One run per cell. Each model gave the same transcript on all six of its clean runs, so the decoding is deterministic here.
On clean audio the gap from base.en to small.en is nine words. Most of them
are the same few terms said more than once. base.en wrote "Postgres SQL" for
all four PostgreSQL, "riddies" for all three Redis, and "iampotency" for
idempotency. Then come homophones: "right" for "write", "too" for "two",
"you" for "use". All four models lost the decimal in "99.9 percent", writing
"99, 9%".
Kubernetes, Prometheus and Kafka came through every model on clean audio. Not all jargon is hard. The words that broke were the ones that sound like other English words.
Does background noise change which model you need?
Steady noise, less than we expected. From clean down to pink noise at 5 dB,
base.en lost 17 more words and small.en 8 more. medium.en lost 10 more,
and ended one word ahead of small.en. The case for medium.en over
small.en on a noisy fan-filled room is one word in 656.
At 0 dB, where noise is as loud as the speech, the ranking finally spreads out: 19.66%, 11.74%, 6.25% and 3.51%.
What happens when the noise is other people talking?
This is where size paid. At 10 dB babble, tiny.en doubled its clean error
rate and medium.en stayed under 2%. At 5 dB, tiny.en and base.en were
both above 20% while medium.en was at 4.88%. Whisper has to decide which
voice is the speaker, and the larger decoders decided better.
At 0 dB every model failed, but not the same way. We split the errors:
| Model at babble 0 dB | Substitutions | Deletions | Insertions | Words output (normalised) |
|---|---|---|---|---|
| tiny.en | 281 | 140 | 115 | 631 |
| base.en | 366 | 96 | 68 | 628 |
| small.en | 426 | 14 | 244 | 886 |
| medium.en | 277 | 8 | 128 | 776 |
small.en wrote 886 words for a 656-word interview. It did not miss the
speaker; it transcribed everyone, splicing the background talkers' sentences
into the answer. Near the start, after the question about traffic, its
transcript already reads "Why do you want to save your current role?" That
question is asked at the end of the interview. A WER above 100% is what this
looks like. medium.en also wrote more words than were spoken, 776, and
scored 62.96%. For a live tool this matters more than the number. A wrong word
is visible. A plausible sentence from someone else's conversation is not.
Why does whisper.cpp misspell Redis and PostgreSQL?
Because nothing told it those words were likely. whisper-cli --prompt feeds
text to the decoder as if it came before the audio, so the vocabulary is
already in context. We used one line:
whisper-cli -m ggml-base.en.bin -f clean.wav -l en \
--prompt "Interview notes: Kubernetes, PostgreSQL, Redis, idempotency keys, Prometheus, Grafana, Kafka, p99 latency."| Model | Clean, no prompt | Clean, prompt | Pink 5 dB, no prompt | Pink 5 dB, prompt |
|---|---|---|---|---|
| tiny.en | 3.66% | 3.96% | 5.64% | 5.49% |
| base.en | 2.74% | 1.22% | 5.34% | 4.42% |
| small.en | 1.37% | 0.76% | 2.59% | 1.98% |
| medium.en | 0.91% | 0.76% | 2.44% | 1.22% |
With the prompt, base.en got all three Redis, the idempotency, Grafana and
three of the four PostgreSQL on clean audio. It is now better than small.en
without a prompt, at half the runtime. small.en and medium.en with a prompt
tie at five errors. tiny.en got slightly worse, and spelled PostgreSQL,
Redis, idempotency and Grafana wrong with or without the hint.
The prompt helped base.en less in noise. At 5 dB it still wrote Redis wrong
all three times. small.en got every term right at 5 dB once prompted.
We cannot separate two causes here. Kokoro may pronounce "Redis" unusually. But the same audio scored 3 of 3 once the word was in the prompt, so the sound was close enough. What was missing was the prior.
It cost nothing measurable: base.en took 6.16 s with the prompt against a
7.09 s median without.
How fast is each model on an M3, and how much memory?
Real-time factor is wall time divided by audio length; below 1 is faster than
real time. These are from /usr/bin/time -l around whisper-cli, median of
three runs on the clean file, including model load:
| Model | Wall time (median) | Real-time factor | Speed | Peak RSS | Model load |
|---|---|---|---|---|---|
| tiny.en | 4.88 s | 0.019 | 52.5x | 338 MB | 76 ms |
| base.en | 7.09 s | 0.028 | 36.1x | 456 MB | 90 ms |
| small.en | 14.57 s | 0.057 | 17.6x | 923 MB | 283 ms |
| medium.en | 43.77 s | 0.171 | 5.9x | 2.22 GB | 785 ms |
Those runs were with the laptop's 1-minute load average between 2 and 9. An earlier pass of the same commands, while other jobs pushed the load average past 60, gave medians of 6.97, 13.81, 24.59 and 60.85 s. If your Mac is busy with a browser, a build and a video call, plan for the slower column.
With Metal switched off (-ng), still four threads, the same file took 21.22 s
on tiny.en, 39.75 s on base.en and 95.03 s on small.en, one run each.
small.en on the CPU is slower than medium.en on the GPU. Check that the
build you run reports using MTL0 backend in its log.
For live transcription, treat these as a lower bound. They are whole-file
numbers, where the model loads once and every 30-second window is full.
A live tool sends short chunks, and Whisper's encoder works on a 30-second
window however little audio is in it, so the cost per second of speech goes
up. We did not measure per-chunk latency. What the whole-file numbers do show
is headroom: small.en used 5.7% of real time on a quiet machine and 9.6% on
a busy one; medium.en used 17.1% and 23.8%.
Score your own recordings with wer.py
To find out whether this holds for your voice and your microphone, record a few minutes, type out exactly what was said, and run each model over it. The script below is the one that produced every WER in this article. It needs Python 3 and nothing else.
#!/usr/bin/env python3
"""wer.py REFERENCE.txt HYPOTHESIS.txt [--errors]"""
import re, sys
ONES = ("zero one two three four five six seven eight nine ten eleven twelve "
"thirteen fourteen fifteen sixteen seventeen eighteen nineteen").split()
TENS = "_ _ twenty thirty forty fifty sixty seventy eighty ninety".split()
def say(n):
if n < 20: return ONES[n]
if n < 100: return TENS[n // 10] + ("" if n % 10 == 0 else " " + ONES[n % 10])
for div, word in ((10**9, "billion"), (10**6, "million"), (1000, "thousand"), (100, "hundred")):
if n >= div:
q, r = divmod(n, div)
return say(q) + " " + word + ("" if r == 0 else " " + say(r))
def number(m):
whole, _, frac = m.group(0).replace(",", "").partition(".")
out = say(int(whole))
if frac: out += " point " + " ".join(ONES[int(d)] for d in frac)
return " " + out + " "
def normalise(text):
t = re.sub(r"^[QA]\|", "", text, flags=re.M) # our reference's speaker tags
t = t.lower().replace("%", " percent ")
t = re.sub(r"(?<=[a-z])(?=\d)|(?<=\d)(?=[a-z])", " ", t) # p99 -> p 99
t = re.sub(r"\d[\d,]*(\.\d+)?", number, t)
t = t.replace("'", "").replace("’", "")
return re.sub(r"[^a-z ]+", " ", t).split()
def align(ref, hyp):
d = [[i + j if i * j == 0 else 0 for j in range(len(hyp) + 1)] for i in range(len(ref) + 1)]
for i in range(1, len(ref) + 1):
for j in range(1, len(hyp) + 1):
d[i][j] = min(d[i-1][j] + 1, d[i][j-1] + 1, d[i-1][j-1] + (ref[i-1] != hyp[j-1]))
i, j, ops = len(ref), len(hyp), []
while i or j:
if i and j and d[i][j] == d[i-1][j-1] + (ref[i-1] != hyp[j-1]):
if ref[i-1] != hyp[j-1]: ops.append(("S", ref[i-1], hyp[j-1]))
i, j = i - 1, j - 1
elif i and d[i][j] == d[i-1][j] + 1:
ops.append(("D", ref[i-1], "")); i -= 1
else:
ops.append(("I", "", hyp[j-1])); j -= 1
return d[-1][-1], ops[::-1]
ref = normalise(open(sys.argv[1]).read())
hyp = normalise(open(sys.argv[2]).read())
errs, ops = align(ref, hyp)
print(f"WER {100 * errs / len(ref):.2f}% ({errs} errors / {len(ref)} words)")
if "--errors" in sys.argv:
for op in ops: print(" ", *op)Then loop over the models, keeping the transcript next to its timing:
ffmpeg -i interview.m4a -ar 16000 -ac 1 -c:a pcm_s16le interview.wav
for m in tiny.en base.en small.en medium.en; do
/usr/bin/time -l whisper-cli -m ggml-$m.bin -f interview.wav -l en \
-otxt -of out-$m 2> time-$m.log
echo "$m $(python3 wer.py reference.txt out-$m.txt) $(grep ' real ' time-$m.log)"
doneRun it once with --errors and read the list. The rate tells you how many
words broke. The list tells you whether a prompt would fix them, which is
usually cheaper than a bigger model.
The normaliser is opinionated in the same way as the one in normalising text for search: decide what counts as "the same" before you compare anything. If you transcribe more than one language, route by language first; a short-text detector is in how do you tell what language a message is in. And if the transcript goes on to a hosted model, its length is what you pay for, as measured in what a token costs.
What we could not measure
One synthetic recording, two synthetic American voices, read cleanly. No
accents, no non-native speakers, no laptop microphone, no room echo, no
cross-talk with a real second person. No languages other than English, and so
no multilingual models. No large-v3 or large-v3-turbo. No Core ML encoder
build. One M3, sharing the machine with other work.
So treat the absolute percentages as a floor. The pattern we would bet on
still holding with real voices: small.en with a vocabulary prompt is the
sensible default, base.en with a prompt is fine for clean audio,
and medium.en earns its 2.2 GB when other people are talking in the room.


