---
title: "Parakeet vs Whisper on Mac: a Meeting-Audio Benchmark — Mumr"
description: "NVIDIA Parakeet and OpenAI Whisper, run on-device on a Mac over an hour of real meetings, recorded two ways. Accuracy, speed, memory, and what each writes during silence. Raw CSVs included."
canonical: "https://mumr.app/research/local-transcription-benchmark-mac"
language: "en-GB"
last-updated: "2026-09-30T14:00:18+01:00"
source: "https://mumr.app — Markdown rendition of the HTML page at `canonical`"
---

Research · Benchmark

# NVIDIA Parakeet vs OpenAI Whisper on a Mac: a meeting-audio benchmark

The best-known speech-to-text scores come from clean audiobook recordings. In a meeting, people talk over each other and sit far from the microphone. We took an hour of real recorded meetings, heard through headsets and through one microphone in the room, and fed it to NVIDIA Parakeet and two sizes of OpenAI Whisper on a Mac. Each engine got the audio in 30-second pieces, the way a meeting recorder hands it over. We measured accuracy, speed, memory, and what each engine writes when nobody is talking.

Last updated: 30 September 2026 · By Mumr

The short version

-   On meeting audio, **Parakeet made half the errors of Whisper large-v3-turbo**: 20.3% word error rate against 40.4% on headset mics, 41.0% against 79.8% on a single room mic.
-   On a clean audiobook the order flips. **Whisper large-v3-turbo scored 3.0%**, and this Parakeet build dropped whole stretches of read speech (a known bug, explained below).
-   **Parakeet is about ten times faster**: an hour of audio in roughly 16 seconds against three to five minutes for Whisper, on an M5 Max.
-   In silence, **Whisper large-v3-turbo wrote 2 words on one run and 123 on the next**, over identical audio. Parakeet wrote four or five stray words.
-   One machine, one run per engine (two for the silence test). The per-file data is below as CSV.

## The results

Word error rate (WER) is the share of words the engine got wrong, left out, or made up, against a human reference transcript. Lower is better. "× real time" is how many seconds of audio the engine gets through per second of work. Apple M5 Max, 64 GB, macOS 26.6.2, every engine fed the same 30-second pieces.

| Engine | Meetings, headset mics | Meetings, one room mic | Audiobook, clean | × real time | Peak memory |
| --- | --- | --- | --- | --- | --- |
| NVIDIA Parakeet TDT 0.6B v3 | **20.3%** | **41.0%** | 25.2% ¹ | **212–227×** | 112–130 MB |
| OpenAI Whisper large-v3-turbo | 40.4% | 79.8% | **3.0%** | 13–20× | 175–177 MB |
| OpenAI Whisper small | 54.9% | — ² | 6.8% | 20–24× | 157–159 MB |

¹ This Parakeet build goes silent for 11 to 13 seconds at a time on read speech; see [the Parakeet gap](#parakeet-gap). ² Whisper small on the room mic ran only in an earlier pass whose raw data we did not keep, so we left it out.

Peak memory is the process footprint after the model is loaded. Whisper large-v3-turbo took 155 seconds and 643 MB to load the first time on a cold cache, and 4 to 5 seconds after that. Parakeet loaded in a tenth of a second.

## The meeting numbers run high

Next to published leaderboards, every meeting figure above looks bad. Two things push them up, and both hit all three engines the same way.

First, the [AMI Meeting Corpus](https://groups.inf.ed.ac.uk/ami/corpus/) is hard on purpose: four people in a room, overlapping speech, false starts, and a reference transcript that counts every "um" the engine left out. Second, we cut the audio every 30 seconds, whether or not someone is mid-word, because that is how a meeting recorder feeds its engine. Leaderboards cut on sentence boundaries. A word sliced in half at a boundary is an error for everyone.

Compare the rows with each other. The absolute numbers tell you how hard the audio is, and the gaps between rows tell you which engine copes with it.

## The room microphone

A meeting recorder on a laptop hears the far end of a call through the Mac's own audio, and whoever is in the room through one built-in microphone. The single-room-mic condition is closest to that, and there the engines differ most: 41.0% errors for Parakeet against 79.8% for Whisper large-v3-turbo.

Most of Whisper's errors there are missing words. On one five-minute piece Whisper large-v3-turbo wrote **nothing at all from 1:35 to 3:30** while Parakeet transcribed the whole stretch. The room audio was only 2 to 5 dB quieter than the headset recording of the same meeting, so volume doesn't explain the gap. Whisper uses its own confidence to decide whether a stretch is speech, and noisy distant speech often fails that test and gets dropped as silence.

## Words written into silence

Meetings have dead air: someone shares a screen, reads a document, waits for a colleague. An engine that invents text during silence puts sentences in your notes that nobody said. Whisper is known for this, and few people publish a number for it, so we built a test: real speech with six silences of 20 to 90 seconds inserted at known points, once as digital silence and once as quiet room tone at −60 dBFS.

| Engine | Digital silence | Room tone, −60 dBFS |
| --- | --- | --- |
| NVIDIA Parakeet TDT 0.6B v3 | 5 words, in 4 of 6 silences | 4 words, in 3 of 6 |
| Whisper large-v3-turbo, run 1 | 2 words, in 1 of 6 | 4 words, in 1 of 6 |
| Whisper large-v3-turbo, run 2 | **123 words, in 2 of 6** | 4 words, in 1 of 6 |

Same audio, same machine, same model. On the second run, one 75-second silence came back as _"ur ur ur…"_, 121 times. The other two words were _"Thank you."_, Whisper's best-known silence phrase, which it also wrote into room tone. Whisper stays quiet in most silences and fills some of them, and you can't tell in advance which run you'll get. An average of the two runs would hide that, so the table shows both.

Parakeet's were single words. Most were the tail of the sentence before the silence, stamped a tenth of a second into it ("the", "no.", "screen."). One, "Real.", appeared half a minute into room tone with nothing around it.

We didn't keep run 1's raw data. The downloadable `hallucination.csv` is run 2.

## Speed and memory

Parakeet ran at 212 to 227 times real time. Whisper large-v3-turbo ran at 13 to 20, and Whisper small at 20 to 24. For an hour-long meeting that is about 16 seconds against three to five minutes. On this machine both keep up with a live recording. You feel the difference when you import a backlog of recordings, and on older Macs, which this run didn't cover.

Memory was modest for all three: 112 to 177 MB once loaded. The number to plan for is Whisper large-v3-turbo's first load, which peaked at 643 MB and took over two minutes on a cold cache.

## The Parakeet gap on read speech

On every audiobook file, Parakeet wrote nothing from 0:30 to 0:43, 1:00 to 1:13 and 1:30 to 1:43, then carried on. That is what its 25.2% measures: on one file it wrote 212 words where Whisper wrote 323, and the missing ones sit in those windows. Resetting the model between pieces changed nothing. The pattern matches [FluidAudio issue #909](https://github.com/FluidInference/FluidAudio/issues/909), a property of the v3 Core ML model on certain inputs.

The meeting audio didn't trigger it: Parakeet's word counts there track Whisper's five seconds at a time. The version measured is FluidAudio 0.14.7; 0.15.7 carries long-form fixes and has not been measured yet. When it has, this page gets a second Parakeet row.

## Languages

Everything above is English, because that is what the corpus is. Coverage differs a lot: Parakeet TDT 0.6B v3 handles 25 European languages; Whisper handles 99. If you hold meetings in Japanese, Hindi or Arabic, Whisper is the only one of the two that will try, and nothing on this page tells you how well it does.

## Which one to use

-   For calls and meetings in English or a major European language, use Parakeet. It made fewer errors on real meeting audio, held on to far-away voices, and ran ten times faster.
-   For clean single-speaker audio, such as a lecture on a good mic or a podcast, use Whisper large-v3-turbo. On clean speech it beat everything else here by a wide margin.
-   For a language Parakeet doesn't cover, use Whisper, at the largest size your Mac handles without slowing down.
-   If your recordings have long silences and you use Whisper, check the transcript around the quiet parts.

## Where Mumr fits

We make Mumr, a free meeting recorder for Mac, and we ran this to decide what it should use. Mumr ships Parakeet as the default engine, with Whisper small and Apple's built-in Speech engine as options, all on-device.

Apple Speech is missing from this page. Its first run here found a bug in Mumr's own code: Mumr accepted the first finished sentence of each 30-second piece and threw away the rest: 27 words where 80 were spoken. The fix is written and ships in the next release; the Apple rows will be measured on that build and added here.

## Method

-   **Engines, as a Mac app ships them.** NVIDIA Parakeet TDT 0.6B v3 through [FluidAudio](https://github.com/FluidInference/FluidAudio) 0.14.7. OpenAI Whisper large-v3-turbo and small through [WhisperKit](https://github.com/argmaxinc/WhisperKit) (argmax-oss-swift 1.0.0). Default settings, English.
-   **Audio.** AMI Meeting Corpus, test split, CC BY 4.0: one hour through individual headset mics, and the same meetings through a single distant mic, as 5-minute files. LibriSpeech test-clean, 10 minutes, as a calibration set with well-known published scores. A 14-minute silence set built from real speech with six 20 to 90 second silences at known offsets.
-   **Input shape.** We cut every file into 30-second pieces before it reaches the engine, with no overlap and no voice-activity detection, because Mumr records that way. We didn't measure whole files.
-   **Scoring.** WER with Whisper's English text normaliser applied to both the engine output and the reference. Speed timed on the engine call alone, model already loaded. Silence scored as words whose timestamps fall inside a known-silent span.
-   **Calibration.** Whisper large-v3-turbo scored 3.0% on LibriSpeech test-clean, in line with its published figure of about 2.5 to 3%. That match is how we know the test harness measures what it claims to.
-   **Machine.** Apple M5 Max, 64 GB, macOS 26.6.2, thermal state nominal throughout. One run per engine, except the silence test, which Whisper large-v3-turbo ran twice.
-   **Not covered yet.** Older Macs (the M1 MacBook Air with 8 GB is the one most readers will want), Apple Speech, and FluidAudio 0.15.7. Each gets added as a dated row.

**Download the data:** [summary.csv](https://mumr.app/research/local-transcription-benchmark-mac/summary.csv) (one row per engine and condition), [per-file.csv](https://mumr.app/research/local-transcription-benchmark-mac/per-file.csv) (every 5-minute file, with substitutions, deletions and insertions), [hallucination.csv](https://mumr.app/research/local-transcription-benchmark-mac/hallucination.csv) (the silence test). `summary.csv` also carries one Apple Speech row, a single two-minute audiobook file (8.1%), too small to put in the table. Results from another Mac are welcome at [hello@mumr.app](https://mumr.app/contact).

## Parakeet on your Mac, for your meetings

Mumr records the call's audio and your mic on your Mac, transcribes on-device with Parakeet by default, and writes the recap with the AI you already use. Free, no account, no bot in the call.

[Download for Mac](https://mumr.app/download)

Free on the Mac App Store · macOS 15 or later

## FAQ

### Is NVIDIA Parakeet more accurate than Whisper on a Mac?

On meeting audio, in this benchmark, yes: 20.3% against 40.4% for Whisper large-v3-turbo on headset mics, and 41.0% against 79.8% on a single room mic. On clean read speech the order flips, and Whisper large-v3-turbo scored 3.0%.

### How fast is Parakeet compared with Whisper on Apple Silicon?

About 210 to 230 times real time for Parakeet, so an hour in roughly 16 seconds. Whisper large-v3-turbo and small ran at 13 to 24 times real time, three to five minutes for the same hour. Apple M5 Max, engine call only.

### Does Whisper hallucinate during silence?

Sometimes, and not repeatably. Over the same six silences, Whisper large-v3-turbo wrote 2 words on one run and 123 on the next. Parakeet wrote four or five stray single words.

### Can I re-run this benchmark?

The per-file results are downloadable as CSV, the audio is the public AMI Meeting Corpus, and the method section lists versions, input shape and scoring. Numbers from other Macs are welcome.
