dictation without a gpu

You do not need a graphics card to dictate.

Local dictation used to mean "bring your own GPU". That stopped being true. On our published benchmark, a 0.6B speech model running on a plain CPU transcribed five real recorded clips at 11.9 percent word error, using zero graphics memory, faster than 11 times realtime. Here is the data, and what a GPU is actually still worth.

Download for WindowsGet it from Microsoft Store

Runs on any modern Windows 10 or 11 PC · no VRAM required · free tier, no card

measured, not asserted

CPU beats the GPU model most people can actually run

Word error rate across five real recorded clips. Lower is better. The comparison that matters is not against the flagship, it is against the model your card could realistically load.

Parakeet v3 Plain CPU · 0 GB VRAM11.9% WER
Whisper large-v3 RTX 3080 Ti · ~4.2 GB VRAM7.8% WER
Whisper medium RTX 3080 Ti · ~2.3 GB VRAM22.8% WER

Whisper large-v3 on a flagship GPU is still the most accurate overall, and it ships inside Voxmelt for exactly that reason. The point of this chart is narrower and more useful: if you do not own a flagship card, the CPU option is not the compromise. It is the better result. Full method, per-clip numbers and hardware.

the clip that looks most like real dictation

On fast natural speech, the CPU wins outright

The test set includes a clip of fast, unedited speech with filler words, which is what real dictation actually sounds like. On that clip Parakeet on the CPU scored 2.0 percent word error against 3.1 percent for Whisper large-v3 on the GPU.

We are not claiming the CPU model is better overall, because it is not: across all five clips the flagship is ahead, 7.8 against 11.9. But the clip that most resembles how people really talk is the one where no graphics card was needed at all.

so what is the gpu for

We moved it to the job that actually needs it

Your CPU does the listening

Parakeet v3 in int8, 0 GB of VRAM, ready in about four seconds. English plus 24 European languages. This is the default, on every machine.

Your GPU does the thinking

If you have an NVIDIA card, it runs a local language model through Ollama that turns a rambling transcript into the email, summary or commit message you actually wanted.

Neither one phones home

Both halves run on your machine. No audio and no transcript leaves the device, with or without a graphics card, online or off.

That split is the whole design. Dictation is a solved problem on commodity hardware, so spending a graphics card on it is a waste of the one component that could be doing something harder.

every clip, no cherry-picking

Including the ones where the CPU loses

ClipParakeet, CPUlarge-v3, GPU
Clean dictation8.0%6.8%
Technical / code speech20.7%16.1%
Fast natural speech with filler2.0%3.1%
Medical jargon21.7%10.9%
Background noise8.3%3.1%

Technical speech and medical jargon are the CPU model’s weak spots, and we would rather you saw that here than discovered it later. That is what the engine switcher is for: if your work is full of jargon, run Whisper instead. Both ship in the same app.

questions

GPU-free dictation, answered

Do you actually need a GPU for voice-to-text?

No, and our own published benchmark is the evidence. NVIDIA's Parakeet TDT v3 running on a plain CPU scored 11.9 percent word error across five real recorded clips, using 0 GB of VRAM, at 11 to 13 times faster than realtime. A graphics card buys you a better ceiling, not a working product. Dictation is no longer the part that needs one.

Is CPU dictation slower than GPU dictation?

Not in any way you would feel while dictating. On the same clips, CPU Parakeet ran at 11 to 13x realtime, meaning a minute of speech is transcribed in roughly five seconds. Both engines finish far faster than you can talk, so the bottleneck is your speaking speed, not the hardware.

How does CPU accuracy compare to a mid-range GPU?

This is the part most people get backwards. On our test set, Parakeet on a CPU (11.9 percent word error) was substantially more accurate than Whisper medium running on an RTX 3080 Ti (22.8 percent). Whisper large-v3 on that GPU is still the most accurate overall at 7.8 percent, but it needs roughly 4 GB of VRAM. If you do not have a flagship card, a modern CPU model beats what a mid-tier GPU can actually run.

What hardware do I actually need?

Any Windows 10 or 11 PC with a reasonably modern processor. There is no VRAM requirement for dictation because no graphics memory is used at all. Integrated graphics, an ultrabook, a work laptop with the GPU locked down, or a mini PC are all fine.

So what is the GPU for, then?

The AI text studio. If you have an NVIDIA card, Voxmelt uses it to run a local language model that rewrites your raw dictation into an email, a summary, a commit message or a translation before you paste it. That is the step that genuinely benefits from a GPU, and it is optional. Dictation never needs it.

Does Windows built-in voice typing work without a GPU?

It does not need a GPU, but Microsoft documents that standard voice typing requires an internet connection, because the recognition happens on their servers rather than on your PC. So it solves the hardware question by moving your audio off the machine. Voxmelt solves it by running a model that fits on the CPU you already have.

Where can I see the full benchmark?

On the benchmark page. It lists every clip, the word error rate and character error rate per engine, processing time, realtime multiple, peak VRAM, and the exact hardware and library versions used. The numbers on this page are generated from that same dataset, so the two can never disagree.

Try it on the PC you already own

No graphics card, no VRAM, no cloud. Free tier with no card, a 15-day trial of everything, and a one-time purchase after that.