Changelog

Every release, on the record

Versions match the public release tags. Auto-updates ship to all tiers silently, so whatever you read here is already on your machine.

v1.19.3latest

Custom vocabulary actually works now, and a screen that explains the app

  • Custom vocabulary was silently doing nothing in 1.19.0 through 1.19.2. The app was fetching an older speech engine than the one it shipped with, and that older engine had never heard of custom vocabulary, so every term you taught it was quietly dropped. The correct engine is now fetched on first run after updating, which is what actually restores the feature. If you added vocabulary terms and wondered why nothing changed, this was why - they are still saved, and they work now.
  • New "How Voxmelt works" screen: a short loop showing your voice entering the machine, the CPU turning it into text, the GPU rewriting it, and the result landing in your document, with nothing leaving for the cloud. Open it any time from the eye button in the title bar; it also appears on the welcome step of first-run setup.
  • The speech model can no longer be unloaded while a recording is running. Doing so could take the model out from under an in-progress transcription; both the Models page and the engine itself now refuse it and say why.
  • Switching engine, device or precision now tells you the change applies to your next recording. Previously that notice only appeared when changing Whisper model size, so moving GPU to CPU, or Whisper to Parakeet, looked like it had taken effect immediately while the old model was still transcribing.
  • The About screen now lists the engines actually doing the work, including the attribution Parakeet requires, instead of a stale stack list that still claimed faster-whisper on CUDA.
  • Release builds now verify their own version before shipping, so the mismatch that caused the vocabulary bug cannot happen again unnoticed.
v1.19.2

Downloads that finish, payments that hold

  • Model downloads no longer stall forever. Store users reported transfers frozen at 84 percent with the speed at zero: the connection had gone quiet and nothing woke it up. Every chunk is now awaited under a timeout, a dead link reconnects and resumes from exactly what is already on disk, and the app tells you it is reconnecting instead of pretending to still be downloading.
  • Pause during a download now responds immediately rather than waiting on a connection that had already stopped talking.
  • Three further model-download and model-selection bugs found in the Store build are fixed.
  • Payment handling hardened: refunds and chargebacks are processed properly, a paid tier can never be downgraded by a stale or out-of-order event, failed entitlement grants are detected instead of silently passing, and trial registration no longer drifts out of sync with the real purchase.
  • All 14 outstanding dependency security advisories cleared.
v1.19.0

Per-app tones, custom vocabulary, one-time pricing

  • Voxmelt is now a one-time purchase. No subscription, no renewal, no monthly meter: buy the tier once and it stays yours. App copy, terms, and the refund policy were all rewritten to match.
  • Studio: per-app tones in Mini Mode. Bind an application to an AI template and anything you dictate into that app is styled automatically - commit messages in your editor, plain prose in your browser. The app you are actually typing into is resolved by process, and unmapped apps fall back to your own default template instead of inheriting the last app's tone.
  • Pro and Studio: custom vocabulary. Teach the transcriber the names it keeps mangling, on either engine. It rejoins names that get split apart ("Box melt" becomes "Voxmelt") and lets you map an exact mis-hearing when a name is too far off for safe matching. Deliberately conservative so it never rewrites a word you actually said, and the transcript header shows what it corrected with one click to see the original.
  • Idle auto-stop: a recording you forget about ends itself after 60 seconds with no human voice, using voice detection rather than raw loudness, so a noisy fan cannot hold the session open. Configurable, and you can switch it off.
  • Fixed: AI post-processing never ran automatically if the AI model was still cold, and that session stayed marked as processed forever.
  • Fixed: a fresh install's first AI run translated the dictation instead of cleaning it up, because the default template was Translate. New installs now start on Vibe Dictation.
  • Fixed: the trial-ended message could appear immediately after the welcome tour, and quoted a 30 minute daily free cap that has been 60 minutes for some time.
  • Fixed: deleting many history sessions at once rewrote the whole index for every single one. It is a single operation now.
  • Timestamps are off by default for new installs. Existing installs keep whatever they had.
v1.18.8

The CPU-first Store release

  • This build is live on the Microsoft Store: dictation runs on your CPU via NVIDIA Parakeet TDT v3 by default - no graphics card required. Your NVIDIA GPU, if you have one, is reserved for the AI text studio.
  • Humming, coughs, and background noise no longer turn into stray filler words: every Parakeet chunk now passes a speech gate before decoding, and filler-only output is dropped at the source.
  • New first-run journey: you land on the Models hub, one guided Parakeet download, then straight to Record. If an NVIDIA GPU is detected, AI polish offers to set itself up too; if not, setup stays minimal on purpose.
  • Mini Mode's live-dictation glow now switches off the moment focus leaves a real text field - clicking the desktop, the taskbar, or another app clears it correctly.
  • The engine chip and the speed stat now name the engine and device actually running - no more "GPU" label on a CPU transcription.
  • Legal docs updated to v2.1 for the CPU-first architecture.
v1.18.0

The hybrid pivot

  • New default dictation engine: NVIDIA Parakeet TDT 0.6B v3 (int8) running entirely on your CPU cores, so dictation no longer needs a graphics card at all. On our published real-clip benchmark it wins outright on fast natural speech (2.0 vs 3.1 percent word error rate) and is the strongest no-GPU option we measured, while Whisper large-v3 on a GPU still leads overall (7.8 vs 11.9 percent) and ships for anyone who has the hardware - see voxmelt.com/benchmark.
  • The GPU's new job is AI: with dictation off the card, your Ollama model stays warm in VRAM for instant rewrites, summaries, and translations. No more model juggling on 12 GB cards.
  • Parakeet tuning for real dictation: boundary silence pads stop edge-word clipping, a CPU thread cap keeps the system smooth while you speak, and the refined pass budgets itself for TDT's attention cost.
  • Whisper is still in the app for what Parakeet cannot do: ~100 languages and the GPU accuracy king, one toggle away.
  • 1.18.1-1.18.7 follow-ups: the Models page became the single model-config hub with a CPU | GPU split, a live microphone test card, a slimmer first-run wizard, and Ollama detection + Mini Mode fixes.
v1.16 series

The Store certification grind

  • Voxmelt entered the Microsoft Store: a full MSIX channel with Store-managed updates, built from the same codebase as the auto-updating EXE.
  • First crack in the GPU wall: with no NVIDIA GPU present, transcription now falls back to a bundled CPU model instead of refusing to run.
  • A long tail of container hardening: model downloads inside the MSIX sandbox, long-path handling, engine load watchdogs, and crisp unplated taskbar icons.
v1.15 series

The 5 MB installer

  • The heavy CUDA engines moved out of the installer and now download on first run, resumable - the installer itself dropped to a few megabytes.
  • Free tier doubled: 30 to 60 minutes of dictation per day.
  • Read Aloud stability fixes, per-module onboarding tour, chunk-processing indicator, and friendlier limit messages.
v1.14.0

One codebase, two channels

  • Dual-channel builds: the auto-updating EXE and the Microsoft Store MSIX now come from the same source, with the app detecting its own channel at runtime.
  • Self-contained Python sidecars - no system Python, no PATH surprises.
v1.12.0

Brain Capsule

  • Every built-in template is now compiled on-device by a three-layer prompt engine (action / format / style) living in the compiled Rust backend. Small local models stopped collapsing instructions, and the recipes never leave the app.
v1.11.0
  • Voiceover grew a persona x tone voice matrix with curated delivery tones.
  • Compose and Record now pick Template and Tone independently, matching how you actually think about output.
v1.10.0

Compose / Voiceover split

  • Compose and Voiceover became their own full surfaces, with the template ladder v2 and a tone-first picker.
  • 1.10.1: an on-canvas "Get started" guide walks new users to their first template, model, and recording after setup.
v1.9.0

The IA overhaul

  • Navigation reorganized around intent: Record is only for recording, configuration lives in Settings, and the AI Studio composer got its own home.
  • Tone-shift chips and text-encoding fixes across the AI pipeline.
v1.8.0

The voice suite

  • TTS Read-Back: a local text-to-speech engine reads your polished text back - 32 voice combinations, WAV export, all offline.
  • Voice Director: hands-free voice commands steer the pipeline while you dictate.
v1.7.0

The stabilization release

  • Language setting now fully honored end to end: Auto-detect runs Whisper language detection, and all 10 language choices apply on the next recording without an engine restart.
  • Long dictations no longer risk a false "engine stuck" error: the final full-session accuracy pass now budgets itself against your GPU's measured speed and steps aside when the streaming transcript is already the right answer.
  • Short utterances like a lone "Yes." are no longer swallowed by the hallucination filter - real speech energy now overrides the phantom-phrase list.
  • Voice commands got a precision overhaul: saying "go" or "copy" alone no longer triggers "go ahead" or "copy that". Match threshold raised, partial matches now require most of the phrase.
  • GPU safety on both sides: loading Whisper now checks free VRAM and evicts the AI model when the card is full, on cold start and resume alike. No more cryptic CUDA errors when a big model is pinned.
  • Plan rebalance: the free tier now includes the flagship Casual Clean cleanup tone, Pro covers every everyday template with models up to 14B, and Studio owns the AI Prompt + Code packs plus the full 24B to 70B model garage.
  • Switching AI models or toggling AI off mid-generation now cancels the run cleanly instead of leaving it wedged.
v1.6.8
  • Silent updates: per-user installs mean no UAC prompt on first install or any auto-update from here on.
  • Keep-both is now the default GPU mode - Whisper and your AI model stay warm together, and the models panel mirrors what is actually resident in VRAM.
  • Refreshed installer wizard art and branding.
v1.6.7
  • Theme parity with voxmelt.com: five dark themes including Candy Red and Mono, each with a matching titlebar logo.
  • Re-rendered the full icon set natively at every size - crisp from the taskbar to the Store tile.
v1.6.6
  • Rebindable global hotkeys: trigger recording from any app, with defaults chosen to never steal system-wide combos.
v1.6.5
  • Mini Mode now mirrors the main window's recording state perfectly across monitors.
  • Closing to tray keeps voice commands listening when Mini Mode is active - "go ahead" works with the main window hidden.
v1.6.4
  • The Mini Mode pill now runs its own audio analyser, so its mic animation matches the main record button frame for frame.
v1.6.3
  • App-wide zoom: Ctrl + mouse wheel or Ctrl +/- /0, from 50% to 200%, with a momentary zoom HUD.
  • Visual polish pass: refreshed call-to-action buttons and consistent plan badge colors everywhere.
v1.6.2
  • AI stall watchdog: a generation that stops producing tokens is auto-recovered with a retryable error instead of hanging on "Processing".
  • Accuracy stack tuning: chunk-overlap dedup, repetition collapse, and the hallucination filter now work as one contract.
v1.6.1
  • Honest tokens-per-second: AI speed now reports pure generation rate, with model load time shown separately.
  • Eliminated up to two hidden model reloads before the first token in the auto-process path.
v1.6.0

GPU-only, guaranteed

  • New invariant: Voxmelt never runs a model partially on the CPU. Every load path now checks the fit before a single byte hits VRAM, and over-budget models get a clear popup instead of a 2 tok/s crawl.
  • Smart swap: when Whisper must yield to a big AI model, it suspends in place and resumes in 2 to 3 seconds instead of paying a full cold start.
  • Co-residency warmup: the AI model starts loading during the final transcription pass, so cleanup begins the instant your transcript is ready.
v1.5.1
  • Fixed: switching the Whisper model size in settings now actually loads the new model on your next recording.
v1.5.0
  • VRAM numbers you can trust: one source of truth for every model footprint, corrected by real measurements from your own card after first load.
  • New int8_float16 precision option: half the Whisper VRAM with near-identical accuracy, for running both engines on tighter cards.
v1.4.0

The model garage

  • Curated AI model catalog from 3B to 70B: Gemma, Qwen, Llama, DeepSeek R1, Mistral, and MoE flagships, each with measured VRAM and strengths.
  • The picker filters to what your card can actually run, and downloads stream in-app with live progress.
v1.3.0

No-Hands Mode

  • Voice commands: start, stop, copy, clear, and load or unload models entirely by voice, powered by an always-on tiny model that coexists with the main engine.
  • Mini Mode: an always-on-top pill with live waveform, perfect for dictating into another app on a second monitor.
v1.2.0

The template library

  • AI templates organized by persona: Voice Polisher, Ghostwriter, Data Architect, Polyglot, and more, with multiple tones per template.
  • Custom templates: write your own system prompt, pick a persona, and it runs 100% on your GPU like every built-in.
  • Template discovery page with search and per-tone live samples.
v1.1.0

Free for everyone, fairly

  • Freemium launch: 15-day full trial of everything, then a free tier that stays genuinely useful - no card required.
  • Pro and Studio subscriptions via Razorpay, with cryptographically signed entitlements verified in compiled code.
  • Trial is anchored to your device, so a reinstall picks up where it left off.
v0.3.0
  • Accounts: email + password and one-click Google sign-in, used only to manage a paid license. Recording never requires an account.
  • Session history: every dictation saved locally, searchable, with optional AI output alongside.
v0.2.0
  • The second brain: local AI post-processing through Ollama - cleanup, summarize, reformat - streamed token by token, on the same GPU.
  • Dark theme system with live accent switching.
v0.1.2
  • System tray with close-to-tray: the X button hides the window and keeps your hotkey armed and models warm.
  • Branded installer, mic permission handled silently.
v0.1.1
  • Zero-terminal first-run wizard: Ollama, Whisper, and your first AI model installed with clicks, not commands.
  • Clear, actionable errors for driver and CUDA version issues.
v0.1.0

First public build

  • Whisper large-v3 dictation on your own NVIDIA GPU: press record, speak, get text in near real time.
  • Continuous dictation with overlapping chunks and boundary dedup for unlimited session length.
  • Global hotkey with auto-copy and auto-paste into whatever app has focus.
  • Zero cloud calls, zero telemetry, works with the network cable unplugged. The whole point, from day one.