Two engines

Two AI text-to-speech engines, because voice is a trade-off.

Kokoro is the "done in half a second" TTS engine — instant speech synthesis, no GPU needed. Chatterbox is the "wait a bit, get something you'd actually play for a client" engine — human-sounding narration with voice cloning. Pick per line, per project, whenever.

Fast mode

Kokoro TTS

Kokoro 82M · ONNX runtime · CPU

Near-instant generation. Sounds clean and consistent — perfect for narration, audiobooks, and anything where you need speed and clarity over emotional range.

  • Latency~0.4s per 30s of audio
  • Voices13 built-in (US + UK, M/F)
  • Model size82 MB (bundled)
  • RequiresCPU only · 4 GB RAM min
  • CostFree forever, unlimited
Natural mode

Chatterbox TTS

Chatterbox · PyTorch · CPU/GPU

Slower, but genuinely human-sounding. Clones any voice from a 6-second reference sample. The one you'd use for a real podcast, ad, or client deliverable.

  • Latency~4s per 10s of audio (CPU)
  • Voice cloningYes, 6s sample minimum
  • Model size~3 GB (first-use download)
  • Requires6 GB RAM min · GPU optional
  • Cost3 free, then Pro unlocks unlimited
What's inside

Open-source AI speech models. No black boxes.

VoxWild is a Python desktop app that bundles a small number of open-source AI text-to-speech models. Here's exactly what they are and where they come from, because you shouldn't install a speech synthesis tool that won't tell you.

PlatformWindows 10/11 x64
Installer size~377 MB
Disk after install~800 MB (Fast mode only) · ~5 GB with Natural mode
Minimum RAM4 GB (Fast mode) · 6 GB (Natural mode)
GPUOptional · speeds up Enhancement if CUDA available
NetworkOnly for license activation and update checks
TelemetryNone

Hear both engines →  ·  What it costs →

Both engines. One download.

Download for Windows
Fast mode free forever · Natural mode: 3 free tries, then Pro