Private beta

Tract

Dictation transcribed on your own machine.

The problem

What Tract exists to change.

Most dictation tools upload audio to a service to transcribe it. If the material is under NDA, covered by a contract, or a clinical note, that upload is the problem, and a privacy policy does not undo it.

Running speech models locally on Apple Silicon is possible. The work is in the plumbing: fetching and storing model weights, wiring Metal and CoreML, resampling audio to what the model expects, and handling the machine going to sleep mid-recording.

Raw speech-to-text is not usable text. It arrives without punctuation or casing, and sending it to a cloud model to fix that hands back exactly the privacy the local transcription just bought.

Approach

How Tract approaches it.

Tract is a Tauri 2 application: a Rust core, a React interface, and a Swift bridge for the macOS APIs with no Rust binding — App Intents, ScreenCaptureKit, the Speech framework.

Transcription runs through whisper.cpp with Metal and CoreML switched on. The default model is Whisper large-v3-turbo at roughly 1.55 GB on disk. Smaller weights are selectable down to 80 MB, and a quantised large-v3-turbo sits in between at 600 MB.

Parakeet TDT 0.6B, about 460 MB, is the second engine. It runs int8 ONNX on the CPU through sherpa-onnx, which matters on an Intel Mac or when the GPU is busy.

The cleanup pass — punctuation, casing, the obvious misheard word — uses Llama 3.2 3B Instruct at 4-bit through Apple's MLX. That stays on the machine too, because moving it off would defeat the point.

How it works

What actually happens when you run it.

  1. 01

    Capture

    Audio comes off the selected input and is resampled to 16 kHz mono, which is what both engines expect.

  2. 02

    Transcribe locally

    The selected model runs on the machine. whisper.cpp uses Metal with flash attention; Parakeet runs on the CPU. Weights are already on disk, so no request goes out while you are talking.

  3. 03

    Clean up

    A local Llama 3.2 3B pass adds punctuation and casing and fixes obvious errors, turning a transcript into something you can paste.

  4. 04

    Hand off

    The text goes where you are working. An optional encrypted sync keeps a history across your own machines.

Features

What’s inside.

Transcription

  • Whisper large-v3-turbo by default, on Metal with CoreML
  • Model sizes from 80 MB to 1.55 GB, including a 600 MB quantised build
  • Parakeet TDT 0.6B on CPU for machines without a usable GPU path
  • 16 kHz mono resampling handled before the model sees anything

What leaves the machine

  • Nothing during transcription or cleanup
  • Model weights, once, from Hugging Face and GitHub releases
  • Optional sync, encrypted on the client with XChaCha20-Poly1305 before upload
  • An optional cloud rewrite with your own API key, off unless you turn it on

Fit

Where it belongs.

Who it’s for

  • People who dictate material they are not allowed to upload — legal drafting, clinical notes, anything under NDA.
  • Writers and developers on Apple Silicon who would sooner spend 1.5 GB of disk than a per-minute transcription bill.
  • Early users who can work with a 0.1.0 build that changes between releases.

What changes

  • Transcription that completes with the network off.
  • A transcript that already has punctuation and casing, without a second tool or a second upload.
  • One install that covers both Apple Silicon and Intel Macs on macOS 14 and later.

Built on

  • Tauri 2.11 with a Rust 1.77 core, React 19, and Vite
  • whisper.cpp through whisper-rs 0.15, built with Metal and CoreML
  • sherpa-onnx 1.13.3 running Parakeet TDT as int8 ONNX
  • Apple MLX with Llama 3.2 3B Instruct at 4-bit for the cleanup pass
  • A Swift bridge for App Intents, ScreenCaptureKit, and the Speech framework
  • Argon2id, HKDF, and XChaCha20-Poly1305 for the optional encrypted sync

What to know before you buy

  • macOS 14 or later, and macOS only. The Rust core links against Cocoa and a Swift bridge, so there is no Windows or Linux build and no near-term path to one.
  • "On-device" describes transcription and cleanup. Model weights are downloaded once from Hugging Face and GitHub. Two optional features use the network: encrypted sync to tract.nexg.dev, which encrypts on your machine before anything is uploaded, and a bring-your-own-key cloud rewrite that is off by default.
  • Version 0.1.0. Auto-update is not enabled, crash logs stay on the machine with nothing collecting them, and some application icons are still placeholders.
  • Several engines were built and then cut. A streaming Zipformer path measured around 45% word error rate and was removed. A Neural Engine path never worked well enough to keep. Kyutai STT and NVIDIA Canary were both evaluated and rejected because neither exports cleanly to the on-device Mac stack.
  • Private beta. Access is limited, behaviour changes between builds, and none of this is a commitment to a release date.

Questions

Before you ask for a quote.

The things people check before asking for a number.

No. Both engines run locally against weights already on disk, and the cleanup pass runs locally too. The network is used to download a model the first time you select it, and by two optional features you can leave switched off: encrypted sync, which encrypts on your machine before upload, and a cloud rewrite that needs your own API key.

Whisper large-v3-turbo is the default and the most accurate, at about 1.55 GB. The quantised build at 600 MB is the usual trade if disk or memory is tight. Parakeet TDT at roughly 460 MB runs on the CPU, which is the one to pick on an Intel Mac.

No, and there is no near-term plan for one. The application depends on macOS frameworks through a Swift bridge, so a port would mean replacing that layer rather than recompiling.

It is at 0.1.0 and in daily use by a small group. It builds signed universal installers, but auto-update is off, crash reporting is local only, and some icons are placeholders. Treat it as a working tool that is still moving.

Access is limited during the private beta. Check https://tract.nexg.dev for current status, or contact NexG.

Get a quotation

Tell us what you need.

Share the workflow, users, and constraints you're working with. We'll follow up with a scoped quotation.

Required

Email us instead
Back to top