6 min read
Transcribe Audio and Video Files Offline on Windows
Turn a recording you already have into text, speaker-labelled notes, or an SRT subtitle file, with Whisper running on your own Windows PC instead of an upload.
For people who need a transcript or subtitles from a file and do not want to upload the recording to get them.
5,000 words per day. No account required. Local by default, cloud optional.
The job this actually does
Most transcription tools want you to upload the file. That is fine for a conference talk and a problem for an interview, a client call, a therapy session, a legal recording, or an unreleased video. Once the file is on someone else’s server you are trusting a policy instead of checking a fact.
PrivateTranscribe has a Transcribe File page that runs the same local Whisper engine used for dictation, pointed at a file on your disk. You get plain text, timestamped text, speaker labels, or an SRT subtitle file, and the recording does not need to leave the machine to produce any of them.
Run it once on a file you already have
Open the Transcribe File page in the sidebar and pick a recording. Audio works as wav, mp3, m4a, ogg, flac, or webm. Video works as mp4, m4v, mov, mkv, avi, or webm, and the audio is pulled out for you, so a screen recording or a camera file goes in directly without a conversion step first.
Local transcription takes files up to 500 MB. If you switch to a cloud provider instead, the limit drops to 25 MB, because that is the provider’s limit rather than ours. Pick a file you know the content of for the first run so you can judge the accuracy rather than guess at it.
Turn on speaker labels
Speaker labels are a toggle on the same page. Leave the speaker count on auto-detect, or set it yourself anywhere from 2 to 6 when you already know how many people are in the room. The output then reads as Speaker 1 and Speaker 2 rather than one undivided wall of text, which is the difference between a transcript you can use and one you have to re-listen to.
The first time you switch it on there is a one-time model download, because speaker separation is a separate model from Whisper itself. There are two engines. The multilingual one is preferred and gets used when it is available. The fallback is English only, so for a non-English recording let the multilingual model finish downloading before you run the file.
Export it as subtitles or notes
The same transcript comes out in four shapes. Plain text for pasting into a document. Timestamped text when you need to find the moment something was said. Speaker-labelled text for interviews and calls. SRT when you want subtitles a video editor can read.
SRT is the one people underestimate. If you are captioning your own footage, the normal route is uploading the video to a captioning service and waiting. Here the file never moves, and the subtitle file lands next to it.
What stays local and what does not
With a local Whisper model selected, the file is processed on your Windows PC and the transcript is written on your Windows PC. Nothing about the recording needs to reach a transcription provider.
Cloud transcription is still available if you deliberately switch to it, and then the audio does go to that provider under their policy, not ours. We would rather say that plainly than claim everything is always offline. If you want the local claim checked rather than believed, we measured it and published the method, the connections we did find, and where the test is weak.
Limits worth knowing before you rely on it
Whisper is good, not perfect. Names, jargon, overlapping speech, and heavy background noise are where it slips, and a larger model helps more than any setting does. Speaker separation is a best guess rather than a fact, so check the labels on a recording where being wrong would matter.
This runs on Windows only for now. Speed depends on your machine, since an NVIDIA card with CUDA is much faster than CPU, and a long file on CPU takes a while. It is a one-time wait on your own hardware rather than a queue you are paying to sit in.
Test it in Starter first
File transcription, speaker labels, timestamps, and SRT export are all in Starter. There is no account, no card, and no separate purchase for them.
Starter includes 5,000 transcribed words per day, every supported local Whisper model, and CUDA acceleration on compatible NVIDIA GPUs. Run one real file through it before deciding anything. Pro is for people who keep hitting the daily limit, and its first concrete upgrade is removing it.
Try local dictation privately
PrivateTranscribe Starter is free for Windows. It includes local Whisper transcription, global hotkeys, automatic paste, and 5,000 transcribed words per day without requiring an account.
Download Starter for Windows