The problem with voice typing has never been that computers cannot hear you. They hear you fine. The problem is that they write down exactly what you said — and what people say, out loud, in real time, is not what writing looks like.
You start a sentence, decide halfway through that it was the wrong sentence, and start again. You say "um" while you think. You correct yourself. Literal transcription faithfully preserves all of it, and hands you a paragraph you now have to edit. For a lot of people that editing takes as long as typing would have.
This guide is about closing that gap: what causes it, what does not fix it, and what does.
Key Takeaways
- The gap is structural — spoken language has fillers and restarts; written language does not. Literal transcription cannot bridge it.
- Speaking more carefully does not work for long, and it costs you the speed advantage that made dictation attractive.
- A cleanup layer does work — the tool removes fillers and false starts before you ever see the text.
- Cleanup should be a choice, not a default — sometimes you want the literal words, so Raw, Tidy, and Formal are separate modes.
- Custom vocabulary matters more than raw accuracy for names, jargon, and product terms.
Why Literal Transcription Creates Work
Say this out loud at normal speed: "so I think we should — actually, can we push the roadmap review, um, to Thursday instead?"
A literal transcriber writes exactly that. What you wanted was: "Can we push the roadmap review to Thursday instead?"
Every dictated paragraph carries some version of this. The fix people reach for first is to speak more carefully — slower, more deliberately, no fillers. It works for about two minutes, it is exhausting, and it removes the reason you started dictating: that talking is faster than typing.
What Does Not Fix It
Better speech recognition. Accuracy improvements make the transcription more faithful to your speech, which makes the problem slightly worse, not better. A perfect transcriber preserves every "um" perfectly.
Punctuation commands. Saying "comma" and "new paragraph" out loud helps with structure and does nothing about fillers and false starts. It also makes speaking feel like programming.
Editing afterwards. This is the status quo, and it is what makes people quietly stop using dictation after a week.
What Does: A Cleanup Layer
The fix is to put a step between recognition and output. The tool hears what you said, works out what you meant, and gives you the written version.
MirrorCaption Typer does this with three polish modes you pick per moment, because the right amount of cleanup is not the same every time:
- Raw speech — exactly what you said, every filler and false start preserved. Useful when the literal words matter: quotes, notes on how someone phrased something, language practice.
- Tidy (the default) — removes fillers like "um" and "uh", drops slips of the tongue and false starts, adds punctuation, and keeps your meaning and tone. This is the mode that makes dictation feel like it should have felt all along.
- Formal — rewrites the passage into clean, professional wording, ready to drop into an email or a document. It does change your phrasing, deliberately, which is why it is opt-in.
Taking the example above: Raw gives you the sentence with the restart and the "um". Tidy gives you "Can we push the roadmap review to Thursday instead?" Formal gives you something closer to "Could we move the roadmap review to Thursday?"
The Other Half: Getting the Text Where You Need It
Clean text is only half the job. If it lands in a transcription app you still have to copy it into the window you actually wanted it in.
Typer runs in a small window that floats above everything else — your mail client, a doc, a chat box. When you stop speaking, the polished text is copied to your clipboard automatically. There is no button to hunt for and no app to switch to. You paste, and you are done.
Because it runs in the browser there is nothing to install, which matters on managed work laptops where installing software means filing a ticket.
Names, Jargon, and the Words That Are Actually Yours
The errors that survive a good cleanup layer are usually proper nouns: colleagues' names, product names, internal jargon, technical terms. No general model has seen them.
Two things help. First, a custom vocabulary — you add the terms you use, and they are recognised first. Second, learning from your edits: when you correct something, the tool remembers your preference rather than making you fix it again next week.
When You Mix Languages
Plenty of people do not speak one language at a time. A Chinese speaker drops in English product names. A German engineer uses English technical terms mid-sentence. Most dictation tools force you to pick one recognition language and switch manually, which breaks the flow badly enough that people give up.
Typer lets you set several recognition languages at once, so a mixed sentence comes out accurately without switching modes. If your problem is instead that other people speak languages you do not, that is transcription and translation rather than dictation — see dictation vs. transcription for which side you are on.
When Literal Is the Right Answer
Cleanup is not always what you want. Use Raw mode when you are quoting someone, when the exact phrasing is the point, when you are practising a language and want to see what you actually said, or when you are capturing a thought in the shape it arrived. The point is that it is a choice rather than the only option.
Frequently Asked Questions
How do I dictate text that does not need editing?
Use a dictation tool that cleans up as it types rather than transcribing literally. MirrorCaption Typer's Tidy mode removes fillers like "um" and "uh", drops false starts and slips of the tongue, and adds punctuation before the text reaches your clipboard, so pasting is the last step rather than the first of several.
Can voice typing remove um and uh automatically?
Built-in tools such as Apple Dictation, Windows Voice Typing, and Google Docs Voice Typing transcribe literally, so fillers stay in. Tools with a cleanup layer remove them automatically. In MirrorCaption Typer this is a mode you pick per moment: Raw keeps everything, Tidy removes fillers and adds punctuation, Formal rewrites the passage into professional wording.
Will cleaning up my speech change what I meant?
It should not. Tidy mode is designed to keep your meaning and tone and only remove the noise — fillers, restarts, and slips. Formal mode deliberately does rewrite the phrasing, which is why it is a separate mode you choose rather than the default.
Can I dictate in more than one language at a time?
With most tools, no — you pick a recognition language and switch when you change. MirrorCaption Typer lets you set several recognition languages at once, so mixing Chinese with English, or German with English, inside one sentence still comes out accurately.
How do I make voice typing recognise names and jargon?
Add them to a custom vocabulary. Typer lets you add the names, product terms, and jargon you use often so they are recognised first, and it learns from the corrections you make to its output, so the same mistake does not repeat.
Does voice typing work in apps other than the browser?
Typer runs in a small browser window that floats above your other apps and copies the finished text to your clipboard automatically, so you can paste into email, docs, chat, or any other application. There is nothing to install, which also means nothing for IT to approve.
The Bottom Line
Voice typing stopped being an accuracy problem a while ago. It is a translation problem — between how people talk and how text is supposed to read — and literal transcription leaves that translation to you.
Tools that do it for you turn dictation from a thing you tried once into a thing you use every day. Compare the options in our best dictation software guide.
Speak Messy. Paste Clean.
Typer removes the fillers and false starts before the text reaches your clipboard. 1 free hour, no credit card, nothing to install.
Try Typer Free