Updated September 5, 2026 · By Sumbat.T

Search this question and you will get the same number from almost every result: dictation is three times faster than typing. It appears in Google’s AI summary, on Logitech’s blog, on a dozen productivity sites, and in the marketing copy of every dictation app including, in various forms, ours.
It traces back to a single study. That study is real, it is careful, and it says something narrower than the sentence it has been compressed into. Almost nobody quoting it has read it, which is how a finding about reading text messages aloud on an iPhone in 2016 became a general claim about writing at a desk.
This piece goes to the four studies that have actually measured both methods against each other, including one that found typing faster, and works out what the answer is for the writing you personally do. The honest version is more useful than the slogan, and it still ends up recommending dictation for most of what most people type.
Raw input rate, in words per minute
These are entry rates before any correction, from four separate studies. The gap they show is real, and the rest of this article is about how much of it survives contact with actual work.
Conversation, silences removed
236 wpm
Yuan, Liberman & Cieri (2006), 2,438 conversations
Conversation, total elapsed time
196 wpm
Same corpus, range 111 to 291 wpm
Dictating short messages
153 wpm
Ruan et al. (2016), 32 participants
Typing, physical keyboard
51.6 wpm
Dhakal et al. (2018), 168,000 people
Typing, phone touchscreen
36.2 wpm
Palin et al. (2019), 37,370 people
The source is a 2016 paper by Sherry Ruan, Jacob Wobbrock, Kenny Liou, Andrew Ng and James Landay, from Stanford, the University of Washington and Baidu. Thirty-two people entered the same short messages twice, once with the built-in iOS keyboard and once by speaking to Baidu’s Deep Speech 2 recogniser, in English and in Mandarin.
The English result was 153 words per minute by voice against 52 by keyboard, a ratio of 2.93. Mandarin came out at 123 against 43, a ratio of 2.87. Those are good numbers from a well-run experiment, and the study deserves its reputation.
There is a detail in its publication history that tells you how the claim drifted. The version posted to arXiv in August 2016 was titled “Speech Is 3x Faster than Typing for English and Mandarin Text Entry on Mobile Devices”. The version that went through peer review was retitled “Comparing Speech and Keyboard Text Entry for Short Messages in Two Languages on Touchscreen Phones”. The headline claim was removed from the title, and the scope conditions were put in: short messages, two languages, touchscreen phones. The internet kept citing the preprint title.
“We found that with speech recognition, the English input rate was 2.93 times faster (153 vs. 52 WPM) [...] than the keyboard for short message transcription under laboratory conditions for both methods.”
Three qualifications are doing real work in that sentence. Read them carefully, because each one changes how far the result travels.
Transcription
Participants were given the message and said it aloud. They never had to decide what to write, which is the slowest part of most real writing.
Touchscreen
The keyboard it beat was an iPhone 6 Plus. Average typing on a physical keyboard is 51.6 wpm against about 36.2 wpm on a phone, so a laptop keyboard is a stiffer opponent.
Short messages
Text-message length, in a lab. Nothing about the study speaks to a 2,000-word document, a spreadsheet, or a codebase.
What the study is still good evidence for
In 1999, Clare-Marie Karat, Christine Halverson, Daniel Horn and John Karat at IBM’s T.J. Watson Research Center ran 24 users through both transcription and composition tasks, using the commercial continuous speech recognition systems of the day, and compared them against keyboard input on the same tasks.
Raw dictation ran at about 105 words per minute, which is in the same territory as the modern figures. Then they counted the corrections. The effective rate, including the time spent fixing what the recogniser got wrong, fell to 25 wpm for transcription and 8 wpm for composition. Typing in the same study came in at 33 wpm and 19 wpm.
| Study | Sample | Task | Speech | Typing | Result |
|---|---|---|---|---|---|
| Ruan et al. (2016), Stanford | 32 participants | Short message transcription, iPhone 6 Plus | 153 wpm | 52 wpm | Speech 2.93x faster |
| Blackley et al. (2020) | 10 physicians | Clinical notes into an EHR | Slightly faster | Baseline | Time similar, notes 78% longer |
| Karat et al. (1999), IBM | 24 users | Transcription, after corrections | 25 wpm | 33 wpm | Typing faster |
| Karat et al. (1999), IBM | 24 users | Composition, after corrections | 8 wpm | 19 wpm | Typing much faster |
On those 1999 numbers, typing beat dictation at both tasks, and beat it more than twice over on composition. That result is not a reason to keep typing in 2026, because the thing it measured has changed enormously: recognition accuracy. But it is the single most useful data point in this whole question, because of what it isolates.
The variable that actually decides this
Where the correction time goes



Fewer fixes is the whole game
BlabbyAI runs Whisper large v3 turbo and punctuates as you speak, so most of what you say lands ready to send. Ctrl+Space in any Windows app, or the Chrome extension in any browser tab. 60 credits a week free, no card.
The 1999 study measured something else worth keeping. Typing fell from 33 wpm when transcribing to 19 wpm when composing, a drop of more than 40%, and that has nothing to do with speech recognition at all. It is the cost of thinking.
This is the part most speed comparisons quietly omit, and it cuts against the marketing in a specific way. If your hands are not the bottleneck, making your hands faster does not help. Someone composing a difficult paragraph is limited by how fast they can work out what they think, and that runs at the same speed whether the words come out through fingers or a microphone.
It also explains the most common complaint about dictation, which is that it does not feel three times faster in daily use. It is not that the software is underperforming. It is that the 3x figure describes the part of writing you were never spending most of your time on.
The practical consequence is a sorting rule. The more of your writing day consists of text whose substance you already know, the more dictation is worth to you. That is a much larger share of most jobs than people assume: replies, updates, notes, tickets, summaries, and the ordinary traffic of work.
The most interesting result in this literature is not about speed at all. In a 2020 controlled observational study at Brigham and Women’s Hospital, ten physicians who had used speech recognition for at least six months documented simulated outpatient encounters, producing two notes per encounter, one dictated and one typed, in randomised order.
Documentation time came out similar between the two methods, with dictation slightly ahead. On a pure stopwatch reading, it was close to a draw. What differed was what came out.
320.6 vs 180.8 words
Dictated notes were 78% longer than typed notes from the same physicians on the same encounters.
170.9 vs 120.4 unique words
Dictated notes used a broader vocabulary, not simply more words.
7.7 vs 6.6 quality
Dictated notes scored higher on a rated scale of clarity, completeness and information sufficiency.
“Quality analysis supports the perception that SR allows for more detailed notes, but whether dictation is objectively faster than typing remains unclear, and participants described some scenarios where typing is still preferred.”
That is a more honest headline than any speed multiple, and it is the one nobody markets on. In the same amount of time, people said substantially more, and what they said was rated better. The return on dictation showed up as output rather than as saved minutes.
It fits the composition finding neatly. Speaking removes the friction that makes people write tersely, so instead of finishing the same note sooner, they wrote a fuller one. Whether that is a benefit depends entirely on whether you wanted a fuller note. For a clinical record or a handover document it plainly is. For an email to a busy colleague it may not be, which is an argument for editing rather than against dictating.
The Stanford study measured errors in two ways, and the split is worth understanding because it explains something people notice in practice. Speech made fewer errors during entry, 5.30% against 11.22%, but left slightly more in the final text, 1.30% against 0.79%.
The reason is that the two failure modes look different on the page. A typing error is usually a misspelling, which is visibly wrong and which a spellchecker underlines. A recognition error is usually a real, correctly spelled word in the wrong place. It reads smoothly, no tool flags it, and your eye slides over it on a reread because the sentence is grammatical.
The practical habit this implies
The 2020 clinical study found the opposite balance, which is worth noting rather than hiding: typed notes contained more uncorrected errors than dictated ones, 2.9 against 1.5 per note, though most were minor misspellings. Neither method is reliably cleaner. They are differently messy.
Putting the four studies together gives a usable sorting rule rather than a single number. Dictation’s advantage is largest when you know what you want to say, when the text is prose, and when you are somewhere you can speak. It shrinks or reverses as those conditions fail.
Dictation wins clearly
Typing still wins
The list on the left is longer than most people expect, and it has been getting longer. A decade ago dictation meant writing documents. Today a large share of everyone’s typing is short prose in a text field: Slack and Teams messages, ticket descriptions, code review comments, and increasingly prompts to AI tools, which are simply English sentences typed into a box. All of that dictates well.
The list on the right is real and we are not going to pretend otherwise. Nobody should dictate a spreadsheet formula. The point is that the two lists describe different parts of a working day, and most people spend more of theirs on the left than they think.
The first few days of dictation feel worse than the numbers promise, and the reason is not the software. Anyone who writes fluently has years of practice thinking at keyboard speed: composing in sentences, pausing mid-clause, glancing back at the previous line, revising a phrase before finishing it. None of those habits transfer directly to speaking.
A 2025 diary study followed twelve academic and creative writers using an LLM-assisted dictation tool over ten days, writing blog posts, diaries, screenplays, notes and fiction. What the researchers recorded was people developing strategies rather than simply getting faster: speaking a loose outline and expanding it, or dumping unstructured thoughts and organising them afterwards, instead of trying to say finished sentences.
“Through a ten-day diary study, we identified the participants’ in-context writing strategies using Rambler, such as how they expanded from an outline or organized their loose thoughts for different writing goals.”
That is the shape of the learning curve, and it is worth knowing before you start so you do not conclude after twenty minutes that dictation does not suit you. You stop trying to dictate the finished text and start dictating the raw material. The editing you then do is work you would have done anyway.
None of the studies above is about you, and the variation between people is large: the typing study of 168,000 volunteers found rates from 4 wpm to 158 wpm. If you type at 120 wpm the calculation is genuinely different from someone typing at 40. The only number that settles it is your own.
For the mechanics of running that test, our companion pieces cover the underlying numbers in more depth: words per minute speaking sets out where the speaking-rate figures come from, and what is a good typing speed covers the other half of the comparison.
Add to ChromeYes for getting words out, and the size of the gap depends on what you are doing. The most direct measurement is a 2016 Stanford study that had 32 people enter the same short messages both ways on an iPhone 6 Plus: speech ran 2.93 times faster in English, 153 words per minute against 52. That is the study behind almost every "3x faster" claim you will read. Two things narrow it. It measured transcription, where participants were given the text and simply said it aloud, so none of the time went on deciding what to write. And the keyboard it beat was a phone touchscreen, not a full-size keyboard, where average typing runs about 51.6 wpm rather than the 36 wpm typical of a phone. For work you already know you want to say, the advantage is large and real. For text you are still working out as you go, it shrinks, because thinking becomes the bottleneck in either mode.
Speaking runs roughly three to four times faster as a raw rate. Conversational speech was measured at 196 words per minute across 2,438 word-aligned telephone conversations, rising to 236 wpm once silences are stripped out. Typing was measured at 51.6 wpm across 168,000 people and 136 million keystrokes, with a standard deviation of 20.2 and a range from 4 to 158 wpm. On phones the typing figure drops to about 36.2 wpm. So the raw gap is about 3.8x against a physical keyboard and 5.4x against a phone. The head-to-head figure of 153 against 52 wpm from the Stanford experiment is the more honest one to quote for dictation specifically, because it measured both methods on the same task with the same people rather than comparing two separate studies.
Because raw entry rate is only one part of writing, and two other things eat the advantage. The first is composition. A 1999 IBM study at CHI measured the same people transcribing text that already existed and composing text themselves: typing fell from 33 wpm to 19 wpm when composing, and speech fell much further. Deciding what to say is slower than producing it in either mode, so the faster your input method, the more of your time is spent thinking rather than entering. The second is correction. In the Stanford comparison speech made fewer errors during entry, 5.30% against 11.22%, but left slightly more in the finished text, 1.30% against 0.79%. Time saved on entry comes back on the tidy-up. The practical result is a large gain on text you already know you want, and a smaller one on text you are still figuring out.
Yes, and it is worth knowing about because nobody quotes it. The 1999 IBM study by Karat, Halverson, Horn and Karat put 24 users through transcription and composition tasks on the commercial speech systems of the day. Raw dictation ran at about 105 wpm, but once the required corrections were made, the effective rate fell to 25 wpm for transcription and 8 wpm for composition. Typing in the same study averaged 33 wpm transcribing and 19 wpm composing. On those numbers typing won both tasks, by a wide margin on composition. The reason that result does not describe 2026 is the correction burden: those systems needed extensive fixing, and error rates have fallen enormously since. It is a useful historical control precisely because it shows the whole question turns on correction time rather than on speaking speed, which has not changed at all.
For most emails, yes, and email is close to the ideal case. A short work email is something you usually already know the substance of before you start, so the composition penalty that hurts dictation on harder writing barely applies. The message is also short enough that a stray recognition error is quick to spot and fix. The awkward part of email has never been the typing anyway, it is the tone and the framing, and that part costs the same in either mode. Where dictation clearly loses on email is anywhere you cannot speak out loud: an open-plan office, a train, a shared room. That is a constraint of the environment rather than of the method.
Not for the syntax itself, and this is the clearest case where typing still wins. Code is dense with punctuation, brackets, camel case and exact identifiers, all of which are slow and error-prone to say aloud and fast to type for anyone fluent with a keyboard. Dedicated voice-coding systems exist and they work, but they require learning a command language and weeks of practice. What has changed is that a great deal of programming in 2026 is prose: prompts to an AI coding agent, pull request descriptions, commit messages, code review comments, documentation and Slack replies to colleagues. That is ordinary English in a text field, and it dictates exactly as well as any other writing.
Yes, and more than most people expect, because the skill involved is not speaking but composing out loud. Fluent writers have years of practice thinking at keyboard speed, in sentences, with the ability to pause mid-clause and reread. Speaking a finished paragraph is a different motor and cognitive habit, and the first few days of dictation feel clumsy for that reason rather than because the software is failing. The 2025 Rambler diary study followed twelve academic and creative writers over ten days and found people developing distinct working strategies, such as speaking a loose outline first and expanding it, rather than trying to produce polished sentences on the first pass. That is the shape of the learning curve: you stop trying to dictate the finished text and start dictating the raw material.
The two make different kinds of mistakes, which matters more than which makes fewer. In the Stanford study, speech produced fewer errors during entry, 5.30% against 11.22% corrected error rate, but left marginally more errors in the final text, 1.30% against 0.79%. The difference is in kind: a typing error is usually a misspelling that looks obviously wrong, while a recognition error is usually a real word in the wrong place, which reads smoothly and is therefore easier to miss on a reread. The practical consequence is that dictated text needs a proofread more than typed text does, and needs it for different reasons. In the 2020 clinical study, typed notes actually contained more uncorrected errors than dictated ones, 2.9 against 1.5 per note, though most were minor misspellings.
It found something more interesting than a speed verdict. A 2020 controlled observational study at Brigham and Women’s Hospital had ten physicians document the same simulated patient encounters twice, once dictated and once typed, in randomised order. Documentation time came out similar between the two methods, with dictation slightly ahead, so on pure speed it was close to a wash. What differed was the output: dictated notes ran 320.6 words against 180.8 for typed notes, used more unique words, 170.9 against 120.4, and scored higher on quality at 7.7 against 6.6. In other words the physicians did not save much time, they produced substantially more complete documentation in the same time. That reframes the whole question: the return on dictation is often more output rather than less time.
It depends on what your writing day is actually made of, and the honest test takes about a week. If most of what you produce is text whose substance you already know before you start, which covers email, Slack and Teams messages, notes, tickets, first drafts, patient notes and prompts to AI tools, dictation will save you real time and the effect shows up quickly. If most of your writing time goes on deciding what you think, expect a smaller gain that arrives in a different form: getting a rough version out of your head faster and editing it afterwards, rather than producing finished prose at 150 words a minute. The other real constraint is whether you can talk out loud where you work. The way to settle it is to try it on your own work rather than on a typing test, which is why a free tier with no card is the sensible way in.
For the product side of this, dictation software is the category overview, Windows speech to text covers the desktop app, and the Chrome extension covers dictation in the browser with nothing to install.
The multiple is arguable. The direction is not.
Every study here agrees that getting words out is faster by voice, and that correction decides how much of that you keep. BlabbyAI is built around the correction half: accurate transcription, punctuation as you speak, and custom modes that format the text the way you wanted it. 60 credits a week free, no card.