Updated September 2, 2026 · By Sumbat.T

Words Per Minute Speaking: What the Research Actually Measured

A speedometer dial next to a studio microphone, illustrating the measurement of speaking speed

The short answer

  • People speak at about 196 words per minute in ordinary conversation. That is the measured figure from 2,438 word-aligned telephone conversations. Individual conversations ranged from 111 to 291 wpm.
  • Take the silences out and it rises to 236 wpm. Same corpus, same people, different definition. Roughly a fifth of conversational time is not speech, so which number is “the average” depends entirely on whether you count the gaps.
  • The 150 wpm figure you see everywhere is a convention, not a measurement. It comes from a National Center for Voice and Speech tutorial page with no sample size and no method attached. It is a reasonable target for a prepared presentation and a low estimate for a conversation.
  • Typing runs at 51.6 wpm across 168,000 people and 136 million keystrokes, and 36.2 wpm on a phone. That gap is the entire argument for dictation, and it is why BlabbyAI exists: a Windows app and a Chrome extension that type what you say into any text field, from $8.49 a month.

There is a number that appears in almost every article about speaking speed, in most presentation guides, and in the AI summary Google shows above the search results: 150 words per minute. It is repeated so consistently that it has the texture of a fact.

It is not exactly wrong. But it is not a measurement either, and the research that did measure conversational speech carefully arrived somewhere quite different. This piece sets out where each number in this topic actually comes from, which ones have a study and a sample size behind them, and which ones are folklore that has been copied between websites for long enough to look authoritative.

The practical reason to care is the gap between speaking and typing. Whatever figure you accept for speech, it is several times faster than anyone types, and that difference is the whole basis of dictation software. It is also routinely overstated, so the second half of this piece is about what the gap is really worth once composition and correction are accounted for.

Input rate by method, in words per minute

Every figure below comes from a named study. Sample sizes vary from 168,000 people down to 12, which is why the source column matters as much as the number.

Conversation, silences excluded

236 wpm

Yuan, Liberman & Cieri (2006) · 2,438 aligned Switchboard conversations

Conversation, total elapsed time

196 wpm

Yuan, Liberman & Cieri (2006) · Same corpus, range 111 to 291 wpm

Reading aloud

183 wpm

Brysbaert (2019) meta-analysis · 77 studies, 5,965 participants

Dictating into a phone

153 wpm

Ruan et al. (2016), Stanford · 32 participants, short messages

Presentation pace (convention)

150 wpm

National Center for Voice and Speech · Rule of thumb, no sample size given

Typing, physical keyboard

51.6 wpm

Dhakal et al. (2018), CHI · 168,000 people, 136M keystrokes

Typing, mobile keyboard

36.2 wpm

Palin et al. (2019), MobileHCI · 37,370 volunteers

Handwriting, copying

22 wpm

Brown (1988), HFES · Small study, 12 subjects

What people actually measured: 196 wpm, and 236 without the pauses

The most careful measurement of conversational speaking rate in English comes from a 2006 Interspeech paper by Jiahong Yuan, Mark Liberman and Christopher Cieri at the University of Pennsylvania and the Linguistic Data Consortium. They used the version of the English Switchboard corpus corrected and aligned at ICSI, comprising 2,438 conversations, which means every word has a timestamp rather than a transcript sitting loosely over an audio file.

That alignment is what makes the study useful, because it lets you calculate speaking rate two ways and see how much the definition matters. In the authors' own words: adding up the number of words spoken by both participants and dividing by the total elapsed time of the conversations gives “an overall average rate of 196 words per minute”. Using the alignments to exclude silences and non-speech noises gives “an average net speaking rate of 236 WPM”.

196 wpm

The gross rate. Total words divided by total time, silences included. Individual conversations ranged from 111 to 291 wpm.

236 wpm

The net rate, with silences and non-speech noises removed. Minimum 158 wpm, maximum 312 wpm.

The distance between those two figures is the interesting part. About 40 words a minute, roughly a fifth of the total, is silence: breathing, thinking, waiting for the other person, the small gaps that make speech sound like speech rather than a recitation. Both numbers describe the same recordings. They differ only in whether the pauses count.

Why this matters if you are quoting a number

“The average person speaks at X words per minute” is an incomplete sentence. It needs to say whether X includes the silences, and whether the setting was a conversation, a presentation or a reading.

The same study found the variation between people is large: 111 to 291 wpm across conversations, a spread of more than two and a half times. Any single average is hiding that.

The same paper found two smaller effects worth knowing. Men spoke slightly faster than women, by about 4 to 5 words per minute or roughly 2%, which is a much smaller difference than the stereotype suggests. Older speakers were slower. And the topic of conversation moved the average meaningfully: across the assigned discussion topics in the corpus, average speaking rate ranged from 152 wpm to nearly 170 wpm, so what you are talking about changes how fast you talk about it.

Where the 150 wpm figure comes from, and what it is good for

The 150 wpm number is real, in the sense that it genuinely appears where people say it does. The National Center for Voice and Speech, a respectable research organisation, states in a tutorial page that the average rate of speech for English speakers in the United States is about 150 words per minute. That is the source almost everyone is ultimately citing, whether or not they know it.

What that page does not include is a sample size, a methodology, a corpus, or a citation to a study. It is a rule of thumb offered in an educational context, and it is a perfectly sensible one. It is simply not the same kind of object as a corpus measurement of 2,438 aligned conversations, and the two get quoted interchangeably as though they were.

There is also a good reason 150 feels right to people, and it is worth saying plainly rather than treating the figure as simply wrong. Presentations are slower than conversations. When you speak to an audience you pause for emphasis, you slow down for the parts that matter, and you leave room for people to follow. Somewhere around 130 to 150 wpm is a comfortable pace for a prepared talk. The number is good advice for the situation it came from. The error is generalising it from the podium to the telephone.

FigureType of claimEvidence behind it
236 wpmConversation, silences removedPeer-reviewed, 2,438 word-aligned conversations
196 wpmConversation, total timeSame study, same corpus
183 wpmReading aloudMeta-analysis, 77 studies, 5,965 participants
150 wpmGeneral conventionInstitutional tutorial page, no sample size stated

One further set of numbers deserves scepticism. Search for speaking rate and you will find confident tables giving separate figures for audiobook narration, podcasting, broadcast news and auctioneering. Those tables contradict each other from site to site, and none of the ones in wide circulation cite a study. They appear to be folklore that propagates between blogs. The one solidly measured figure in that territory is reading aloud at 183 wpm, from a 2019 meta-analysis by Marc Brysbaert covering 77 studies and 5,965 participants, which is a reasonable anchor for any kind of scripted delivery.

The other half of the comparison: typing at 51.6 wpm

The typing side of this comparison is on much firmer ground, because somebody ran the largest study of the question ever attempted. In 2018, Vivek Dhakal, Anna Feit, Per Ola Kristensson and Antti Oulasvirta published “Observations on Typing from 136 Million Keystrokes” at CHI, based on 168,000 volunteers from more than 200 countries.

The headline figure is a mean of 51.56 words per minute, with a standard deviation of 20.20. The range ran from 4 wpm to 158 wpm, and the fastest participants reached around 120 wpm. As with speech, the spread is the story: the average describes almost nobody, and the difference between a slow typist and a fast one is far larger than the difference between a slow speaker and a fast one.

51.6 wpm

The average across 168,000 people. Standard deviation 20.2, so the typical person is somewhere between 31 and 72 wpm.

36.2 wpm

Phone typing, from a separate study of 37,370 volunteers. Roughly 70% of physical keyboard speed.

40 to 70%

The share of keystrokes fast typists perform with rollover, pressing the next key before releasing the last.

Two findings from that study are worth repeating because they contradict what most people assume. The first is that formal training barely mattered: participants who had taken a typing course performed very similarly to those who never had, despite the trained group using more fingers. The second is what did separate the fast group, which was rollover, overlapping one keystroke with the next. Fast typists were not moving their fingers faster so much as running them in parallel.

The mobile figure comes from a companion study, “How do People Type on Mobile Devices?”, covering 37,370 volunteers. It found an average of 36.2 wpm with 2.3% uncorrected errors, with over 74% of people typing with both thumbs, which was significantly faster than any other technique. The fastest mobile typists were aged 10 to 19, which will surprise nobody who has watched a teenager text.

So how much faster is speaking? About three to four times

Put the two sides together and the arithmetic is straightforward. Against the 196 wpm conversational figure, speech runs roughly 3.8 times faster than the average physical keyboard and about 5.4 times faster than a phone keyboard. Against the more conservative 150 wpm convention it is closer to 2.9 times.

Those are back-of-envelope comparisons between separate studies, though, and there is a better number available. In 2016 a team from Stanford, Baidu and the University of Washington ran the comparison properly: same people, same task, both input methods. Sherry Ruan, Jacob Wobbrock, Kenny Liou, Andrew Ng and James Landay had 32 participants enter short messages on an iPhone 6 Plus, by voice and by keyboard, and measured both.

The controlled comparison

  • English: speech was 2.93 times faster, at 153 wpm against 52 wpm for the touchscreen keyboard.
  • Mandarin: speech was 2.87 times faster, at 123 wpm against 43 wpm.
  • Errors went both ways. Speech produced fewer errors during entry (5.30% against 11.22%) but left slightly more in the finished text (1.30% against 0.79%).

There is a detail in that study which almost every article citing it gets wrong, and it is worth stating clearly. The keyboard being beaten three-to-one was a smartphone touchscreen keyboard, not a physical one, and the task was transcribing short messages rather than composing original text. The paper is frequently cited under its original working title, “Speech Is 3x Faster than Typing”, which invites people to read it as a claim about desktop typing in general. It is not that. The finding is solid and the study is good; the scope is narrower than the headline suggests.

The gap, in practice

ChatGPTGoogle DocsGmailWhatsAppMicrosoft Word

Microphone

Type at the speed you talk

BlabbyAI turns speech into text in any Windows application on Ctrl+Space, or in any browser tab through the Chrome extension. Whisper large v3 turbo, no voice training, no per-minute billing. 60 credits a week free, no card.

Add BlabbyAI to Chrome

Why dictation does not feel three times faster in practice

Anyone who has actually used dictation for real work will have noticed that it does not deliver a clean 3x. The honest explanation has two parts, and understanding them is what separates people who get value out of dictation from people who try it for a week and stop.

The first is composition. Transcribing text that already exists is a different task from working out what to say, and the gap is enormous in both modalities. A 1999 IBM study by Karat and colleagues measured keyboard users at around 33 wpm when transcribing but only 19 wpm when composing. The bottleneck moves from your fingers to your head. If you are staring at a blank document deciding what you think, dictation cannot make you think faster, and the input rate stops being the limiting factor at all.

The second is correction. Every error you fix eats into the time you saved. This is where modern speech recognition has changed the arithmetic most, because the 1999 numbers came from an era when dictation meant training a voice profile for twenty minutes and correcting constantly. Today's models need no training at all. But the principle holds: the speed you get is entry speed minus cleanup, and cleanup depends on your microphone, your accent, and how much background noise you are working in.

What this means for whether dictation will help you

The speedup is largest on text you already know you want to write: replies, notes on something you just did, messages, documentation of work already done, a first pass at something you have thought about.

It is smallest on text you are still working out while you write it. That is not a flaw in the software, it is where the bottleneck moved to.

A note on accuracy, because the marketed numbers are optimistic

Speech recognition accuracy is usually quoted as a word error rate, and the figure that circulates is around 2%. That number is real but it comes from LibriSpeech test-clean, a benchmark of clearly read audiobook passages recorded in good conditions. It is the easy case by design.

Across the mixed datasets of the Open ASR Leaderboard, OpenAI's Whisper large-v3 averages a word error rate of 7.44. On AMI, a corpus of real meeting recordings with overlapping speakers and room noise, the same model reaches 15.95. So a single accuracy percentage tells you almost nothing without knowing what it was measured on, and anyone quoting 99% accuracy is quoting the best case.

The good news for dictation specifically is that dictation is close to the easy case. One person, speaking deliberately, into a microphone near their mouth, in a room they chose. That is much closer to the audiobook benchmark than to the meeting-room one, which is why dictation accuracy in practice tends to be good and why accuracy is rarely the thing that makes people give up.

The finding that reframes the whole question: 39 bits per second

One more study is worth knowing about, because it suggests that words per minute is a slightly odd thing to measure in the first place.

In 2019, Christophe Coupé, Yoon Mi Oh, Dan Dediu and François Pellegrino published a paper in Science Advances examining 17 languages from 9 families, using 170 native adult speakers and around 240,000 syllables. They measured two things: how fast people produce syllables, and how much information each syllable carries.

Syllable rate varied enormously, from 4.3 to 9.1 syllables per second depending on the language and speaker. But the languages spoken quickly turned out to carry less information per syllable, and the slower ones carried more. The two effects cancelled. The rate of information actually transmitted landed at roughly 39 bits per second across all 17 languages, with a standard deviation of just 5.1.

Why this matters for a dictation question

It suggests the ceiling on speaking is not really in the mouth. If humans converge on the same information rate regardless of how fast the language is spoken, then the constraint is cognitive rather than mechanical. Which is the same conclusion the composition research reaches from the other direction: the limit on how fast you can produce text is how fast you can decide what it should say.

BlabbyAI logo

Closing the gap: dictation on Windows and in the browser

If the numbers above make the case for talking instead of typing, the practical question is what actually types the words. BlabbyAI is built for exactly the case this article describes: not transcribing a recording after the fact, but replacing the keyboard while you work.

On Windows it is a desktop app with a global shortcut. Press Ctrl+Space in any application, speak, press it again, and the text appears in whatever field the cursor is in. That works in Word, in Slack, in a browser tab, in an IDE, in a hospital records system, anywhere Windows will accept typed characters. The transcription runs on Whisper large v3 turbo, which handles punctuation and casing without being told, and there is no voice profile to train before you start.

The BlabbyAI shortcut settings window showing Ctrl+Space bound to a dictation mode

The shortcut panel. Ctrl+Space by default, and you can bind additional keys to specific modes.

The second product is a Chrome extension, and it is not a lesser version of the first. It is the right answer for a locked-down work laptop where you cannot install desktop software, for a Chromebook, and for anyone on macOS. It does the same job inside the browser, which for a lot of people is where most of their typing happens anyway: Gmail, Google Docs, Notion, Linear, a CRM, a support queue.

The feature that matters most for the composition problem described above is custom modes. A mode is a plain-English instruction applied to what you said, so instead of transcribing you verbatim it can clean up filler, restructure a rambling thought into bullet points, or format the same spoken sentence as a formal email in one mode and a short Slack message in another. That is aimed squarely at the gap between speaking quickly and producing usable text, which is the real bottleneck rather than the raw word rate.

The BlabbyAI mode editor, showing a written instruction that transforms dictated speech

A custom mode is written as an instruction, not programmed as a macro.

On pricing, the free tier is 60 credits a week, which is roughly 2,000 words of transcription and enough to find out whether dictation suits how you work. Unlimited use is $8.49 a month. There is no per-minute billing, because this is a dictation tool rather than a transcription service.

What else is available, and what it costs

Dictation software is a reasonably crowded category now, and the options divide fairly cleanly by platform and by price. The free built-in option is genuinely worth trying first, because if plain transcription is all you need then you already have it.

ToolPricePlatformNotes
BlabbyAI logoBlabbyAI
Free tier, then $8.49/moWindows app + Chrome extensionWhisper large v3 turbo, custom modes, Ctrl+Space
Windows Win+H logoWindows Win+H
Free, built inWindows 11 and 10Unlimited, but no vocabulary or formatting control
Wispr Flow logoWispr Flow
Free tier, then $15/momacOS, Windows, mobileFree tier caps at 2,000 words a week
Superwhisper logoSuperwhisper
Free tier, Pro $8.49/momacOS, Windows, iOSCan run models locally
Dragon Professional v16 logoDragon Professional v16
Quote only, no published priceWindowsSold through a sales conversation

Windows dictation on Win+H is free, unlimited and already installed, and for straight transcription it is decent. What it does not give you is any control over the output: no custom vocabulary for the terms you use all day, no formatting rules, and no way to transform what you said rather than simply transcribe it. That is the line between the built-in tool and a paid one, and it is worth testing before you spend anything.

The Windows Win+H dictation bar open on the desktop

Windows dictation on Win+H. Free and unlimited, but the words arrive raw.

How to measure your own rate, and what to do with it

Every number in this article is an average over other people. Your own rate is more useful, and it takes two minutes to find.

Speak for one minute about something you actually need to write, not a prepared passage, and count the words. Then type the same content for one minute and count those. The ratio between them is your personal version of the comparison this whole article is about, and it will be more informative than 196 against 51.6, because it includes your accent, your keyboard, your subject matter and your habits.

  • Measure spontaneous speech, not reading. Reading aloud has its own benchmark, 183 wpm, and it is faster and steadier than working out a thought while you speak.
  • Count a full minute. Speech is bursty, and a ten-second sample will flatter or punish you depending on where the pauses fall.
  • Do it on real work. The rate that matters is the one you hit on the actual emails, notes and documents you produce, not on a passage chosen to be easy to say.

If the gap you find is large and most of what you write is content you already know you want to say, dictation will pay for itself quickly. If you spend most of your writing time deciding what you think, the gain will be smaller and it will come from somewhere else: getting a messy first version out of your head faster, and tidying it afterwards.

Add to Chrome

Common questions

How many words per minute do people speak?

It depends on whether you count the silences, and that single decision moves the answer by about 40 words a minute. The most carefully measured figure comes from a 2006 Interspeech study by Yuan, Liberman and Cieri, which analysed 2,438 word-aligned English telephone conversations from the Switchboard corpus. Counting every word spoken by both participants and dividing by the total elapsed time, the overall average was 196 words per minute, ranging from 111 to 291 wpm across individual conversations. When the researchers used the word alignments to strip out silences and non-speech noises, the average "net" speaking rate rose to 236 wpm, with a minimum of 158 and a maximum of 312. The widely quoted 150 wpm figure comes from the National Center for Voice and Speech and is best described as a convention rather than a measurement: it appears in a tutorial page with no sample size and no methodology attached. Both numbers are in circulation, and the honest way to use them is to say which one you mean.

Is 150 words per minute a good speaking speed?

For a presentation, yes, and that is the context the number really belongs to. Speaking to an audience is slower than talking to a friend on the phone, deliberately: you pause for emphasis, you let a point land, and you give people time to follow. Around 130 to 150 wpm is a comfortable delivery pace for a prepared talk. The confusion arises because that presentation figure gets quoted as the average rate of human speech in general, which the conversational data does not support. In ordinary two-way conversation the measured rate is closer to 196 wpm, and above 236 wpm once you discount the gaps. So 150 wpm is a good target for a speech and a low estimate for a conversation, and treating those two situations as one number is where most articles on this topic go wrong.

How much faster is speaking than typing?

Roughly three to four times, depending on which typing you compare it to. The largest typing study available measured 168,000 volunteers across 136 million keystrokes and found an average of 51.6 words per minute, with a standard deviation of 20.2 and a full range of 4 to 158 wpm. A separate mobile study of 37,370 people found 36.2 wpm on phones. Against the 196 wpm conversational figure, speech runs about 3.8 times faster than a physical keyboard and 5.4 times faster than a phone. The most direct measurement of the gap comes from a controlled 2016 study at Stanford, which had people enter the same short messages both ways on an iPhone: speech came out 2.93 times faster in English, 153 wpm against 52 wpm. That last comparison is the honest one to quote, because it measured both modalities on the same task with the same people. Capturing that gap in real work is what a dictation tool like BlabbyAI is for: it types what you say into whatever app you are already in.

What is the average typing speed?

51.6 words per minute, from the largest dataset ever collected on the question. The 2018 CHI paper "Observations on Typing from 136 Million Keystrokes" by Dhakal, Feit, Kristensson and Oulasvirta recorded 168,000 volunteers from more than 200 countries and found a mean of 51.56 wpm with a standard deviation of 20.20. The spread matters more than the average: participants ranged from 4 wpm to 158 wpm. The fastest people in the study reached about 120 wpm. One finding surprised the researchers: people who had taken a formal typing course typed at very similar speeds to those who never had, despite using fewer fingers. What actually separated fast typists was rollover, pressing the next key before releasing the previous one, which accounted for 40 to 70% of keystrokes among the quickest participants.

Does everyone speak at the same rate?

No, and the variation is larger than most summaries admit. In the Switchboard analysis, individual conversations ranged from 111 wpm to 291 wpm, which is a spread of more than two and a half times between the slowest and fastest. The same study found that males speak slightly faster than females, by around 4 to 5 words per minute or about 2%, and that older speakers are slower. Conversation topic mattered too: average rates across different assigned topics ranged from 152 to nearly 170 wpm within the same corpus. There is also a finding at the language level worth knowing. A 2019 Science Advances study of 17 languages and 170 native speakers found syllable rates varying from 4.3 to 9.1 syllables per second, and yet the rate of information transmitted converged on roughly 39 bits per second across all of them. Languages spoken quickly tend to pack less meaning into each syllable, so the throughput evens out.

If speaking is so much faster, why does dictation not feel three times faster?

Because raw entry speed is not the whole task, and the honest answer here is more useful than the marketing one. Two things eat the advantage. The first is composition: thinking about what to write is slower than the mechanical act of producing it, in either modality. A 1999 IBM study measured keyboard users at 33 wpm when transcribing text that already existed but only 19 wpm when composing it themselves, and the same collapse applies to speech. The second is correction. In the Stanford comparison, speech produced fewer errors during entry (5.30% against 11.22%) but left slightly more in the final text (1.30% against 0.79%), so some of the time saved goes back into fixing things. The practical result is that dictation gives you a real and substantial speedup on anything you already know you want to say, and a smaller one on text you are still working out as you go.

How accurate is speech recognition now?

Good enough that accuracy is rarely the thing that stops people, though the numbers marketed are usually better than the ones you will get. Modern systems are built on models like Whisper, trained by OpenAI on 680,000 hours of audio for the original release and far more for later versions. The figure you often see quoted, around 2% word error rate, comes from LibriSpeech test-clean, a benchmark of clearly read audiobook passages that is deliberately easy. Across the mixed datasets of the Open ASR Leaderboard the same Whisper large-v3 model averages 7.44 word error rate, and on AMI meeting recordings it reaches 15.95. So accuracy depends heavily on your microphone, your accent and your background noise, and a single headline percentage tells you very little. For one person dictating deliberately into a decent microphone, which is what dictation software actually is, real-world accuracy sits at the good end of that range.

How do I measure my own speaking rate?

Read something aloud for exactly one minute and count the words, then do it again while talking normally rather than reading, because the two will differ. Reading aloud has its own measured benchmark: a 2019 meta-analysis by Brysbaert, covering 77 studies and 5,965 participants, put average oral reading at 183 wpm. Spontaneous speech is messier, with restarts and pauses, so counting a full minute of it is more representative than counting a sentence. If you want the number that matters for dictation specifically, measure the second one, and measure it on your own real work rather than on a prepared passage. The rate you speak at when you are working out a thought is the rate that determines how much dictation will actually save you.

Does speaking faster mean communicating more?

Not across languages, which is one of the more surprising findings in the research. The 2019 Science Advances study measured both syllable rate and information density across 17 languages from 9 families, using 170 native adult speakers and around 240,000 syllables. Speech rate varied widely: Japanese and Spanish speakers produce syllables quickly, while languages like Vietnamese and Thai are slower. But the faster languages carry less information per syllable, and the slower ones carry more, so the information rate landed at about 39 bits per second everywhere, with a standard deviation of just 5.1. Within a single language the picture is different: speaking faster does move more words, up to the point where listeners stop keeping up. For dictation the relevant limit is not how fast you can physically talk but how fast you can think of what to say next.

What speaking rate should I use for a podcast or a video?

The rates that get quoted for narration, podcasting and broadcast, usually somewhere between 140 and 180 wpm, are worth treating with some caution: they circulate widely between blogs without a study behind them, and different sources give contradictory ranges. The one solidly measured number in that territory is oral reading at 183 wpm, from Brysbaert's 2019 meta-analysis of 77 studies. That is a reasonable anchor for scripted delivery. Beyond it, the practical guidance is straightforward: slower for complex or unfamiliar material, faster for light material and for audiences who can rewind. If you are producing something and want a target, record two minutes at your natural pace, count the words, and adjust from there rather than from a number you read somewhere.

Sources

For more on turning the speaking-speed advantage into working software, dictation software is the category overview, Windows speech to text covers the desktop side, and the Chrome extension covers dictation in the browser.

You already talk at three times the speed you type

BlabbyAI is the part that turns that into text: Ctrl+Space on Windows, or the Chrome extension anywhere you cannot install software. No voice training, no per-minute billing. 60 credits a week free.

Add BlabbyAI to Chrome