Updated September 6, 2026 · By Sumbat.T

There is a specific reason this question is harder to answer than it looks, and it is not that the information is missing. It is that the word “voice to text” covers two different jobs, most of the available advice quietly answers the wrong one, and the evidence people cite about what works has been passed around for so long that it no longer says what it is used to say. This page separates the two jobs, reports what each of the recommended plugins actually does according to its own documentation, and is honest about the one question nobody has actually settled.
Start here, because getting this wrong is what makes people install three plugins and conclude that dictation in Obsidian is bad.
Dictation and transcription are not the same product
Dictation means your words appear as text at your cursor while you are speaking. The note is finished when you stop talking.
Transcription means you record audio, stop, hand the file to a model, wait, and the text arrives afterwards. It is a good workflow for a voice memo recorded away from your desk, and a bad one for the ordinary act of writing a note by thinking out loud.
Both are legitimate. They are simply different tools, and the search results for this topic do not distinguish between them at all. Someone who wants to talk their way through a daily note gets recommended a plugin that requires them to record a clip, stop, and wait for an API round trip. It works exactly as documented, and it is not what they wanted.
So the useful version of the question is not what is the best voice to text plugin for Obsidian. It is do you want to talk while you write, or process a recording afterwards, and then, separately, is Obsidian the only place you need this. Those two answers pick the tool between them.
These are the four plugins that come up most often for this query, including in Google's own AI summary of it. Every entry below is from the plugin's own listing or README rather than from a review of it, because the difference between the two turns out to matter.
| Plugin | API key | Cost | Live dictation |
|---|---|---|---|
| Speech Kit | No | Free, runs locally | Yes |
| Speech to Text | Yes | Plugin free, you pay the API | No |
| Whisper | Yes | MIT plugin, you pay the API | No |
| Voice MD | Yes | Plugin free, you pay OpenAI | No |
Read the last column first. Three of the four do not do the thing most people are asking for. They are transcription tools, correctly built and clearly documented, that have been collectively filed under a word that means something else.
The only one of the four with true live dictation. Desktop only: Windows, macOS, Linux. Works offline once models are installed.
OpenAI Whisper or Deepgram. Its own listing quotes Deepgram Nova-3 at $0.0043/min. Live microphone transcription is described as not included in the published 1.0.20 release.
Works with OpenAI, Groq, Azure or any Whisper-compatible endpoint. Alt+Q starts and stops the recording, then the transcription lands at the cursor.
OpenAI key stored in Obsidian SecretStorage. Buffers to IndexedDB so a failed upload can be retried. 25 MB upload limit per file.
Credit where it is due: Speech Kit is good
If you want dictation that lives inside Obsidian and you want it free, Speech Kit is the answer, and it is worth saying so plainly. No API key, no account, no per-minute billing, runs on your own machine, works offline once the models are downloaded, and the words stream in as you speak.
Its limitation is not quality or price. It is scope: it works in Obsidian, and only in Obsidian. That is the right trade if Obsidian is where your writing happens. It is the wrong one if your notes are one window among many, which is the case this page turns to next.
The three key-requiring plugins are all free as software, which is how they get described as free. The setup they actually ask for is worth spelling out, because it is usually presented as a single step and it is not:
The per-minute rates are genuinely small. Speech to Text's own listing quotes Deepgram Nova-3 at $0.0043/min, which is a few cents an hour. The friction is not the money, it is the four steps and the ongoing responsibility for a billing relationship that exists so you can talk to your own notes.
No key, no billing, no plugin



Dictate into Obsidian without setting up an API account
BlabbyAI runs at the operating-system level, so the same shortcut types into an Obsidian note, your email, your browser and your chat. Custom modes turn a spoken thought into bullets or a task list as it arrives. Windows desktop app and Chrome extension. 60 credits a week free, no card.
If you are on Windows, the obvious thought is to skip plugins entirely and use the voice typing built into the operating system. Search for whether that works in Obsidian and you will find confident answers in both directions. Having gone through the primary sources, here is what is actually established, because the honest answer is more useful than a confident one.
Nearly every claim that dictation fails in Obsidian traces back to one forum thread. Here is its opening post, in full, because the framing matters:
“I would like to dictated (speech to text) using the Mac OS dictation feature. It works in a lot of places and applications in Mac OS but not obsidian. Is there a special way to get this to work?”
Obsidian Forum, “Dictation - speech to text”, 25 January 2021
Mac OS dictation feature. This is a macOS report, from January 2021, and it says nothing whatsoever about Windows. It gets quoted as though it were a general statement about Obsidian, and Windows users read it and conclude their own case is settled. It is not.
The replies under it also disagree with each other, which is more informative than a clean failure would have been. One user reports it working “around 10% of the time in Obsidian and 100% of the time everywhere else”. Another says they get dictation in Obsidian just fine. A third describes it working once or twice and then stopping. A separate bug report from November 2020 describes a long delay during which Obsidian appears unresponsive, and it was archived without a documented resolution.
Intermittent is worse than broken
A feature that fails cleanly teaches you in one attempt. A feature that works ten percent of the time cannot be relied on and cannot be abandoned, and in note-taking specifically it has a nasty failure mode: you do not find out the thought was never captured until you go looking for it later.
There is a 2024 Obsidian forum thread that gets linked as proof Win+H fails. Read it and it is about remapping the shortcut, not about dictation not working. The author says they use Windows voice dictation a lot and that it is activated with Win+H, and their complaint is that PowerToys and AutoHotkey cannot rebind that shortcut to another key inside Obsidian while the same rebinding works in Notepad and Word. That is a keyboard-hook problem in a desktop application window. It is not evidence about whether dictated text reaches the note.
The nearest thing to a direct first-hand Windows comment on the forum criticizes Win+H for being weak on accuracy and for supporting few languages besides English. That is a complaint about quality, which implies the thing runs, but implication is not confirmation and this page is not going to present it as one.
Microsoft's own documentation for voice typing requires only that you have your cursor in a text box, plus a microphone and an internet connection. They publish no application compatibility list and make no “works in every app” guarantee, so their documentation cannot be used to prove the Obsidian case in either direction.
The honest state of it
Whether Win+H reliably types into an Obsidian note on Windows today is not established by any primary source we could find. The macOS evidence is real but is from 2021 to 2023 and is contradictory. The Windows evidence is a misread thread and one comment about accuracy.
It is also a sixty-second test on your own machine, which beats every citation available: open a note, put the cursor in it, press Win+H, and say a sentence. That is the actual answer for your setup, and no article can give it to you.
The reason to lay this out rather than pick a side is that it points at the real decision. If a dictation route can be silently unreliable inside one specific application, the durable fix is a tool that does not depend on that application at all. BlabbyAI has its own Windows desktop app and puts text into whichever field has focus, so its behavior in Obsidian is the same as its behavior in Word, in a browser or in a terminal, rather than depending on how one application handles system dictation.
Worth one short section, because the contrast explains the whole desktop problem in a sentence. On a phone, the keyboard does the dictation, not the app. You tap the microphone on the keyboard, the system transcribes, and the keyboard inserts characters as though you had tapped them, so Obsidian is never in the audio path and has nothing to get wrong.
Desktop has no equivalent universal keyboard layer, which is precisely why the desktop version of this question is a mess and the phone version is not. That missing layer is the job a system-level dictation tool does on a computer. BlabbyAI puts the text into the focused field itself rather than asking the application to cooperate, which is the same mechanism your phone keyboard uses and the reason it does not care which window is in front. It is a Windows desktop app and a Chrome extension, so this is the desktop answer rather than a phone one.
Every option so far has been scoped to Obsidian. Before settling on one, it is worth asking how much of your writing day Obsidian actually is.
For most people who keep a personal knowledge base, the notes are the part they think of as writing, and they are a minority of the words they type. The rest is email, chat, issue trackers, documents, browser forms, commit messages and search boxes. A plugin, by definition, covers none of that. Install Speech Kit and you can talk to your notes and then go back to typing everything else. If you are weighing that trade, our comparison of the best dictation apps covers the system-level options side by side, and whether dictation is actually faster than typing goes through what the studies measured.
| Approach | Live dictation | Works outside Obsidian | Formats as you speak | Setup | Cost |
|---|---|---|---|---|---|
| BlabbyAI | Yes | Yes | Yes | Install, no API key | Free tier, $8.49/mo unlimited |
| Speech Kit plugin | Yes | No | No | Install, download models | Free |
| Whisper / Voice MD / Speech to Text | No | No | No | API account, billing, key | Metered per minute |
| Windows voice typing (Win+H) | System-wide; unverified in Obsidian | System-wide; unverified in Obsidian | No | Built in | Free |
Win+H is in that table honestly: it is free, it is already installed, and it is system-wide. Its weaknesses are the ones its own users name, which are accuracy on unusual vocabulary and thin language support, plus the unresolved question above about how it behaves in this particular application. For a note full of proper nouns, book titles, technical terms and the names of people you know, accuracy on unusual vocabulary is close to the entire job. There is more detail on that option in our piece on Windows voice typing, if you want to know what it handles before you test it.
There is one more gap between transcription and a usable note, and it is the one that decides whether dictated notes actually get used. Speech, transcribed faithfully, is a wall of text. Notes are useful because they have shape: a heading, some bullets, a task, a link to another note. Cleaning that up by hand afterwards gives back the time the dictation saved.
BlabbyAI custom modes are free-form AI instructions applied to what you said, so the text can arrive already shaped. You pick the mode before you speak. Three examples, with the spoken input on the left:
Recap, typed out after a call
You said: “quick note to myself after the pricing call, they want the enterprise tier but they pushed back hard on the seat minimum, marcus is sending over their current contract this week and we need to come back with a revised quote before the fifteenth”
**Pricing call** - Interested in enterprise tier - Pushed back on the seat minimum - Marcus sending current contract this week - [ ] Revised quote before the 15th
Daily note, cleaned up
You said: “spent most of today on the migration script, the tricky part was that the old records don't have a created timestamp so I had to infer it from the id sequence, it works but I'm not confident about the rows from before the two thousand nineteen import”
Spent most of today on the migration script. The tricky part: old records have no created timestamp, so I inferred it from the ID sequence. It works, but I am not confident about rows from before the 2019 import.
Literature note
You said: “the argument in chapter four is basically that attention is a finite resource and that the interface designs we take for granted are optimized to consume it, which connects to the thing I wrote about notification batching last month”
Chapter 4 argues attention is a finite resource, and that common interface designs are optimized to consume it. Connects to: [[Notification batching]]
Because a mode is a free-form instruction rather than a fixed feature, it also covers spoken commands you define yourself. If you want saying “new paragraph” to break the line rather than type the words, that is a line in the mode. If you want technical vocabulary left strictly alone, that is another.
Given everything above, here is the order that makes sense for most people, and the reasoning behind each step.
If you only take one thing from this page
Check whether the tool you are about to install does dictation or transcription, and check whether it works only in Obsidian. Those two questions eliminate most of the disappointment on this topic, and neither of them is about accuracy scores.
No. Obsidian ships no built-in dictation or speech-to-text feature, and its official help documentation does not document any dictation feature. Every route to voice in Obsidian is something you add: either a community plugin, or a tool that runs outside Obsidian at the operating-system level and types into whatever field has focus. That second route is worth understanding before you start installing plugins, because it is the one that also covers every other application you type in during the day. BlabbyAI works that way: you press a shortcut, speak, and the words are typed where your caret is, whether that caret is in an Obsidian note, a browser tab, Slack or your terminal. There is a Windows desktop app and a Chrome extension, and the free tier is 60 credits a week with no card required.
It depends on whether a plugin is what you want, because a plugin only ever works inside Obsidian. If you want the same shortcut everywhere you type, BlabbyAI does it outside Obsidian too, with no API key and no per-app setup. If you do want a plugin specifically, Speech Kit (listed in the community directory as local-dictation) is the strongest of the four commonly recommended ones, and it is the only one of them that does live dictation where words appear as you speak. It is free, it runs locally and offline once the models are downloaded, it needs no API key, and it supports Windows, macOS and Linux on desktop. The other three commonly recommended plugins, Speech to Text, Whisper by nikdanilov, and Voice MD, all require you to bring your own paid API key and all three are record-then-transcribe rather than live dictation. The honest limitation of any plugin, Speech Kit included, is scope: it works in Obsidian and nowhere else. If the notes in Obsidian are only part of what you write each day, a system-level tool such as BlabbyAI covers Obsidian plus your email, your browser, your chat and your documents with one shortcut.
Yes, several, and they differ in ways worth knowing. Speech Kit is a free community plugin that runs entirely on your own machine with no API key and does live dictation inside Obsidian. Windows includes voice typing on Win+H and macOS includes system dictation, both free, though both are noticeably weaker than a Whisper-class model on unusual vocabulary, and macOS dictation specifically has a long and unresolved history of trouble inside Obsidian that is covered further down this page. BlabbyAI has a free tier of 60 credits a week with no card required, which is enough to find out whether dictating your notes suits how you think before you decide anything. The Chrome extension needs no desktop installation at all.
It is the difference between talking while you write and talking first then processing a file afterwards, and it is the single most common misunderstanding on this topic. Dictation means your words appear as text at your cursor while you speak, so the note is finished when you stop talking. Transcription means you record audio, stop, hand the file to a model, wait, and then the text arrives. Most of the Obsidian plugins recommended for voice to text are actually transcription tools: you press a key to start recording, press it again to stop, and the plugin then sends the audio to an API and inserts the result. That is a genuinely useful workflow for a long voice memo captured away from the keyboard, and a poor one for the ordinary act of writing a note by talking. If what you want is to think out loud and watch the note fill in, you want dictation, which means either Speech Kit or a system-level tool like BlabbyAI.
Three of the four most commonly recommended ones do. Speech to Text requires a key from OpenAI Whisper or Deepgram, and its own listing quotes Deepgram Nova-3 at $0.0043 a minute. Whisper by nikdanilov works with OpenAI, Groq, Azure or any Whisper-compatible endpoint, and the plugin is MIT licensed while the API usage is billed to you. Voice MD requires an OpenAI key stored in Obsidian SecretStorage and has a 25 MB upload limit per file. Speech Kit is the exception and requires no key, no account and no subscription, because it runs the model locally. The practical consequence of a bring-your-own-key plugin is that you sign up for a developer account, generate a credential, add billing to it and paste the key into a note-taking app, which is a fair amount of setup before the first word is transcribed. A tool like BlabbyAI has no key step at all: install it, press the shortcut, speak.
This is genuinely unsettled, and most pages that answer it confidently are quoting evidence that does not say what they think it says. The famous forum quote, that dictation "works in a lot of places and applications but not obsidian", comes from a January 2021 thread whose author is explicitly describing macOS dictation, not Windows. Microsoft documents Win+H as needing only a cursor in a text box, and publishes no application compatibility list, so their documentation cannot settle the Obsidian case either way. The nearest first-hand Windows comment on the Obsidian forum criticizes Win+H for accuracy and thin language support rather than for failing to run. The genuinely useful answer is that this is a sixty-second test on your own machine, and that a tool which does its own text insertion removes the question. BlabbyAI has a native Windows desktop app that types into any focused text field, so it behaves the same in Obsidian as it does in Word or in a browser.
The reports are real, they are old, and no vendor has ever published a cause. The original January 2021 forum thread describes macOS dictation working across the system but not in Obsidian, and the replies underneath it disagree with each other in a way that is more informative than a flat failure would be: one user reports it working "around 10% of the time in Obsidian and 100% of the time everywhere else", another reports it working fine, and a third reports it working once or twice then stopping. A separate 2020 bug report describes a long delay during which Obsidian appears unresponsive, and it was archived without a documented resolution. The community theory is that Obsidian being an Electron application is the cause, and Obsidian's own developer documentation does confirm the desktop app exposes Electron APIs, but no primary source establishes that Electron is why dictation struggles, so it is worth treating as a plausible guess rather than an explanation. The last substantive post in that thread is from March 2023, so nobody has documented the current state either. What is certain is that intermittent dictation is worse than none, because you cannot tell whether a thought was captured. That unpredictability is the argument for a tool that does its own text insertion rather than relying on the system dictation route: BlabbyAI types into whichever field has focus, so it behaves the same in Obsidian as in any other application.
Yes, and mobile works differently from desktop in a way that explains why it is more reliable. On a phone the dictation is performed by the keyboard, not by Obsidian: you tap the microphone on your keyboard, the operating system transcribes, and the keyboard inserts the text as though you had typed it. Obsidian is not involved in the audio path at all, which is why Android users report Google voice typing working well inside Obsidian Mobile. There are exceptions tied to specific keyboards rather than to Obsidian, with two forum threads reporting a bug that appears when SwiftKey is paired with Google voice-to-text. Desktop has no equivalent mechanism, which is the whole reason this question is complicated on a computer and simple on a phone. On a Windows desktop, BlabbyAI plays the role the phone keyboard plays: it puts the text into the focused field itself.
Raw transcription gives you a wall of text, which is the wrong shape for a note that is meant to be searched and linked later, and cleaning it up by hand undoes the time you saved. This is where the tool choice matters more than the accuracy figures. BlabbyAI custom modes are free-form AI instructions, so a mode can turn a spoken ramble into a set of Markdown bullets, another can produce a short summary followed by action items, and another can leave technical vocabulary strictly alone. You choose the mode before you speak and the text arrives already shaped, so the note is filed rather than merely captured. Custom modes also handle spoken commands you define yourself, so saying "new paragraph" can break the line rather than typing the words. That is a different job from transcription accuracy and it is usually the one that decides whether dictated notes actually get used.
For the prose part of a note, usually yes, and the size of the gap depends on what you are comparing. Comfortable sustained speech runs meaningfully faster than most people type, and the advantage is largest for the loose exploratory writing that a personal knowledge base is mostly made of: a paragraph working out what you think about something, a summary of a conversation, a note to your future self. The advantage shrinks or reverses for anything dense with syntax, so links, tags, nested lists and code fences are still faster with your hands. In practice most people who dictate into a note-taking app end up mixed rather than hands-free: speak the paragraphs, type the structure. We have looked at the underlying studies in more detail in a separate piece on whether dictation is faster than typing, including the one where typing won. The gap also depends on what arrives: raw transcription still needs tidying, whereas BlabbyAI custom modes can return the paragraph already formatted as bullets or a summary, so the note is finished rather than merely captured.
Talk to your notes, and to everything else you write
BlabbyAI types into any focused text field, so one shortcut covers Obsidian, your email, your browser and your chat, with no plugin and no API key. Custom modes turn a spoken thought into bullets, a summary or a task list as it arrives. Windows app and Chrome extension, 60 credits a week free, no card.