Clinical Study

Give ‘Em Something to Talk About

Demystifying Speech BCIs

Speech decoding is the process by which a brain-computer interface (BCI) reads neural activity from speech-related regions of the brain, interprets the words a person is trying to say, and turns them into text or audible speech. For people who have lost the ability to speak, this transformative technology offers a path to communicate by directly using the intention to talk rather than the coordinated movement of the lips, tongue, and larynx. 

But when we think about using BCIs to restore speech for people who have lost it, it almost seems like magic. How is it that a device can read out your own words simply from your attempt to speak? 

The reality is that it isn’t magic, but mechanism: an intricate bond between biology and technology that has resulted in a device with complex capabilities. Here, we’ll dive into the science behind speech decoding to demystify the process.

What is a Speech Brain-Computer Interface (BCI)?

A brain-computer interface is a computer-based system that records brain signals, analyzes them, and then converts them into intended action (Shih et al. 2012)

In the case of a speech BCI, it reads signals from regions of the brain associated with speech, decodes their word intent, and then converts them into the user’s intended speech. 

How Does a Speech BCI Work? Recording, Decoding, and Action

There are three key processes through which a speech BCI translates thoughts into speech in this manner: Recording, Decoding, and Action

To examine this, we will go through one example of a BCI system: brain-to-text-to-speech, where brain signals are translated to text and then converted to speech (Willett et al. 2023; Card et al. 2024). 

Recording involves acquiring neural signals, processing them, and getting them out of the brain and onto a computer. Decoding involves extracting the actual intended words behind neural signals, by passing these signals through various advanced statistical models, including large language models (LLMs). Action involves turning these intended words into a form of communication, such as synthesized text or speech. 

One key thing to keep in mind is how each of these steps contributes to latency, the delay between the user’s attempted speech and the actual text and speech synthesis. 

Speech BCIs should ultimately be seamlessly integrated into everyday life, which means allowing the user to speak at approximately 160 words per minute, the speed of natural conversation (Yuan et al. 2006). Thus, latency is a key consideration in building a BCI that is actually useful in everyday life. 

Stage 1: Recording Neural Signals

During the recording process, the BCI device captures the electrical signals generated by neurons during attempted speech. 

Neurons communicate with each other through action potentials, or electrical spikes across their cell membranes. Some BCIs record single-neuron action potentials directly, while others can only measure the averaged activity of many neurons’ action potentials. 

After recording these signals, the speech BCI processes them by amplifying relevant signals from the neurons of interest and filtering out unwanted components that don’t help the BCI.

After this processing, the data will be passed onto the computerized part of the system. However, computers don’t speak in action potentials; they speak in 1’s and 0’s. 

Hence, these signals are digitized, meaning they are converted into the language of computers (binary language). The signals get further processed to enhance against the noise, and they are then passed onto the decoding process. 

Digitization of action potentials into binary language to be passed onto computerized system.

During recording, delays can be introduced when filtering the digitized signal from the implant. These delays are small (on the order of milliseconds), but still contribute to the total latency that accumulates across the whole brain-to-speech pipeline.

The recording stage also sets a ceiling on the whole system’s information transfer rate (ITR) — a measure, in bits per second (bps), of how much information actually makes it from intention to decoded action. The decoder can only work with the amount of speech-relevant information recorded at this stage, as you cannot reverse engineer this information downstream from what was never recorded

Stage 2: Neural Decoding

Decoding is the stage through which recorded neural signals are translated into intended speech. For brain-to-text-to-speech, this commonly involves something called phonemes, which are the building block sounds of speech –  the individual little sound bites that make up a whole word. 

For example, the word “cat” has 3 phonemes: “/k/”, “/a/”, and “/t/”

“Sun” also has 3 phonemes: “/s/” “/u/” “/n/” 

In the English language, there are ~40 different phonemes which can be combined to make any English word.

For the first step in decoding brain-to-text-to-speech, a machine learning model (such as a neural network) analyzes the data and determines which phonemes are most likely represented from the neural data.

Digitized neural data passed through neural network to predict most likely phenomes.

These different phoneme probabilities are then fed into one or more language models that determine the most likely sentence from these phonemes. 

Predicted phonemes passed to large language model to predict the most likely sentence.

At this point, latency can be introduced depending on the types of language models that are used. Some require a lot of computing power and computing time, but may yield higher accuracy. For example, language models with a higher working memory or context window can lead to higher accuracy, fewer hallucinations, and more coherent responses, but can also take extra processing time. The key to minimizing latency here is to dial in the right model in order to have high accuracy without introducing too much delay. 

Stage 3: Action 

The action stage is where the user’s speech is finally realized. 

After the language models determine the most likely word, phrase, or sentence, this is then passed to a text-to-speech software program. Finally, the intended speech will be transmitted out of a speaker. 

Most likely sentence transcribed into speech and transmitted out of a speaker.

The latency that can be introduced here also depends on the speed of the decoding models and the software; the decoder model must have confidence to generate the most likely sentence on the screen, and the text-to-speech software has to accurately generate that into sound. 

At this stage, if a user’s voice has been recorded earlier through voice-banking software, the speech can actually be spoken in their own voice. Additionally, if this action system is combined with other software, a personalized digital avatar can speak for users as well, in the case that their facial movements are also significantly limited and they’d like to be more expressive in their speech. 

Taken together, the goal across all three stages is to keep total latency low enough that conversation feels natural. 

Just the written transcript of natural speech (excluding other important aspects of communication, like emotion and emphasis) conveys roughly 39 bits of information per second (Coupé et al. 2019) and runs at about 160 words per minute (Yuan et al., 2006), so a speech BCI aims to recognize, decode, and action intended words quickly enough to keep pace with the rhythm of natural conversation rather than a noticeable delay behind it. 

For perspective, the Connexus® BCI measured an information transfer rate of over 200 bits per second of brain state information in preclinical benchmark testing — significantly above the rate at which speech itself conveys information with plenty of headroom, both for accuracy and for advanced applications down the line.

Using Our Inside Voices: Inner Speech Decoding

While speech BCIs are a transformative technology that can drastically improve the quality of life for people in need, we should also address the elephant in the room with this technology: what about inner speech? Can my BCI broadcast private thoughts? 

First, let’s define what inner speech is. Inner speech is often described as the inner monologue or the “little voice inside our heads” (Alderson-Day and Fernyhough 2015). It can be both voluntary and involuntary; for example, if you are counting out a pile of dollar bills, you might voluntarily recite “one, two, three” in your inner monologue. If you were walking and came across a snake, you might involuntarily scream an expletive in your head.  

There are concerns about if speech BCIs decode every word you think all the time, like in the unfortunate circumstance your speech BCI could decode “yuck” while trying grandma’s heirloom fruitcake at Thanksgiving. 

This concern has been addressed in a paper by Kunz and colleagues, where they found that inner speech is discernible from intended, outer speech in neural data (Kunz et al. 2025). They also found that they could block inner speech from being decoded, which could then only be unlocked if a user “thought” the specific passkey for the decoder. 

Being able to switch inner speech decoding on and off is important; you could switch inner speech decoding on in the case that you wanted to passively take mental notes during a boring meeting or keep note of your deep shower thoughts. In a clinical context, decoding inner speech could also lead to less mental and physical fatigue for the users. 

Instances such as these are a successful example of what the BCI research community should strive to do: take feedback from actual BCI users to improve the performance, security, and comfort of those who will actually use these devices.

The Future of Speech Restoration

For the hundreds of thousands of people living with conditions like ALS, brainstem stroke, and spinal cord injury, the loss of speech is also a loss of independence, connection, and the ability to advocate for themselves. 

Restoring even a portion of that — at the speed and nuance of real conversation — would transform the lives of these people in need. It would mean reclaiming a fundamental part of how a person moves through the world. These are the stakes behind every bit per second and every millisecond of latency discussed above.

Several speech BCI systems are currently in clinical trials, including our Connexus® BCI. As the field advances, the work ahead is not only technical but human: building systems that are faster, more accurate, and more powerful in partnership with the people who will actually rely on them. 

Demystifying how speech decoding works is part of that partnership – because the people we hope to help deserve to understand and inform the design of the technology being built for them.

FAQs

Can a BCI Read Your Thoughts?

What speech BCIs record aren’t actually formed “thoughts” per se, but neural data, and it has to be trained to even determine what that data means for each specific user. For example, when training a speech BCI, the device reads neural data, translates it into speech, and then displays this text on a screen for the user to approve as correct or incorrect. 

By design, the user is in control of what gets decoded when; the BCI user interface clearly displays if the decoder is on or off and if it is decoding inner speech or not; both processes are directly controlled by the user. Thus, instead of passively reading thoughts and broadcasting them, BCIs perform a highly specific task that requires specialized training, and it only does this at the user’s directive. 

How Fast Can a Speech BCI Work?

Intracortical BCIs with a direct brain-to-speech decoding pathway have been recorded to synthesize speech at a near-instantaneous speed (Wairagkar et al. 2025), going from acquiring neural signals to producing sounds with a delay of 25 milliseconds. This is below what is perceived as a natural “gap” in conversation (~120 ms) (Hedner 2011). 

Is a Speech BCI Safe?

Intracortical BCIs have been used for many different applications over decades with an excellent safety profile. One of the longest-running intracortical research programs (BrainGate) reported zero intracranial infections, zero device-related deaths, and zero permanent disability events across 33.4 person-years of follow-up (Rubin et al. 2023). Additionally, established procedures analogous to intracortical (such as deep brain stimulation and stereoelectroencephalography) require implantation of devices with a much larger footprint at a greater depth than an intracortical device; these methods also have an excellent safety record over decades of use (Sarica et al., 2023; Muthiah et al., 2025; Rasiah et al., 2023).