Methods and devices for real-time word and speech decoding from neural activity
Patent Information
- Application Number
- JP2023572722
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-05-26
- Filing Date
- 2022-05-26
- Publication Date
- 2025-05-30
AI Technical Summary
Existing methods for restoring communication in individuals with dysarthria, such as those with stroke, traumatic brain injury, or amyotrophic lateral sclerosis, are slow and tedious, failing to provide a natural and efficient means of speech restoration.
A method and system that decodes words and sentences directly from neural activity using a deep learning computational model, incorporating language models to predict word sequences, and utilizes neural recording devices to capture brain electrical signals from the sensorimotor cortex, processing them to restore communication.
Enables real-time decoding of speech and sentences from brain activity, improving autonomy and quality of life for individuals with dysarthria by providing a faster and more direct means of communication.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit under 35 U.S.C. §119(e) of Provisional Application No. 63 / 193,351, filed May 26, 2021, which is incorporated herein by reference in its entirety.
[0002] Government Support Statement This invention was made with government support under Grant No. U01 NS098971-01 awarded by the National Institutes of Health (NIH). The government has certain rights in this invention.
[0003] Introduction Dysarthria is the loss of the ability to speak. Dysarthria can result from a variety of conditions, including stroke, traumatic brain injury, and amyotrophic lateral sclerosis (Beukelman et al. (2007) Augmentative and Alternative Communication 23(3):230-242). For paralyzed individuals with severe motor disabilities, dysarthria can interfere with communication with family, friends, and caregivers, reducing self-reported quality of life (Felgoise et al. (2016) Amyotrophic Lateral Sclerosis and Frontotemporal Degeneration 17(3-4):179-183). Neurotechnologies designed to restore communication to paralyzed patients who have lost the ability to speak have the potential to improve autonomy and quality of life. However, most existing approaches are slow and tedious compared to natural speech. Therefore, there remains a need for better methods to restore communication skills to patients with dysarthria. Summary of the Invention
[0004] Methods, devices, and systems are provided for assisting an individual's communication. Specifically, methods, devices, and systems are provided for decoding words and sentences directly from an individual's neural activity. In the disclosed methods, cortical activity from brain regions involved in speech processing is recorded while the individual attempts to speak or spell words (even though the words or spelled letters are not spoken). A deep learning computational model is used to detect and classify words from the recorded brain activity. The decoding of speech from brain activity is aided by the use of a language model that predicts how a particular word sequence is likely to appear. Additionally, decoding of trial non-speech movements from neural activity can be used to further assist communication. The neurotechnologies described herein can be used to restore communication to patients who have lost the ability to speak, potentially improving autonomy and quality of life.
[0005] In one aspect, a method of assisting a subject in communication is provided, the method including: positioning a neurorecording device comprising electrodes at a location within a sensorimotor cortical region of the subject's brain to record electrical brain signal data associated with trial utterances by the subject; positioning an interface in communication with a computing device at a location on the subject's head, the interface being connected to the neurorecording device; recording the electrical brain signal data associated with the trial utterances by the subject using the neurorecording device, the interface receiving the electrical brain signal data from the neurorecording device and transmitting the electrical brain signal data to a processor; and decoding, using the processor, words, phrases, or sentences from the recorded electrical brain signal data.
[0006] In certain embodiments, the subject has difficulty communicating due to dysarthria, stroke, traumatic brain injury, brain tumor, or amyotrophic lateral sclerosis. In some embodiments, the subject is paralyzed.
[0007] In certain embodiments, the location of the neurorecording device is within the ventral sensorimotor cortex. For example, electrodes may be positioned on the surface of or within the sensorimotor cortical region. In some embodiments, electrodes are positioned on the surface of the sensorimotor cortical region of the brain within the subdural space.
[0008] In certain embodiments, the method includes recording electrical brain signal data from a sensorimotor cortical region selected from the precentral gyrus, postcentral gyrus, posterior middle frontal gyrus, posterior superior frontal gyrus, or posterior inferior frontal gyrus region, or any combination thereof.
[0009] In certain embodiments, the neurorecording device comprises a brain-invasive electrode array or an electrocorticography (ECoG) electrode array.
[0010] In certain embodiments, the electrodes are deep electrodes or surface electrodes.
[0011] In certain embodiments, the features used by the processor are high gamma frequency features contained in the electrical signal data, and in some embodiments, the high gamma frequency electrical signal data may include neural oscillations in the range of 70 Hz to 150 Hz.
[0012] In certain embodiments, the method further includes mapping the subject's brain to identify optimal locations for positioning electrodes to record electrical brain signals associated with trial speech by the subject.
[0013] In certain embodiments, the interface comprises a percutaneous pedestal connector attached to the subject's skull, hi some embodiments, the interface further comprises a detachable headstage connected to the percutaneous pedestal connector.
[0014] In certain embodiments, the processor is provided by a computer or a handheld device (e.g., a mobile phone or tablet).
[0015] In certain embodiments, the processor is programmed to automate speech detection, word classification, and sentence decoding using machine learning algorithms based on identifying neural activity patterns of electrical signals in the recorded electrical brain signal data associated with trial word productions by the subject. In some embodiments, the machine learning algorithms use artificial neural network (ANN) models for speech detection and word classification and natural language processing techniques, such as, but not limited to, hidden Markov models (HMMs) or Viterbi decoding models, for sentence decoding.
[0016] In certain embodiments, the processor is programmed to automate the detection of the onset and end of word production during trial utterances by the subject. In some embodiments, the method further includes assigning speech event labels for preparation, utterance, and pause to time points during the recording of the electrical brain signal data. In some embodiments, the processor is programmed to use the electrical brain signal data recorded within a time window around the detected onset for word classification.
[0017] In certain embodiments, the subject is restricted to a specified set of words for the trial utterance.
[0018] In certain embodiments, the processor is programmed to calculate the probability that a word in the word set is the intended word that the subject attempted to produce during the trial utterance. In some embodiments, the processor is programmed to calculate, for every word in the word set, the probability that the word in the word set is the intended word that the subject attempted to produce during the trial utterance, and select the word in the word set that has the highest probability of being the intended word that the subject attempted to produce during the trial utterance.
[0019] In a particular embodiment, the word set may include am, are, bad, bring, clean, closer, comfortable, coming, computer, do, faith, family, feel, glasses, going, good, goodbye, have, hello, help, here, hope, and how. Examples include: hungry, I, is, it, like, music, my, need, no, not, nurse, okay, outside, please, right, success, tell, that, they, thirsty, tired, up, very, what, where, yes, and you.
[0020] In certain embodiments, the subject may use words from the word set without restriction to create sentences, while in other embodiments the subject is restricted to a specified sentence set in the trial utterance.
[0021] In certain embodiments, the processor is programmed to calculate the probability that a word sequence is an intended sentence that the subject attempted to produce during the trial utterance. In some embodiments, the processor is programmed to calculate, for every sentence in the sentence set, the probability that the sentence in the sentence set is an intended sentence that the subject attempted to produce during the trial utterance. In some embodiments, the processor is programmed to calculate the probability that a number of possible sentences composed entirely of words from the specified word set are intended sentences that the subject attempted to produce during the trial utterance. In some embodiments, the processor is programmed to retain the sentence composed entirely of words from the specified word set that the subject most likely attempted to produce during the trial utterance, as well as other such sentences that are less likely. In some embodiments, the processor is programmed to track the probabilities of the first, second, and third most likely sentences at any given time. The most likely sentence may change as new word events are processed. For example, a second most likely sentence based on processing a word event may become the most likely sentence after one or more additional word events are processed.
[0022] In certain embodiments, the sentence set includes sentences that the subject can select to communicate with the caregiver regarding a task the caregiver would like the caregiver to perform. In some embodiments, the sentences that can be composed entirely of words from the specified word set include sentences that the subject can use to communicate with the caregiver regarding a task the caregiver would like the caregiver to perform.
[0023] In a particular embodiment, the sentence set may include: Are you going outside, Are you tired, Bring my glasses here, Bring my glasses please, Do not feel bad, Do you feel comfortable, Faith is good, Hello how are you, Here is my computer, How do you feel, How do you like my music, I am going outside, I am not going, I am not hungry, I am not okay, I am okay, I am outside, I am thirsty, I do not feel comfortable, I feel very comfortable, I feel I am very hungry, I hope it is clean, I like my nurse, I need my glasses, I need you, It is comfortable, It is good, It is okay, It is right here, My computer is clean, My family is here, My family isExamples include: outside, My family is very comfortable, My glasses are clean, My glasses are comfortable, My nurse is outside, My nurse is right outside, No, Please bring my glasses here, Please clean it, Please tell my family, That is very clean, They are coming here, They are coming outside, They are going outside, They have faith, What do you do, Where is it, Yes, and You are not right.
[0024] In particular embodiments, the processor is programmed to use a language model that provides the probability of a next word given a previous word or phrase in a word sequence to aid decoding by determining predicted word sequence probabilities, for example, according to the language model, more frequently occurring words are assigned a higher weight than less frequently occurring words.
[0025] In certain embodiments, the processor is programmed to use a hidden Markov model (HMM) or a Viterbi decoding model to determine the most likely word sequences in the subject's intended utterance given electrical brain signal data associated with the trial utterance, predicted word probabilities from word classification using a machine learning algorithm, and word sequence probabilities using a language model.
[0026] In certain embodiments, the method further includes recording electrical brain signal data associated with a subject's trial non-speech movement, wherein the subject performs a trial non-speech movement to indicate the start or end of a trial speech utterance or to control an external device, and analyzing the electrical brain signal data using a non-speech movement classification model that identifies a pattern of electrical signals in the recorded electrical brain signal data associated with the trial non-speech movement and calculates a probability that the subject attempted the non-speech movement. In some embodiments, the trial non-speech movement includes a trial movement of the head, arm, hand, foot, or leg.
[0027] In certain embodiments, the processor is further programmed to automate detection of a trial non-speech movement of the subject based on identifying a neural activity pattern of electrical signals within the recorded electrical brain signal data that is associated with the trial non-speech movement, hi some embodiments, the processor is further programmed to assign an event label for the trial non-speech movement to a time point during the recording of the electrical brain signal data.
[0028] In certain embodiments, the method further includes evaluating the accuracy of the decoding.
[0029] In another aspect, there is provided a computer-implemented method for decoding sentences from recorded electrical brain signal data associated with trial utterances by a subject, the computer performing the steps including: a) receiving recorded electrical brain signal data from the subject; b) analyzing the recorded electrical brain signal data using a speech detection model to calculate a probability that trial utterances are occurring at any time during the recording of the electrical brain signal data and to detect the start and end of word productions during the trial utterances by the subject; c) analyzing the electrical brain signal data using a word classification model to identify patterns of electrical signals in the recorded electrical brain signal data associated with the trial word productions by the subject and to calculate predicted word probabilities; d) performing sentence decoding by using the calculated word probabilities from the word classification model in combination with predicted word sequence probabilities within the sentence using a language model that provides the probability of a next word given a previous word or phrase in a word sequence to calculate predicted word sequence probabilities, and determining a most likely word sequence within the sentence based on the predicted word probabilities determined using the word classification model and the language model; and e) displaying the sentence decoded from the recorded electrical brain signal data.
[0030] In certain embodiments, the processor is programmed to automate speech detection, word classification, and sentence decoding using machine learning algorithms based on identifying neural activity patterns of electrical signals in the recorded electrical brain signal data associated with trial word productions by the subject. In some embodiments, the machine learning algorithms use artificial neural network (ANN) models for speech detection and word classification and natural language processing techniques, such as, but not limited to, hidden Markov models (HMMs) or Viterbi decoding models, for sentence decoding.
[0031] In certain embodiments, the subject is restricted to a specified word set for the trial utterance, and in some embodiments, the processor is further programmed to: calculate, for every word in the word set, a probability that the word in the word set is the intended word that the subject was trying to produce during the trial utterance; and select the word in the word set that has the highest probability of being the intended word that the subject was trying to produce during the trial utterance.
[0032] In certain embodiments, the subject may use words from the word set without restriction to create sentences. In other embodiments, the subject is restricted to a specified sentence set in the trial utterance. In some embodiments, the processor is further programmed to calculate a probability that the word sequence is an intended sentence that the subject attempted to produce during the trial utterance. In some embodiments, the processor is further programmed to calculate a probability that a sentence from the sentence set is an intended sentence that the subject attempted to produce during the trial utterance.
[0033] In certain embodiments, the computer-implemented method further includes assigning speech event labels for preparation, speech, and pauses to time points during the recording of the electrical brain signal data.
[0034] In certain embodiments, the computer-implemented method further includes analyzing electrical brain signal data recorded within a time window around the detected onset of the word classification (e.g., from 1 second before the detected onset of the word classification to 3 seconds after the detected onset).
[0035] In certain embodiments, the computer-implemented method further includes assigning a higher weight to more frequently occurring words than to less frequently occurring words according to the language model.
[0036] In certain embodiments, the computer-implemented method further includes receiving recorded electrical brain signal data associated with a subject's attempted non-speech movement, wherein the subject performs a trial non-speech movement to indicate the start or end of a trial speech utterance or to control an external device; and analyzing the electrical brain signal data using a non-speech movement classification model to identify a pattern of electrical signals in the recorded electrical brain signal data associated with the trial non-speech movement and calculate a probability that the subject attempted the non-speech movement. In some embodiments, the trial non-speech movement includes a trial movement of the head, arm, hand, foot, or leg. In some embodiments, the computer-implemented method further includes assigning an event label for the trial non-speech movement to a time point during the recording of the electrical brain signal data.
[0037] In certain embodiments, the computer-implemented method further includes storing a user profile of the subject that includes information regarding patterns of electrical signals in the recorded electrical brain signal data that are associated with trial word productions by the subject.
[0038] In another aspect, a non-transitory computer-readable medium is provided that includes program instructions that, when executed by a processor in a computer, cause the processor to perform the computer-implemented method described herein for decoding sentences from recorded electrical brain signal data associated with trial utterances by a subject.
[0039] In another aspect, a kit is provided comprising a non-transitory computer-readable medium and instructions for decoding electrical brain signal data associated with trial utterances by a subject.
[0040] In another aspect, a system for assisting a subject in communication is provided, the system comprising: a neural recording device comprising electrodes adapted to be positioned at locations within a sensorimotor cortical region of the subject's brain to record electrical brain signal data associated with trial utterances by the subject; a processor programmed to decode sentences from the recorded electrical brain signal data in accordance with the computer-implemented methods described herein; an interface in communication with the computing device adapted to be positioned at locations on the subject's head, the interface receiving the electrical brain signal data from the neural recording device and transmitting the electrical brain signal data to the processor; and a display component for displaying the sentences decoded from the recorded electrical brain signal data.
[0041] In certain embodiments, the subject has difficulty communicating due to dysarthria, stroke, traumatic brain injury, brain tumor, or amyotrophic lateral sclerosis.
[0042] In certain embodiments, the location of the neurorecording device is within the ventral sensorimotor cortex.
[0043] In certain embodiments, the electrodes are adapted to be positioned on the surface of or within a sensorimotor cortical region, hi some embodiments, the electrodes are adapted to be positioned on the surface of a sensorimotor cortical region of the brain within the subdural space.
[0044] In certain embodiments, the neurorecording device comprises a brain-invasive electrode array or an electrocorticography (ECoG) electrode array.
[0045] In certain embodiments, the electrodes are deep electrodes or surface electrodes.
[0046] In certain embodiments, the electrical signal data includes high gamma frequency components characteristic of the electrical signal. In some embodiments, the high gamma frequency electrical signal data includes neural oscillations in the range of 70 Hz to 150 Hz.
[0047] In certain embodiments, the interface comprises a percutaneous pedestal connector attached to the subject's skull, hi some embodiments, the interface further comprises a headstage connectable to the percutaneous pedestal connector.
[0048] In certain embodiments, the processor is provided by a computer or a handheld device (e.g., a mobile phone or tablet).
[0049] In certain embodiments, the processor is programmed to automate speech detection, word classification, and sentence decoding using machine learning algorithms based on identifying neural activity patterns of electrical signals in the recorded electrical brain signal data associated with trial word productions by the subject. In some embodiments, the machine learning algorithms use artificial neural network (ANN) models for speech detection and word classification and natural language processing techniques, such as, but not limited to, hidden Markov models (HMMs) or Viterbi decoding models, for sentence decoding.
[0050] In certain embodiments, the processor is further programmed to assign speech event labels for preparation, speech, and pauses to time points during the recording of the electrical brain signal data, hi some embodiments, the processor is further programmed to use the electrical brain signal data recorded within a time window around the detected onset of a word classification.
[0051] In certain embodiments, the subject is restricted to a specified word set for the trial utterance, and in some embodiments, the processor is further programmed to: calculate, for every word in the word set, a probability that the word in the word set is the intended word that the subject was trying to produce during the trial utterance; and select the word in the word set that has the highest probability of being the intended word that the subject was trying to produce during the trial utterance.
[0052] In a particular embodiment, the word set may include am, are, bad, bring, clean, closer, comfortable, coming, computer, do, faith, family, feel, glasses, going, good, goodbye, have, hello, help, here, hope, and how. Examples include: hungry, I, is, it, like, music, my, need, no, not, nurse, okay, outside, please, right, success, tell, that, they, thirsty, tired, up, very, what, where, yes, and you.
[0053] In certain embodiments, the subject may use words from the word set without restriction to create a sentence. In other embodiments, the subject is restricted to a specified sentence set in the trial utterance. In some embodiments, the processor is further programmed to calculate a probability that a word sequence is an intended sentence that the subject attempted to produce during the trial utterance. In some embodiments, the processor is further programmed to calculate a probability that a sentence from the sentence set is an intended sentence that the subject attempted to produce during the trial utterance. In some embodiments, the sentence set includes sentences that the subject can select to communicate with the caregiver regarding a task the caregiver desires the caregiver to perform.
[0054] In a particular embodiment, the sentence set may include: Are you going outside, Are you tired, Bring my glasses here, Bring my glasses please, Do not feel bad, Do you feel comfortable, Faith is good, Hello how are you, Here is my computer, How do you feel, How do you like my music, I am going outside, I am not going, I am not hungry, I am not okay, I am okay, I am outside, I am thirsty, I do not feel comfortable, I feel very comfortable, I feel I am very hungry, I hope it is clean, I like my nurse, I need my glasses, I need you, It is comfortable, It is good, It is okay, It is right here, My computer is clean, My family is here, My family isExamples include: outside, My family is very comfortable, My glasses are clean, My glasses are comfortable, My nurse is outside, My nurse is right outside, No, Please bring my glasses here, Please clean it, Please tell my family, That is very clean, They are coming here, They are coming outside, They are going outside, They have faith, What do you do, Where is it, Yes, and You are not right.
[0055] In certain embodiments, the processor is further programmed to automate detection of a trial non-speech movement of the subject based on identifying a neural activity pattern of electrical signals within the recorded electrical brain signal data that is associated with the trial non-speech movement, hi some embodiments, the processor is further programmed to assign an event label for the trial non-speech movement to a time point during the recording of the electrical brain signal data.
[0056] In another aspect, a kit is provided that includes a system described herein for assisting a subject in communicating, a non-transitory computer-readable medium, and instructions for using the system to record and decode electrical brain signal data associated with trial utterances by the subject.
[0057] In another aspect, a method of assisting a subject in communication is provided, the method including: positioning a neurorecording device having electrodes at a location within a sensorimotor cortical region of the subject's brain to record electrical brain signal data associated with trial spellings of letters of words of an intended sentence by the subject; positioning an interface in communication with a computing device at a location on the subject's head, the interface being connected to the neurorecording device; recording, using the neurorecording device, the electrical brain signal data associated with the subject's trial spellings, the interface receiving the electrical brain signal data from the neurorecording device and transmitting the electrical brain signal data to a processor of the computing device; and decoding, using the processor, the spelled words of the intended sentence from the recorded electrical brain signal data.
[0058] In certain embodiments, the electrical signal data includes high gamma frequency component features (eg, 70 Hz to 150 Hz) and low frequency component features (eg, 0.3 Hz to 100 Hz).
[0059] In certain embodiments, recording the electrical brain signal data comprises recording the electrical brain signal data from a sensorimotor cortical region selected from the precentral gyrus, postcentral gyrus, posterior middle frontal gyrus, posterior superior frontal gyrus, or posterior inferior frontal gyrus region, or any combination thereof.
[0060] In certain embodiments, the method further includes mapping the subject's brain to identify optimal locations for positioning electrodes to record brain electrical signals associated with the subject's attempted spelling of words.
[0061] In certain embodiments, the processor is programmed to automate the detection of brain activity associated with the spelling attempts, letter classification, word classification, and sentence decoding based on identifying neural activity patterns of electrical signals within the recorded electrical brain signal data associated with the subject's spelling attempts of the word.
[0062] In particular embodiments, the processor is programmed to use machine learning algorithms for speech detection, character classification, word classification, and sentence decoding. In some embodiments, the machine learning algorithms may use natural language processing techniques.
[0063] In certain embodiments, the processor is further programmed to constrain word classifications from character sequences decoded from neural activity associated with attempted spellings of words by the subject to only words within the vocabulary of the language used by the subject.
[0064] In certain embodiments, the processor is programmed to automate the detection of the start and end of letter production during a spelling attempt by the subject.
[0065] In certain embodiments, the processor is further programmed to assign speech event labels for preparation, speech, and pauses to time points during the recording of the electrical brain signal data.
[0066] In certain embodiments, the processor is programmed to use electrical brain signal data recorded within a time window around the detected onset of the subject's attempted spelling of the letter.
[0067] In certain embodiments, the method further includes providing the subject with a series of go cues indicating when the subject should begin a spelling trial of each letter of the word of the intended sentence. In some embodiments, the series of go cues are visually presented on a display. In some embodiments, each go cue is preceded by a countdown to the presentation of the go cue, and a countdown to the next letter to be spelled is visually presented on the display and begins automatically after each go cue. In some embodiments, the series of go cues are provided with set time intervals between each go cue. In some embodiments, the subject can control the set time intervals between each go cue. In some embodiments, the processor is programmed to use electrical brain signal data recorded within a time window following the go cues.
[0068] In certain embodiments, the processor is programmed to calculate a probability that a decoded word sequence from the decoded letter sequence is the intended sentence that the subject attempted to produce during the subject's trial spelling of the letters of the words of the intended sentence.
[0069] In particular embodiments, the processor is programmed to use a language model that provides the probability of a next word given a previous word or phrase in a word sequence to aid decoding by determining predicted word sequence probabilities. In some embodiments, more frequently occurring words are assigned a higher weight than less frequently occurring words according to the language model.
[0070] In certain embodiments, the processor is further programmed to use the sequence of predicted character probabilities to calculate potential sentence candidates and to automatically insert spaces in the character sequences between predicted words in the sentence candidates.
[0071] In certain embodiments, the method further includes recording electrical brain signal data associated with a subject's attempted non-speech movement, wherein the subject performs the trial non-speech movement to indicate the start or end of an attempted spelling of a word of an intended sentence or to control an external device, and analyzing the electrical brain signal data using a classification model that identifies patterns of electrical signals in the recorded electrical brain signal data associated with the attempted non-speech movement and calculates a probability that the subject attempted the non-speech movement.
[0072] In certain embodiments, the attempted non-speech movement comprises an attempted head, arm, hand, foot, or leg movement. In some embodiments, the attempted hand movement comprises an imaginary hand gesture or an imaginary hand grasp.
[0073] In certain embodiments, the processor is programmed to automate detection of a subject's attempted non-speech movement signaling the subject's termination of a spelling attempt based on identifying a neural activity pattern of electrical signals within the recorded electrical brain signal data associated with the attempted non-speech movement, hi some embodiments, the processor is further programmed to assign an event label for the attempted non-speech movement to a time point during the recording of the electrical brain signal data.
[0074] In one aspect, the method further includes recording electrical brain signal data associated with trial utterances by the subject using a neuro-recording device, wherein the interface receives the electrical brain signal data from the neuro-recording device and transmits the electrical brain signal data to a processor of a computing device; and decoding, using the processor, words, phrases, or sentences from the recorded electrical brain signal data associated with the trial utterances by the subject as described herein.
[0075] In certain embodiments, the method further includes evaluating the accuracy of the decoding.
[0076] In another aspect, a computer-implemented method for decoding a sentence from recorded electrical brain signal data associated with a subject's trial spellings of letters of a word of the intended sentence is provided, the computer comprising the steps of: a) receiving recorded electrical brain signal data associated with the subject's trial spellings of letters of a word of the intended sentence; b) analyzing the recorded electrical brain signal data using a speech detection model to calculate a probability that the trial spelling is occurring at any point in time and to detect the start and end of letter productions during the subject's trial spellings; and c) identifying patterns of electrical signals in the recorded electrical brain signal data associated with the subject's trial letter productions and generating a system of predicted letter probabilities. a) analyzing the electrical brain signal data using a character classification model to calculate a sequence of predicted character probabilities; b) calculating potential sentence candidates based on the sequence of predicted character probabilities and automatically inserting spaces in the character sequence between predicted words in the sentence candidate, wherein decoded words in the character sequence are constrained to only be words in the vocabulary of the language used by the subject; c) analyzing the potential sentence candidates using a language model that provides the probability of the next word given the previous word or phrase in the word sequence to calculate a predicted word sequence probability and determining the most likely sequence of words in the sentence; and f) displaying the sentence decoded from the recorded electrical brain signal data.
[0077] In certain embodiments, recorded electrical brain signal data is used only within a time window around the detected onset of the subject's attempted spelling of the letter.
[0078] In certain embodiments, the method further includes displaying a series of go cues to the subject indicating when the subject should begin a spelling trial of each letter of the word of the intended sentence. In some embodiments, each go cue is preceded by a countdown to the presentation of the go cue, and a countdown to the next letter to be spelled begins automatically after each go cue. In some embodiments, the series of go cues are provided with set time intervals between each go cue. In some embodiments, the subject can control the set time intervals between each go cue. In some embodiments, electrical brain signal data recorded within a time window following the go cues is used for character classification.
[0079] In certain embodiments, the computer-implemented method further includes receiving recorded electrical brain signal data associated with a subject's attempted non-speech movement, wherein the subject performs a trial non-speech movement to indicate the start or end of an attempted spelling of a word of an intended sentence or to control an external device; and analyzing the electrical brain signal data using a movement classification model to identify a pattern of electrical signals in the recorded electrical brain signal data associated with the attempted non-speech movement and calculate a probability that the subject attempted the non-speech movement. In some embodiments, the attempted non-speech movement includes an attempted movement of the head, arm, hand, foot, or leg. In some embodiments, the attempted hand movement includes an imaginary hand gesture or an imaginary hand grasp.
[0080] In certain embodiments, machine learning algorithms are used for speech detection and character classification.
[0081] In certain embodiments, the computer-implemented method further includes assigning a higher weight to more frequently occurring words than to less frequently occurring words according to the language model.
[0082] In certain embodiments, the computer-implemented method further includes storing a user profile for the subject that includes information regarding patterns of electrical signals within the recorded electrical brain signal data that are associated with letter productions during spelling attempts by the subject.
[0083] In certain embodiments, the electrical signal data includes high gamma frequency component features (eg, 70 Hz to 150 Hz) and low frequency component features (eg, 0.3 Hz to 100 Hz).
[0084] In certain embodiments, the computer-implemented method further includes evaluating the accuracy of the decoding.
[0085] In certain embodiments, the computer-implemented method further includes decoding sentences from recorded electrical brain signal data associated with trial utterances by the subject, wherein the computer performs the steps including: a) receiving recorded electrical brain signal data associated with trial utterances by the subject; b) analyzing the recorded electrical brain signal data using a speech detection model to calculate a probability that the trial utterance is occurring at any point in time and to detect the start and end of word productions during the trial utterances by the subject; c) analyzing the electrical brain signal data using a word classification model to identify patterns of electrical signals in the recorded electrical brain signal data associated with the trial word productions by the subject and calculate predicted word probabilities; d) performing sentence decoding by using the calculated word probabilities from the word classification model in combination with predicted word sequence probabilities within the sentence using a language model that provides the probability of the next word given a previous word or phrase in the word sequence to calculate predicted word sequence probabilities, and determining the most likely word sequence within the sentence based on the predicted word probabilities determined using the word classification model and the language model; and e) displaying the sentence decoded from the recorded electrical brain signal data. In some embodiments, machine learning algorithms are used for speech detection, word classification, and sentence decoding. In some embodiments, artificial neural network (ANN) models are used for speech detection and word classification, and hidden Markov models (HMMs), Viterbi decoding models, or other natural language processing techniques are used for sentence decoding.
[0086] In another aspect, a non-transitory computer-readable medium is provided, the non-transitory computer-readable medium including program instructions that, when executed by a processor in a computer, cause the processor to perform the computer-implemented methods described herein.
[0087] In another aspect, a kit is provided comprising a non-transitory computer-readable medium and instructions for decoding electrical brain signal data associated with attempted spellings of letters of words of an intended sentence by a subject.
[0088] In another aspect, a system for assisting a subject in communication is provided, the system comprising: a neural recording device comprising electrodes adapted to be positioned at a location within a sensorimotor cortical region of the subject's brain to record electrical brain signal data associated with attempted speech, attempted spelling of letters of words of an intended sentence, or attempted non-speech movements by the subject, or a combination thereof; a processor programmed to decode sentences from the recorded electrical brain signal data in accordance with computer-implemented methods described herein; an interface in communication with the computing device, the interface adapted to be positioned at a location on the subject's head, the interface receiving the electrical brain signal data from the neural recording device and transmitting the electrical brain signal data to the processor; and a display component for displaying the sentences decoded from the recorded electrical brain signal data.
[0089] In certain embodiments, the electrodes are adapted to be located on the surface of or within the sensorimotor cortical area.
[0090] In certain embodiments, the electrodes are adapted to lie on the surface of the sensorimotor cortical region of the brain within the subdural space.
[0091] In certain embodiments, the neural recording device comprises a brain-penetrating electrode array.
[0092] In certain embodiments, the neurorecording device includes an electrocorticography (ECoG) electrode array.
[0093] In certain embodiments, the electrodes are deep electrodes or surface electrodes.
[0094] In certain embodiments, the electrical signal data includes high gamma frequency component features (eg, 70 Hz to 150 Hz) and low frequency component features (eg, 0.3 Hz to 100 Hz).
[0095] In certain embodiments, the interface comprises a percutaneous pedestal connector attached to the subject's skull.
[0096] In certain embodiments, the interface further comprises a headstage connectable to the percutaneous pedestal connector.
[0097] In certain embodiments, the processor is provided by a computer or a handheld device (e.g., a mobile phone or tablet).
[0098] In another aspect, a kit comprising a system described herein, a non-transitory computer-readable medium, and instructions for using the system to record and decode electrical brain signal data associated with speech attempts, word spelling attempts, or non-speech movement attempts by a subject, or a combination thereof.
[0099] Methods for assisting a subject in communicating through decoding neural activity associated with trial speech, trial spelling of words, or trial non-speech movements can be combined. These techniques are complementary. In some cases, decoding trial spellings may allow a larger vocabulary to be used than decoding trial speech. However, decoding trial utterances may be easier and more convenient for the subject because it allows for faster, more direct word decoding, which may be preferable for expressing frequently used words. To assist with decoding, trial non-speech movements can be used to signal that the subject is beginning or finishing spelling out the trial utterance or intended message. [Brief explanation of the drawings]
[0100] [Figure 1]Schematic of the direct speech BCI. Neural activity acquired from an investigational electrocorticography (ECoG) electrode array implanted in a clinical trial participant with severe paralysis is used to directly decode words and sentences in real time. In a conversational demonstration, participants are visually prompted with a question (A) and instructed to attempt a response using words from a predefined 50-word vocabulary. Simultaneously, cortical signals are acquired from the brain's surface via the ECoG device (B) and processed in real time (C). A speech detection model analyzes the processed neural signals sample-by-sample to detect the participant's speech attempts (D). A classifier calculates word probabilities (across 50 possible words) from each detected window of associated neural activity (E). A Viterbi decoding algorithm uses these probabilities, along with word sequence probabilities from a separately trained language model, to decode the most likely sentence given the ECoG data (F). A predicted sentence, updated each time a word is decoded, is displayed as feedback to the participant (G). [Figure 2A]Neural signal processing and language modeling enable real-time decoding of various sentences. Figure 2A shows the word error rate for word sequences decoded from participants' cortical activity during sentence task blocks. The word error rate quantifies the frequency with which decoding errors were made (lower word error rates indicate better performance). When decoding words with and without a language model (LM), word error rates were significantly lower than chance, and performance significantly improved when an LM was used during decoding (*all P<0.001, 3-way Holm-Bonferroni correction). Figure 2B shows the decoded words per minute values across all trials, either including or excluding incorrectly decoded words. Each distribution was created using kernel density estimation with Scott bandwidth estimation, with a thick horizontal line indicating the median and a smaller horizontal line indicating the range (excluding outliers more than four standard deviations below or above the mean). Figure 2C shows a summary of the difference between the number of words detected and the actual number of words in each trial. The percentage of trials with the correct sentence length is shown in black, and the incorrect sentence length is shown in dark red. Figure 2D shows the edit distance (number of decoding errors made) of decoded sentences with and without LM across all trials and all 50 sentence targets, sorted by increasing edit distance for LM prediction (lower edit distance indicates better performance). Each small vertical dash represents the edit distance of one trial (there are three trials per target sentence; marks with the same edit distance are horizontally offset for visualization). Each point represents the average edit distance for that target sentence. The histogram at the bottom shows the edit distance count across all trials. Figure 2E shows the target sentences for seven different trials and the decoded sentences with and without LM. Correctly decoded words are shown in black, and incorrect words are shown in red. [Figure 2B]Neural signal processing and language modeling enable real-time decoding of various sentences. Figure 2A shows the word error rate for word sequences decoded from participants' cortical activity during sentence task blocks. The word error rate quantifies the frequency with which decoding errors were made (lower word error rates indicate better performance). When decoding words with and without a language model (LM), word error rates were significantly lower than chance, and performance significantly improved when an LM was used during decoding (*all P<0.001, 3-way Holm-Bonferroni correction). Figure 2B shows the decoded words per minute values across all trials, either including or excluding incorrectly decoded words. Each distribution was created using kernel density estimation with Scott bandwidth estimation, with a thick horizontal line indicating the median and a smaller horizontal line indicating the range (excluding outliers more than four standard deviations below or above the mean). Figure 2C shows a summary of the difference between the number of words detected and the actual number of words in each trial. The percentage of trials with the correct sentence length is shown in black, and the incorrect sentence length is shown in dark red. Figure 2D shows the edit distance (number of decoding errors made) of decoded sentences with and without LM across all trials and all 50 sentence targets, sorted by increasing edit distance for LM prediction (lower edit distance indicates better performance). Each small vertical dash represents the edit distance of one trial (there are three trials per target sentence; marks with the same edit distance are horizontally offset for visualization). Each point represents the average edit distance for that target sentence. The histogram at the bottom shows the edit distance count across all trials. Figure 2E shows the target sentences for seven different trials and the decoded sentences with and without LM. Correctly decoded words are shown in black, and incorrect words are shown in red. [Figure 2C]Neural signal processing and language modeling enable real-time decoding of various sentences. Figure 2A shows the word error rate for word sequences decoded from participants' cortical activity during sentence task blocks. The word error rate quantifies the frequency with which decoding errors were made (lower word error rates indicate better performance). When decoding words with and without a language model (LM), word error rates were significantly lower than chance, and performance significantly improved when an LM was used during decoding (*all P<0.001, 3-way Holm-Bonferroni correction). Figure 2B shows the decoded words per minute values across all trials, either including or excluding incorrectly decoded words. Each distribution was created using kernel density estimation with Scott bandwidth estimation, with a thick horizontal line indicating the median and a smaller horizontal line indicating the range (excluding outliers more than four standard deviations below or above the mean). Figure 2C shows a summary of the difference between the number of words detected and the actual number of words in each trial. The percentage of trials with the correct sentence length is shown in black, and the incorrect sentence length is shown in dark red. Figure 2D shows the edit distance (number of decoding errors made) of decoded sentences with and without LM across all trials and all 50 sentence targets, sorted by increasing edit distance for LM prediction (lower edit distance indicates better performance). Each small vertical dash represents the edit distance of one trial (there are three trials per target sentence; marks with the same edit distance are horizontally offset for visualization). Each point represents the average edit distance for that target sentence. The histogram at the bottom shows the edit distance count across all trials. Figure 2E shows the target sentences for seven different trials and the decoded sentences with and without LM. Correctly decoded words are shown in black, and incorrect words are shown in red. [Figure 2D]Neural signal processing and language modeling enable real-time decoding of various sentences. Figure 2A shows the word error rate for word sequences decoded from participants' cortical activity during sentence task blocks. The word error rate quantifies the frequency with which decoding errors were made (lower word error rates indicate better performance). When decoding words with and without a language model (LM), word error rates were significantly lower than chance, and performance significantly improved when an LM was used during decoding (*all P<0.001, 3-way Holm-Bonferroni correction). Figure 2B shows the decoded words per minute values across all trials, either including or excluding incorrectly decoded words. Each distribution was created using kernel density estimation with Scott bandwidth estimation, with a thick horizontal line indicating the median and a smaller horizontal line indicating the range (excluding outliers more than four standard deviations below or above the mean). Figure 2C shows a summary of the difference between the number of words detected and the actual number of words in each trial. The percentage of trials with the correct sentence length is shown in black, and the incorrect sentence length is shown in dark red. Figure 2D shows the edit distance (number of decoding errors made) of decoded sentences with and without LM across all trials and all 50 sentence targets, sorted by increasing edit distance for LM prediction (lower edit distance indicates better performance). Each small vertical dash represents the edit distance of one trial (there are three trials per target sentence; marks with the same edit distance are horizontally offset for visualization). Each point represents the average edit distance for that target sentence. The histogram at the bottom shows the edit distance count across all trials. Figure 2E shows the target sentences for seven different trials and the decoded sentences with and without LM. Correctly decoded words are shown in black, and incorrect words are shown in red. [Figure 2E]Neural signal processing and language modeling enable real-time decoding of various sentences. Figure 2A shows the word error rate for word sequences decoded from participants' cortical activity during sentence task blocks. The word error rate quantifies the frequency with which decoding errors were made (lower word error rates indicate better performance). When decoding words with and without a language model (LM), word error rates were significantly lower than chance, and performance significantly improved when an LM was used during decoding (*all P<0.001, 3-way Holm-Bonferroni correction). Figure 2B shows the decoded words per minute values across all trials, either including or excluding incorrectly decoded words. Each distribution was created using kernel density estimation with Scott bandwidth estimation, with a thick horizontal line indicating the median and a smaller horizontal line indicating the range (excluding outliers more than four standard deviations below or above the mean). Figure 2C shows a summary of the difference between the number of words detected and the actual number of words in each trial. The percentage of trials with the correct sentence length is shown in black, and the incorrect sentence length is shown in dark red. Figure 2D shows the edit distance (number of decoding errors made) of decoded sentences with and without LM across all trials and all 50 sentence targets, sorted by increasing edit distance for LM prediction (lower edit distance indicates better performance). Each small vertical dash represents the edit distance of one trial (there are three trials per target sentence; marks with the same edit distance are horizontally offset for visualization). Each point represents the average edit distance for that target sentence. The histogram at the bottom shows the edit distance count across all trials. Figure 2E shows the target sentences for seven different trials and the decoded sentences with and without LM. Correctly decoded words are shown in black, and incorrect words are shown in red. [Figure 3A]Different neural activity patterns underlie word generation attempts. Figure 3A shows the effect of the amount of training data on word classification accuracy using cortical activity recorded during participants' isolated word generation attempts. Each point represents the mean ± standard deviation across 10 cross-validation folds. Chance accuracy is depicted as a horizontal dashed line. Figure 3B shows a participant's brain reconstruction overlaid with the location of implanted electrodes and their contributions to the speech detection and word classification models. The size (area) and opacity of the plotted electrodes are scaled by their relative contribution (significant electrodes appear larger and more opaque than others). Each set of contributions is normalized to sum to 1. For anatomical reference, the precentral gyrus is highlighted in light blue. Figure 3C shows word confusions from the classification results, showing how frequently the classifier predicted each of the 50 words the participant was attempting to produce given the identity of the target word (values along the diagonal correspond to correct classifications). [Figure 3B] Different neural activity patterns underlie word generation attempts. Figure 3A shows the effect of the amount of training data on word classification accuracy using cortical activity recorded during participants' isolated word generation attempts. Each point represents the mean ± standard deviation across 10 cross-validation folds. Chance accuracy is depicted as a horizontal dashed line. Figure 3B shows a participant's brain reconstruction overlaid with the location of implanted electrodes and their contributions to the speech detection and word classification models. The size (area) and opacity of the plotted electrodes are scaled by their relative contribution (significant electrodes appear larger and more opaque than others). Each set of contributions is normalized to sum to 1. For anatomical reference, the precentral gyrus is highlighted in light blue. Figure 3C shows word confusions from the classification results, showing how frequently the classifier predicted each of the 50 words the participant was attempting to produce given the identity of the target word (values along the diagonal correspond to correct classifications). [Figure 3C]Different neural activity patterns underlie word generation attempts. Figure 3A shows the effect of the amount of training data on word classification accuracy using cortical activity recorded during participants' isolated word generation attempts. Each point represents the mean ± standard deviation across 10 cross-validation folds. Chance accuracy is depicted as a horizontal dashed line. Figure 3B shows a participant's brain reconstruction overlaid with the location of implanted electrodes and their contributions to the speech detection and word classification models. The size (area) and opacity of the plotted electrodes are scaled by their relative contribution (significant electrodes appear larger and more opaque than others). Each set of contributions is normalized to sum to 1. For anatomical reference, the precentral gyrus is highlighted in light blue. Figure 3C shows word confusions from the classification results, showing how frequently the classifier predicted each of the 50 words the participant was attempting to produce given the identity of the target word (values along the diagonal correspond to correct classifications). [Figure 4A]Neural activity recorded during speech trials exhibits long-term stability. Figure 4A shows neural activity from a single electrode across all trials of a participant uttering the word "Goodbye" during an isolated-word task, spanning 18 months of recording. Figure 4B shows word classification results from training and testing detectors and classifiers on subsets of isolated-word data sampled from four non-overlapping date ranges. Each subset contains data from 20 attempted productions of each word. Each solid bar shows results from cross-validated evaluation within a single subset, while each dotted bar shows results from training on data from all subsets except the one being evaluated. Each bar represents the mean ± standard error across 10 evaluation folds. Chance accuracy is depicted as a horizontal dashed line. Also shown are significant differences between the four same-subset assessments for each test subset (*P<0.01, two-tailed Fisher exact test, 10-way Holm-Bonferroni correction) and between two assessments (*P<0.01, two-tailed exact McNemar test, 10-way Holm-Bonferroni correction). Electrode contributions calculated during cross-validated assessments within a single subset are shown at the top (with the most dorsal and posterior electrodes oriented in the upper right corner). The size (area) and opacity of the plotted electrodes are scaled by their relative contributions. Each set of contributions is normalized to sum to 1. [Figure 4B]Neural activity recorded during speech trials exhibits long-term stability. Figure 4A shows neural activity from a single electrode across all trials of a participant uttering the word "Goodbye" during an isolated-word task, spanning 18 months of recording. Figure 4B shows word classification results from training and testing detectors and classifiers on subsets of isolated-word data sampled from four non-overlapping date ranges. Each subset contains data from 20 attempted productions of each word. Each solid bar shows results from cross-validated evaluation within a single subset, while each dotted bar shows results from training on data from all subsets except the one being evaluated. Each bar represents the mean ± standard error across 10 evaluation folds. Chance accuracy is depicted as a horizontal dashed line. Also shown are significant differences between the four same-subset assessments for each test subset (*P<0.01, two-tailed Fisher exact test, 10-way Holm-Bonferroni correction) and between two assessments (*P<0.01, two-tailed exact McNemar test, 10-way Holm-Bonferroni correction). Electrode contributions calculated during cross-validated assessments within a single subset are shown at the top (with the most dorsal and posterior electrodes oriented in the upper right corner). The size (area) and opacity of the plotted electrodes are scaled by their relative contributions. Each set of contributions is normalized to sum to 1. [Figure 5A] Participant MRI results. Figure 5A shows a sagittal MRI of a participant with brainstem atrophy (labeled in blue) caused by encephalomalacia and pontine stroke (labeled in red). Figure 5B shows two additional MRI scans demonstrating the absence of cerebral atrophy, suggesting that cortical neuronal populations (including those recorded in this study) should be relatively unaffected by the participant's pathology. [Figure 5B]Participant MRI results. Figure 5A shows a sagittal MRI of a participant with brainstem atrophy (labeled in blue) caused by encephalomalacia and pontine stroke (labeled in red). Figure 5B shows two additional MRI scans demonstrating the absence of cerebral atrophy, suggesting that cortical neuronal populations (including those recorded in this study) should be relatively unaffected by the participant's pathology. [Figure 6] Real-time neural data acquisition hardware infrastructure. Electrocorticography (ECoG) data acquired from the implanted array and transcutaneous pedestal connector are processed and sent to a Neuroport digital signal processor (DSP). Simultaneously, microphone data is acquired, amplified, and sent to the DSP. Signals from the DSP are sent to a real-time computer, which controls the tasks presented to the participant, including any decoded sentences provided in real time as feedback. Speaker output from the real-time computer is also sent to the DSP and synchronized with the neural signals (not shown). During the previous session, a human patient cable connected to the pedestal acquired ECoG signals, which were then processed by a front-end amplifier before being sent to the DSP (the human patient cable and front-end amplifier are not shown here but would have been replaced by a digital headstage and digital hub in this pipeline when in use). [Figure 7]Real-time neural signal processing pipeline. Using a data acquisition headstage and rig, participants' electrocorticography (ECoG) signals were acquired at 30 kHz, filtered by a wideband filter, adjusted by a software-based line noise cancellation technique, low-pass filtered at 500 Hz, and streamed at 1 kHz to a real-time computer. On the real-time computer, custom software was used to perform common-average referencing, multiband high-gamma bandpass filtering, analytical amplitude estimation, multiband averaging, and z-scoring on the ECoG signals. The resulting signals were then used as a measure of high-gamma activity for the remaining analyses. [Figure 8] Data collection timeline. If multiple data types were collected on a single day, the bars are stacked vertically (the height of the stacked bars for any given day is equal to the total number of tests collected on that day). Irregularities in the data collection schedule are caused in part by external and clinical time constraints unrelated to the implanted device. The gaps between 55 and 88 weeks are due to clinical guidelines regarding the COVID-19 pandemic. [Figure 9] Schematic of the speech detection model. Z-scored high gamma activity across all electrodes is processed at each time point by an artificial neural network consisting of a stack of three long-short-term memory (LSTM) layers and a single dense (fully connected) layer. The dense layer projects the latent dimensions of the final LSTM layer onto a probability space of three event classes: utterance, preparation, and pause. The predicted speech event probability time series is smoothed and then thresholded with probability and time thresholds to yield the start time (t*) and end time of the detected speech event. During sentence decoding, each time a speech event was detected, a window of neural activity spanning -1 to 3 seconds relative to the detected onset (t*) was passed to the word classifier. The neural activity, predicted speech probability time series (top right), and detected speech event (bottom right) shown are actual neural data and detection results over a 7-second time window of an isolated word trial in which participants attempted to produce the word "family." [Figure 10]Word classification model schematic. For each classification, a 4-second time window of high gamma activity is processed by an ensemble of 10 artificial neural network (ANN) models. Within each ANN, high gamma activity is processed by a temporal convolution followed by two bidirectionally gated recurrent unit (GRU) layers. A dense layer projects the latent dimension from the final GRU layer into probability space, which contains the probability that each word from a 50-word set is the target word during the speech production trial associated with the neural time window. The 10 probability distributions from the ensembled ANN models are averaged together to obtain a final vector of predicted word probabilities. [Figure 11] Sentence Decoding Hidden Markov Model. This Hidden Markov Model (HMM) describes the relationship between the word a participant attempts to produce (hidden state qi) and the associated time window of detected neural activity (observed state yi). The HMM emission probability p(y0|q0) can be simplified to p(ωi|yi) (the word likelihood provided by the word classifier), and the HMM transition probability p(qi|qi-1) can be simplified to p(ωi|ci) (the word sequence prior provided by the language model). [Figure 12A]Auxiliary modeling results using isolated word data. Figure 12A shows the effect of the amount of training data on word classification accuracy (left) and cross-entropy loss (right) using cortical activity recorded during participants' isolated word generation trials. Lower cross-entropy indicates better performance. Each point represents the mean ± standard deviation across 10 cross-validation folds (error bars in the cross-entropy plots were typically too small to be seen alongside the circular markers). Chance performance is shown as a horizontal dashed line in each plot (chance cross-entropy loss is calculated as the negative logarithm (base 2) of the reciprocal of the number of word targets). Performance improved more rapidly during the first 4 hours of training data, and less rapidly during the following 5 hours, but did not plateau. When using all available isolated word data, the information transfer rate was 25.1 bits / min (not shown). Figure 12B shows the effect of the amount of training data on the frequency of detection errors during speech detection and detected event curation using isolated word data. Lower error rates indicate better performance. False positives are detected events not associated with word-generation trials, and false negatives are word-generation trials not associated with detected events. Each point represents the mean ± standard deviation across 10 cross-validation folds. While not all available training data was used to fit each speech detection model, each model consistently used 47–83 minutes of data (not shown). Figure 12C shows the distribution of detected onsets from neural activity across 9,000 isolated-word trials in response to a go cue (100 ms histogram bin size). This histogram was created using the results of the final analysis set of a learning curve scheme (all available trials were included in the cross-validation evaluation). The distribution of detected speech onsets had a mean of 308 ms after the associated go cue and a standard deviation of 1,017 ms. This distribution was likely influenced to some extent by behavioral variability in participants' response times.During detected event curation, 429 studies required curation to select a detected event from multiple candidates (420 studies had two candidates, and 9 studies had three candidates). [Figure 12B]Auxiliary modeling results using isolated word data. Figure 12A shows the effect of the amount of training data on word classification accuracy (left) and cross-entropy loss (right) using cortical activity recorded during participants' isolated word generation trials. Lower cross-entropy indicates better performance. Each point represents the mean ± standard deviation across 10 cross-validation folds (error bars in the cross-entropy plots were typically too small to be seen alongside the circular markers). Chance performance is shown as a horizontal dashed line in each plot (chance cross-entropy loss is calculated as the negative logarithm (base 2) of the reciprocal of the number of word targets). Performance improved more rapidly during the first 4 hours of training data, and less rapidly during the following 5 hours, but did not plateau. When using all available isolated word data, the information transfer rate was 25.1 bits / min (not shown). Figure 12B shows the effect of the amount of training data on the frequency of detection errors during speech detection and detected event curation using isolated word data. Lower error rates indicate better performance. False positives are detected events not associated with word-generation trials, and false negatives are word-generation trials not associated with detected events. Each point represents the mean ± standard deviation across 10 cross-validation folds. While not all available training data was used to fit each speech detection model, each model consistently used 47–83 minutes of data (not shown). Figure 12C shows the distribution of detected onsets from neural activity across 9,000 isolated-word trials in response to a go cue (100 ms histogram bin size). This histogram was created using the results of the final analysis set of a learning curve scheme (all available trials were included in the cross-validation evaluation). The distribution of detected speech onsets had a mean of 308 ms after the associated go cue and a standard deviation of 1,017 ms. This distribution was likely influenced to some extent by behavioral variability in participants' response times.During detected event curation, 429 studies required curation to select a detected event from multiple candidates (420 studies had two candidates, and 9 studies had three candidates). [Figure 12C]Auxiliary modeling results using isolated word data. Figure 12A shows the effect of the amount of training data on word classification accuracy (left) and cross-entropy loss (right) using cortical activity recorded during participants' isolated word generation trials. Lower cross-entropy indicates better performance. Each point represents the mean ± standard deviation across 10 cross-validation folds (error bars in the cross-entropy plots were typically too small to be seen alongside the circular markers). Chance performance is shown as a horizontal dashed line in each plot (chance cross-entropy loss is calculated as the negative logarithm (base 2) of the reciprocal of the number of word targets). Performance improved more rapidly during the first 4 hours of training data, and less rapidly during the following 5 hours, but did not plateau. When using all available isolated word data, the information transfer rate was 25.1 bits / min (not shown). Figure 12B shows the effect of the amount of training data on the frequency of detection errors during speech detection and detected event curation using isolated word data. Lower error rates indicate better performance. False positives are detected events not associated with word-generation trials, and false negatives are word-generation trials not associated with detected events. Each point represents the mean ± standard deviation across 10 cross-validation folds. While not all available training data was used to fit each speech detection model, each model consistently used 47–83 minutes of data (not shown). Figure 12C shows the distribution of detected onsets from neural activity across 9,000 isolated-word trials in response to a go cue (100 ms histogram bin size). This histogram was created using the results of the final analysis set of a learning curve scheme (all available trials were included in the cross-validation evaluation). The distribution of detected speech onsets had a mean of 308 ms after the associated go cue and a standard deviation of 1,017 ms. This distribution was likely influenced to some extent by behavioral variability in participants' response times.During detected event curation, 429 studies required curation to select a detected event from multiple candidates (420 studies had two candidates, and 9 studies had three candidates). [Figure 13] Acoustic Contamination Study. Each blue curve shows the average correlation between a spectrogram from a single electrode and the corresponding spectrogram from a time-aligned microphone signal as a function of frequency. The red curve shows the average power spectral density (PSD) of the microphone signal. The vertical dashed lines mark the 60 Hz line noise frequency and its harmonics. Highlighted in green is the high-gamma frequency band (70–150 Hz), which was the frequency band from which neural features used during decoding were extracted. Across all frequencies, the correlation between the electrode and microphone signal is small. There is a slight increase in correlation at the lower end of the high-gamma frequency range, but this increase in correlation occurs as the microphone PSD decreases. Because the correlation is low and does not increase or decrease with the microphone PSD, the observed correlation is likely due to factors other than acoustic contamination, such as shared electrical noise. After comparing these results with those observed in a study describing acoustic contamination (which informed the contamination analysis used here)
[39] , we concluded that our decoding performance was not artificially improved by acoustic contamination of our electrophysiological recordings. [Figure 14A]Long-term stability of speech-evoked signals. Figure 14A shows neural activity from a single electrode across all trials of a participant uttering the word "Goodbye" during an isolated-word task, spanning 81 weeks of recording. Figure 14B shows the participant's brain reconstruction overlaid with electrode locations. The electrodes shown in panel A are filled in black. For anatomical reference, the precentral gyrus is highlighted in light blue. Figure 14C shows word classification results from training and testing the detector and classifier on subsets of isolated-word data sampled from four non-overlapping date ranges. Each subset contains data from 20 attempted productions of each word. Each solid bar shows results from cross-validated evaluation within a single subset, while each dotted bar shows results from training on data from all subsets except the one being evaluated. Each error bar shows a 95% confidence interval of the mean, calculated across the cross-validation fold. Chance accuracy is depicted as a horizontal dashed line. Electrode contributions calculated during cross-validated assessments within a single subset are shown at the top (with the most dorsal and posterior electrodes oriented in the upper right corner). The size (area) and opacity of the plotted electrodes are scaled by their relative contributions. Each contribution set is normalized to sum to 1. These results suggest that speech-evoked cortical responses remained relatively stable throughout the study period, although model recalibration every 2–3 months may still be beneficial for decoding performance. [Figure 14B]Long-term stability of speech-evoked signals. Figure 14A shows neural activity from a single electrode across all trials of a participant uttering the word "Goodbye" during an isolated-word task, spanning 81 weeks of recording. Figure 14B shows the participant's brain reconstruction overlaid with electrode locations. The electrodes shown in panel A are filled in black. For anatomical reference, the precentral gyrus is highlighted in light blue. Figure 14C shows word classification results from training and testing the detector and classifier on subsets of isolated-word data sampled from four non-overlapping date ranges. Each subset contains data from 20 attempted productions of each word. Each solid bar shows results from cross-validated evaluation within a single subset, while each dotted bar shows results from training on data from all subsets except the one being evaluated. Each error bar shows a 95% confidence interval of the mean, calculated across the cross-validation fold. Chance accuracy is depicted as a horizontal dashed line. Electrode contributions calculated during cross-validated assessments within a single subset are shown at the top (with the most dorsal and posterior electrodes oriented in the upper right corner). The size (area) and opacity of the plotted electrodes are scaled by their relative contributions. Each contribution set is normalized to sum to 1. These results suggest that speech-evoked cortical responses remained relatively stable throughout the study period, although model recalibration every 2–3 months may still be beneficial for decoding performance. [Figure 14C]Long-term stability of speech-evoked signals. Figure 14A shows neural activity from a single electrode across all trials of a participant uttering the word "Goodbye" during an isolated-word task, spanning 81 weeks of recording. Figure 14B shows the participant's brain reconstruction overlaid with electrode locations. The electrodes shown in panel A are filled in black. For anatomical reference, the precentral gyrus is highlighted in light blue. Figure 14C shows word classification results from training and testing the detector and classifier on subsets of isolated-word data sampled from four non-overlapping date ranges. Each subset contains data from 20 attempted productions of each word. Each solid bar shows results from cross-validated evaluation within a single subset, while each dotted bar shows results from training on data from all subsets except the one being evaluated. Each error bar shows a 95% confidence interval of the mean, calculated across the cross-validation fold. Chance accuracy is depicted as a horizontal dashed line. Electrode contributions calculated during cross-validated assessments within a single subset are shown at the top (with the most dorsal and posterior electrodes oriented in the upper right corner). The size (area) and opacity of the plotted electrodes are scaled by their relative contributions. Each contribution set is normalized to sum to 1. These results suggest that speech-evoked cortical responses remained relatively stable throughout the study period, although model recalibration every 2–3 months may still be beneficial for decoding performance. [Figure 15]Schematic of the spelling pipeline. A. At the beginning of the sentence spelling trial, participants silently attempt to speak a word to spontaneously activate the speller. B. Neural features (high gamma activity and low-frequency signals) are extracted in real time from cortical data recorded throughout the task. Features from a single electrode (electrode 0 shown in Figure 19A) are shown. For visualization, the trace was smoothed via convolution with a Gaussian kernel with a standard deviation of 150 ms. The microphone signal shows the absence of speech output during the task. C. A speech detection model, consisting of a recurrent neural network (RNN) and thresholding operations, processes the neural features sample-by-sample to detect silent speech attempts. Once an attempt is detected, the detection model is deactivated and the spelling procedure begins. D. During the spelling procedure, participants spell out the intended message through a complete letter-decoding cycle, which occurs every 2.5 seconds. During each cycle, participants are visually presented with a countdown, followed by a go cue. At the go cue, participants attempt to silently utter a codeword representing the desired letter. E. High-gamma activity and low-frequency signals are calculated for all electrode channels throughout the spelling procedure and divided into 2.5-second non-overlapping time windows corresponding to letter-decoding cycles. F. An RNN-based letter classification model processes each of these neural time windows to predict the probability that the participant was attempting to silently utter each of the 26 possible codewords or to perform a manual command (see G). If the classifier predicts that the participant performed a manual command at least 80% of the time, the spelling procedure ends and the sentence is finalized (see I). Otherwise, the predicted letter probabilities are processed in real time by a beam search algorithm, and the most likely sentence is displayed to the participant. G. After the participant spells out the intended message, they attempt to complete the spelling procedure by clenching their right hand during the next letter-decoding cycle and complete the sentence. H. The neural time windows associated with the manual commands are passed to the classification model.I. If the classifier confirms that the participant attempted a manual command, a neural network-based language model ("DistilGPT-2") rescores sentences consisting of only whole words, and the system uses the most likely sentence after rescoring as the final prediction. [Figure 16A]Summary of spelling system performance during a copy-typing task. Figure 16A shows the character error rate (CER) observed during real-time sentence spelling (labeled "+LM (Real-time results)") and offline simulation in which portions of the spelling system were omitted. In the "Chance" condition, sentences were created by replacing the output from the neural classifier with randomly generated character probabilities without modifying the rest of the spelling pipeline. In the "Only neural decoding" condition, sentences were created solely by concatenating the most likely characters from each of the classifier's predictions during sentence testing (without any whitespace). In the "+Vocab. constraints" condition, predicted character probabilities from the neural classifier were used in conjunction with a beam search that constrained predicted character sequences to form words from a 1,152-word vocabulary. The final condition, labeled "+LM (Real-time results)," shows real-time results during testing by participants and incorporates language modeling during beam search and after sentences were finalized. Sentences decoded in real time by the perfect system exhibited lower CERs than sentences decoded in the other conditions (***P<0.0001, two-tailed Wilcoxon rank-sum test with a six-way Holm-Bonferroni correction). Figure 16B. Word error rate (WER) for the real-time results from Figure 16A and the corresponding offline omission simulation. Figure 16C. Number of characters decoded per minute during real-time testing. Figure 16D. Number of words decoded per minute during real-time testing. In Figures 16A-16D, the distributions are shown as boxplots calculated over n = 34 real-time blocks (in each block, participants attempted spelling 2-5 sentences). Each boxplot shows the quartile of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. In Figures 16A and 16B, each boxplot corresponds to n=34 blocks (in each of these blocks, participants attempted spelling 2-5 sentences).In Figure 16C, each boxplot corresponds to n = 9 blocks (in each of these blocks, participants attempted to spell 2–4 conversational responses). Figure 16E. Number of excess letters in each decoded sentence. Decoded sentences with 0 excess letters indicate that a manual command (to disengage from the speller) was successfully identified from the participant's neural activity immediately after they spelled the last letter of the sentence. Figure 16F. Example sentence spelling trials using decoded sentences from each non-contingency condition. Incorrect letters are colored red. 1 and 2 mark trials in which the real-time decoded sentence contained at least one error. The target sentences for these two trials are given at the bottom of the panel. All other example sentences contained no real-time decoding errors. [Figure 16B]Summary of spelling system performance during a copy-typing task. Figure 16A shows the character error rate (CER) observed during real-time sentence spelling (labeled "+LM (Real-time results)") and offline simulation in which portions of the spelling system were omitted. In the "Chance" condition, sentences were created by replacing the output from the neural classifier with randomly generated character probabilities without modifying the rest of the spelling pipeline. In the "Only neural decoding" condition, sentences were created solely by concatenating the most likely characters from each of the classifier's predictions during sentence testing (without any whitespace). In the "+Vocab. constraints" condition, predicted character probabilities from the neural classifier were used in conjunction with a beam search that constrained predicted character sequences to form words from a 1,152-word vocabulary. The final condition, labeled "+LM (Real-time results)," shows real-time results during testing by participants and incorporates language modeling during beam search and after sentences were finalized. Sentences decoded in real time by the perfect system exhibited lower CERs than sentences decoded in the other conditions (***P<0.0001, two-tailed Wilcoxon rank-sum test with a six-way Holm-Bonferroni correction). Figure 16B. Word error rate (WER) for the real-time results from Figure 16A and the corresponding offline omission simulation. Figure 16C. Number of characters decoded per minute during real-time testing. Figure 16D. Number of words decoded per minute during real-time testing. In Figures 16A-16D, the distributions are shown as boxplots calculated over n = 34 real-time blocks (in each block, participants attempted spelling 2-5 sentences). Each boxplot shows the quartile of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. In Figures 16A and 16B, each boxplot corresponds to n=34 blocks (in each of these blocks, participants attempted spelling 2-5 sentences).In Figure 16C, each boxplot corresponds to n = 9 blocks (in each of these blocks, participants attempted to spell 2–4 conversational responses). Figure 16E. Number of excess letters in each decoded sentence. Decoded sentences with 0 excess letters indicate that a manual command (to disengage from the speller) was successfully identified from the participant's neural activity immediately after they spelled the last letter of the sentence. Figure 16F. Example sentence spelling trials using decoded sentences from each non-contingency condition. Incorrect letters are colored red. 1 and 2 mark trials in which the real-time decoded sentence contained at least one error. The target sentences for these two trials are given at the bottom of the panel. All other example sentences contained no real-time decoding errors. [Figure 16C]Summary of spelling system performance during a copy-typing task. Figure 16A shows the character error rate (CER) observed during real-time sentence spelling (labeled "+LM (Real-time results)") and offline simulation in which portions of the spelling system were omitted. In the "Chance" condition, sentences were created by replacing the output from the neural classifier with randomly generated character probabilities without modifying the rest of the spelling pipeline. In the "Only neural decoding" condition, sentences were created solely by concatenating the most likely characters from each of the classifier's predictions during sentence testing (without any whitespace). In the "+Vocab. constraints" condition, predicted character probabilities from the neural classifier were used in conjunction with a beam search that constrained predicted character sequences to form words from a 1,152-word vocabulary. The final condition, labeled "+LM (Real-time results)," shows real-time results during testing by participants and incorporates language modeling during beam search and after sentences were finalized. Sentences decoded in real time by the perfect system exhibited lower CERs than sentences decoded in the other conditions (***P<0.0001, two-tailed Wilcoxon rank-sum test with a six-way Holm-Bonferroni correction). Figure 16B. Word error rate (WER) for the real-time results from Figure 16A and the corresponding offline omission simulation. Figure 16C. Number of characters decoded per minute during real-time testing. Figure 16D. Number of words decoded per minute during real-time testing. In Figures 16A-16D, the distributions are shown as boxplots calculated over n = 34 real-time blocks (in each block, participants attempted spelling 2-5 sentences). Each boxplot shows the quartile of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. In Figures 16A and 16B, each boxplot corresponds to n=34 blocks (in each of these blocks, participants attempted spelling 2-5 sentences).In Figure 16C, each boxplot corresponds to n = 9 blocks (in each of these blocks, participants attempted to spell 2–4 conversational responses). Figure 16E. Number of excess letters in each decoded sentence. Decoded sentences with 0 excess letters indicate that a manual command (to disengage from the speller) was successfully identified from the participant's neural activity immediately after they spelled the last letter of the sentence. Figure 16F. Example sentence spelling trials using decoded sentences from each non-contingency condition. Incorrect letters are colored red. 1 and 2 mark trials in which the real-time decoded sentence contained at least one error. The target sentences for these two trials are given at the bottom of the panel. All other example sentences contained no real-time decoding errors. [Figure 16D]Summary of spelling system performance during a copy-typing task. Figure 16A shows the character error rate (CER) observed during real-time sentence spelling (labeled "+LM (Real-time results)") and offline simulation in which portions of the spelling system were omitted. In the "Chance" condition, sentences were created by replacing the output from the neural classifier with randomly generated character probabilities without modifying the rest of the spelling pipeline. In the "Only neural decoding" condition, sentences were created solely by concatenating the most likely characters from each of the classifier's predictions during sentence testing (without any whitespace). In the "+Vocab. constraints" condition, predicted character probabilities from the neural classifier were used in conjunction with a beam search that constrained predicted character sequences to form words from a 1,152-word vocabulary. The final condition, labeled "+LM (Real-time results)," shows real-time results during testing by participants and incorporates language modeling during beam search and after sentences were finalized. Sentences decoded in real time by the perfect system exhibited lower CERs than sentences decoded in the other conditions (***P<0.0001, two-tailed Wilcoxon rank-sum test with a six-way Holm-Bonferroni correction). Figure 16B. Word error rate (WER) for the real-time results from Figure 16A and the corresponding offline omission simulation. Figure 16C. Number of characters decoded per minute during real-time testing. Figure 16D. Number of words decoded per minute during real-time testing. In Figures 16A-16D, the distributions are shown as boxplots calculated over n = 34 real-time blocks (in each block, participants attempted spelling 2-5 sentences). Each boxplot shows the quartile of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. In Figures 16A and 16B, each boxplot corresponds to n=34 blocks (in each of these blocks, participants attempted spelling 2-5 sentences).In Figure 16C, each boxplot corresponds to n = 9 blocks (in each of these blocks, participants attempted to spell 2–4 conversational responses). Figure 16E. Number of excess letters in each decoded sentence. Decoded sentences with 0 excess letters indicate that a manual command (to disengage from the speller) was successfully identified from the participant's neural activity immediately after they spelled the last letter of the sentence. Figure 16F. Example sentence spelling trials using decoded sentences from each non-contingency condition. Incorrect letters are colored red. 1 and 2 mark trials in which the real-time decoded sentence contained at least one error. The target sentences for these two trials are given at the bottom of the panel. All other example sentences contained no real-time decoding errors. [Figure 16E]Summary of spelling system performance during a copy-typing task. Figure 16A shows the character error rate (CER) observed during real-time sentence spelling (labeled "+LM (Real-time results)") and offline simulation in which portions of the spelling system were omitted. In the "Chance" condition, sentences were created by replacing the output from the neural classifier with randomly generated character probabilities without modifying the rest of the spelling pipeline. In the "Only neural decoding" condition, sentences were created solely by concatenating the most likely characters from each of the classifier's predictions during sentence testing (without any whitespace). In the "+Vocab. constraints" condition, predicted character probabilities from the neural classifier were used in conjunction with a beam search that constrained predicted character sequences to form words from a 1,152-word vocabulary. The final condition, labeled "+LM (Real-time results)," shows real-time results during testing by participants and incorporates language modeling during beam search and after sentences were finalized. Sentences decoded in real time by the perfect system exhibited lower CERs than sentences decoded in the other conditions (***P<0.0001, two-tailed Wilcoxon rank-sum test with a six-way Holm-Bonferroni correction). Figure 16B. Word error rate (WER) for the real-time results from Figure 16A and the corresponding offline omission simulation. Figure 16C. Number of characters decoded per minute during real-time testing. Figure 16D. Number of words decoded per minute during real-time testing. In Figures 16A-16D, the distributions are shown as boxplots calculated over n = 34 real-time blocks (in each block, participants attempted spelling 2-5 sentences). Each boxplot shows the quartile of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. In Figures 16A and 16B, each boxplot corresponds to n=34 blocks (in each of these blocks, participants attempted spelling 2-5 sentences).In Figure 16C, each boxplot corresponds to n = 9 blocks (in each of these blocks, participants attempted to spell 2–4 conversational responses). Figure 16E. Number of excess letters in each decoded sentence. Decoded sentences with 0 excess letters indicate that a manual command (to disengage from the speller) was successfully identified from the participant's neural activity immediately after they spelled the last letter of the sentence. Figure 16F. Example sentence spelling trials using decoded sentences from each non-contingency condition. Incorrect letters are colored red. 1 and 2 mark trials in which the real-time decoded sentence contained at least one error. The target sentences for these two trials are given at the bottom of the panel. All other example sentences contained no real-time decoding errors. [Figure 16F]Summary of spelling system performance during a copy-typing task. Figure 16A shows the character error rate (CER) observed during real-time sentence spelling (labeled "+LM (Real-time results)") and offline simulation in which portions of the spelling system were omitted. In the "Chance" condition, sentences were created by replacing the output from the neural classifier with randomly generated character probabilities without modifying the rest of the spelling pipeline. In the "Only neural decoding" condition, sentences were created solely by concatenating the most likely characters from each of the classifier's predictions during sentence testing (without any whitespace). In the "+Vocab. constraints" condition, predicted character probabilities from the neural classifier were used in conjunction with a beam search that constrained predicted character sequences to form words from a 1,152-word vocabulary. The final condition, labeled "+LM (Real-time results)," shows real-time results during testing by participants and incorporates language modeling during beam search and after sentences were finalized. Sentences decoded in real time by the perfect system exhibited lower CERs than sentences decoded in the other conditions (***P<0.0001, two-tailed Wilcoxon rank-sum test with a six-way Holm-Bonferroni correction). Figure 16B. Word error rate (WER) for the real-time results from Figure 16A and the corresponding offline omission simulation. Figure 16C. Number of characters decoded per minute during real-time testing. Figure 16D. Number of words decoded per minute during real-time testing. In Figures 16A-16D, the distributions are shown as boxplots calculated over n = 34 real-time blocks (in each block, participants attempted spelling 2-5 sentences). Each boxplot shows the quartile of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. In Figures 16A and 16B, each boxplot corresponds to n=34 blocks (in each of these blocks, participants attempted spelling 2-5 sentences).In Figure 16C, each boxplot corresponds to n = 9 blocks (in each of these blocks, participants attempted to spell 2–4 conversational responses). Figure 16E. Number of excess letters in each decoded sentence. Decoded sentences with 0 excess letters indicate that a manual command (to disengage from the speller) was successfully identified from the participant's neural activity immediately after they spelled the last letter of the sentence. Figure 16F. Example sentence spelling trials using decoded sentences from each non-contingency condition. Incorrect letters are colored red. 1 and 2 mark trials in which the real-time decoded sentence contained at least one error. The target sentences for these two trials are given at the bottom of the panel. All other example sentences contained no real-time decoding errors. [Figure 17A]Characterization of high gamma activity (HGA) and low frequency signals (LFS) during silent speech trials. Figure 17A shows the 10-fold cross-validation classification accuracy of NATO codewords tried in silence using HGA alone, LFS alone, and both HGA and LFS simultaneously. Classification accuracy using LFS alone is significantly higher than that using HGA alone, and classification accuracy using HGA and LFS results is significantly higher than that using either feature type alone (**P<0.001, two-tailed Wilcoxon rank-sum test with three-way Holm-Bonferroni correction). The chance accuracy is 3.7%. Each boxplot shows the quartiles of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. Each boxplot corresponds to n=10 cross-validation folds. Figure 17B shows electrode contributions from a classification model trained using only HGA features. The size and opacity of the plotted electrodes were scaled by their relative contribution, with larger, more opaque-appearing electrodes providing more significant features to the classification model. Figure 17C shows electrode contributions associated with HGA features from a classification model trained using the combined HGA+LFS feature set. Figure 17D shows electrode contributions from a classification model trained using only LFS features. Figure 17E shows electrode contributions associated with LFS features from a classification model trained using the combined HGA+LFS feature set. In Figures 17B-17E, the size and opacity of the plotted electrodes were scaled by their relative contribution, with larger, more opaque-appearing electrodes providing more significant features to the classification model. Figure 17F shows the minimum number of principal components (PCs) required to explain more than 80% of the variance in the spatial dimensions of each feature set across 100 or more bootstrap iterations. The number of PCs required differed significantly for each feature set (***P<0.0001, two-tailed Wilcoxon rank sum test with three-way Holm-Bonferroni correction; *P<0.01, two-tailed Wilcoxon rank sum test with three-way Holm-Bonferroni correction).Figure 17G. The minimum number of PCs required to explain more than 80% of the variance in the temporal dimension for each feature set across over 100 bootstrap iterations. In Figures 17F and 17G, the number of PCs required for each feature set is shown as a histogram, where the x-axis is the percentage of bootstrap iterations that required a particular number of PCs. Figure 17H. The effect of temporal smoothing on classification accuracy. Each point represents the median, and the error bars represent the 99% confidence interval around the bootstrapped estimate of the median. [Figure 17B]Characterization of high gamma activity (HGA) and low frequency signals (LFS) during silent speech trials. Figure 17A shows the 10-fold cross-validation classification accuracy of NATO codewords tried in silence using HGA alone, LFS alone, and both HGA and LFS simultaneously. Classification accuracy using LFS alone is significantly higher than that using HGA alone, and classification accuracy using HGA and LFS results is significantly higher than that using either feature type alone (**P<0.001, two-tailed Wilcoxon rank-sum test with three-way Holm-Bonferroni correction). The chance accuracy is 3.7%. Each boxplot shows the quartiles of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. Each boxplot corresponds to n=10 cross-validation folds. Figure 17B shows electrode contributions from a classification model trained using only HGA features. The size and opacity of the plotted electrodes were scaled by their relative contribution, with larger, more opaque-appearing electrodes providing more significant features to the classification model. Figure 17C shows electrode contributions associated with HGA features from a classification model trained using the combined HGA+LFS feature set. Figure 17D shows electrode contributions from a classification model trained using only LFS features. Figure 17E shows electrode contributions associated with LFS features from a classification model trained using the combined HGA+LFS feature set. In Figures 17B-17E, the size and opacity of the plotted electrodes were scaled by their relative contribution, with larger, more opaque-appearing electrodes providing more significant features to the classification model. Figure 17F shows the minimum number of principal components (PCs) required to explain more than 80% of the variance in the spatial dimensions of each feature set across 100 or more bootstrap iterations. The number of PCs required differed significantly for each feature set (***P<0.0001, two-tailed Wilcoxon rank sum test with three-way Holm-Bonferroni correction; *P<0.01, two-tailed Wilcoxon rank sum test with three-way Holm-Bonferroni correction).Figure 17G. The minimum number of PCs required to explain more than 80% of the variance in the temporal dimension for each feature set across over 100 bootstrap iterations. In Figures 17F and 17G, the number of PCs required for each feature set is shown as a histogram, where the x-axis is the percentage of bootstrap iterations that required a particular number of PCs. Figure 17H. The effect of temporal smoothing on classification accuracy. Each point represents the median, and the error bars represent the 99% confidence interval around the bootstrapped estimate of the median. [Figure 17C]Characterization of high gamma activity (HGA) and low frequency signals (LFS) during silent speech trials. Figure 17A shows the 10-fold cross-validation classification accuracy of NATO codewords tried in silence using HGA alone, LFS alone, and both HGA and LFS simultaneously. Classification accuracy using LFS alone is significantly higher than that using HGA alone, and classification accuracy using HGA and LFS results is significantly higher than that using either feature type alone (**P<0.001, two-tailed Wilcoxon rank-sum test with three-way Holm-Bonferroni correction). The chance accuracy is 3.7%. Each boxplot shows the quartiles of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. Each boxplot corresponds to n=10 cross-validation folds. Figure 17B shows electrode contributions from a classification model trained using only HGA features. The size and opacity of the plotted electrodes were scaled by their relative contribution, with larger, more opaque-appearing electrodes providing more significant features to the classification model. Figure 17C shows electrode contributions associated with HGA features from a classification model trained using the combined HGA+LFS feature set. Figure 17D shows electrode contributions from a classification model trained using only LFS features. Figure 17E shows electrode contributions associated with LFS features from a classification model trained using the combined HGA+LFS feature set. In Figures 17B-17E, the size and opacity of the plotted electrodes were scaled by their relative contribution, with larger, more opaque-appearing electrodes providing more significant features to the classification model. Figure 17F shows the minimum number of principal components (PCs) required to explain more than 80% of the variance in the spatial dimensions of each feature set across 100 or more bootstrap iterations. The number of PCs required differed significantly for each feature set (***P<0.0001, two-tailed Wilcoxon rank sum test with three-way Holm-Bonferroni correction; *P<0.01, two-tailed Wilcoxon rank sum test with three-way Holm-Bonferroni correction).Figure 17G. The minimum number of PCs required to explain more than 80% of the variance in the temporal dimension for each feature set across over 100 bootstrap iterations. In Figures 17F and 17G, the number of PCs required for each feature set is shown as a histogram, where the x-axis is the percentage of bootstrap iterations that required a particular number of PCs. Figure 17H. The effect of temporal smoothing on classification accuracy. Each point represents the median, and the error bars represent the 99% confidence interval around the bootstrapped estimate of the median. [Figure 17D]Characterization of high gamma activity (HGA) and low frequency signals (LFS) during silent speech trials. Figure 17A shows the 10-fold cross-validation classification accuracy of NATO codewords tried in silence using HGA alone, LFS alone, and both HGA and LFS simultaneously. Classification accuracy using LFS alone is significantly higher than that using HGA alone, and classification accuracy using HGA and LFS results is significantly higher than that using either feature type alone (**P<0.001, two-tailed Wilcoxon rank-sum test with three-way Holm-Bonferroni correction). The chance accuracy is 3.7%. Each boxplot shows the quartiles of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. Each boxplot corresponds to n=10 cross-validation folds. Figure 17B shows electrode contributions from a classification model trained using only HGA features. The size and opacity of the plotted electrodes were scaled by their relative contribution, with larger, more opaque-appearing electrodes providing more significant features to the classification model. Figure 17C shows electrode contributions associated with HGA features from a classification model trained using the combined HGA+LFS feature set. Figure 17D shows electrode contributions from a classification model trained using only LFS features. Figure 17E shows electrode contributions associated with LFS features from a classification model trained using the combined HGA+LFS feature set. In Figures 17B-17E, the size and opacity of the plotted electrodes were scaled by their relative contribution, with larger, more opaque-appearing electrodes providing more significant features to the classification model. Figure 17F shows the minimum number of principal components (PCs) required to explain more than 80% of the variance in the spatial dimensions of each feature set across 100 or more bootstrap iterations. The number of PCs required differed significantly for each feature set (***P<0.0001, two-tailed Wilcoxon rank sum test with three-way Holm-Bonferroni correction; *P<0.01, two-tailed Wilcoxon rank sum test with three-way Holm-Bonferroni correction).Figure 17G. The minimum number of PCs required to explain more than 80% of the variance in the temporal dimension for each feature set across over 100 bootstrap iterations. In Figures 17F and 17G, the number of PCs required for each feature set is shown as a histogram, where the x-axis is the percentage of bootstrap iterations that required a particular number of PCs. Figure 17H. The effect of temporal smoothing on classification accuracy. Each point represents the median, and the error bars represent the 99% confidence interval around the bootstrapped estimate of the median. [Figure 17E]Characterization of high gamma activity (HGA) and low frequency signals (LFS) during silent speech trials. Figure 17A shows the 10-fold cross-validation classification accuracy of NATO codewords tried in silence using HGA alone, LFS alone, and both HGA and LFS simultaneously. Classification accuracy using LFS alone is significantly higher than that using HGA alone, and classification accuracy using HGA and LFS results is significantly higher than that using either feature type alone (**P<0.001, two-tailed Wilcoxon rank-sum test with three-way Holm-Bonferroni correction). The chance accuracy is 3.7%. Each boxplot shows the quartiles of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. Each boxplot corresponds to n=10 cross-validation folds. Figure 17B shows electrode contributions from a classification model trained using only HGA features. The size and opacity of the plotted electrodes were scaled by their relative contribution, with larger, more opaque-appearing electrodes providing more significant features to the classification model. Figure 17C shows electrode contributions associated with HGA features from a classification model trained using the combined HGA+LFS feature set. Figure 17D shows electrode contributions from a classification model trained using only LFS features. Figure 17E shows electrode contributions associated with LFS features from a classification model trained using the combined HGA+LFS feature set. In Figures 17B-17E, the size and opacity of the plotted electrodes were scaled by their relative contribution, with larger, more opaque-appearing electrodes providing more significant features to the classification model. Figure 17F shows the minimum number of principal components (PCs) required to explain more than 80% of the variance in the spatial dimensions of each feature set across 100 or more bootstrap iterations. The number of PCs required differed significantly for each feature set (***P<0.0001, two-tailed Wilcoxon rank sum test with three-way Holm-Bonferroni correction; *P<0.01, two-tailed Wilcoxon rank sum test with three-way Holm-Bonferroni correction).Figure 17G. The minimum number of PCs required to explain more than 80% of the variance in the temporal dimension for each feature set across over 100 bootstrap iterations. In Figures 17F and 17G, the number of PCs required for each feature set is shown as a histogram, where the x-axis is the percentage of bootstrap iterations that required a particular number of PCs. Figure 17H. The effect of temporal smoothing on classification accuracy. Each point represents the median, and the error bars represent the 99% confidence interval around the bootstrapped estimate of the median. [Figure 17F]Characterization of high gamma activity (HGA) and low frequency signals (LFS) during silent speech trials. Figure 17A shows the 10-fold cross-validation classification accuracy of NATO codewords tried in silence using HGA alone, LFS alone, and both HGA and LFS simultaneously. Classification accuracy using LFS alone is significantly higher than that using HGA alone, and classification accuracy using HGA and LFS results is significantly higher than that using either feature type alone (**P<0.001, two-tailed Wilcoxon rank-sum test with three-way Holm-Bonferroni correction). The chance accuracy is 3.7%. Each boxplot shows the quartiles of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. Each boxplot corresponds to n=10 cross-validation folds. Figure 17B shows electrode contributions from a classification model trained using only HGA features. The size and opacity of the plotted electrodes were scaled by their relative contribution, with larger, more opaque-appearing electrodes providing more significant features to the classification model. Figure 17C shows electrode contributions associated with HGA features from a classification model trained using the combined HGA+LFS feature set. Figure 17D shows electrode contributions from a classification model trained using only LFS features. Figure 17E shows electrode contributions associated with LFS features from a classification model trained using the combined HGA+LFS feature set. In Figures 17B-17E, the size and opacity of the plotted electrodes were scaled by their relative contribution, with larger, more opaque-appearing electrodes providing more significant features to the classification model. Figure 17F shows the minimum number of principal components (PCs) required to explain more than 80% of the variance in the spatial dimensions of each feature set across 100 or more bootstrap iterations. The number of PCs required differed significantly for each feature set (***P<0.0001, two-tailed Wilcoxon rank sum test with three-way Holm-Bonferroni correction; *P<0.01, two-tailed Wilcoxon rank sum test with three-way Holm-Bonferroni correction).Figure 17G. The minimum number of PCs required to explain more than 80% of the variance in the temporal dimension for each feature set across over 100 bootstrap iterations. In Figures 17F and 17G, the number of PCs required for each feature set is shown as a histogram, where the x-axis is the percentage of bootstrap iterations that required a particular number of PCs. Figure 17H. The effect of temporal smoothing on classification accuracy. Each point represents the median, and the error bars represent the 99% confidence interval around the bootstrapped estimate of the median. [Figure 17G]Characterization of high gamma activity (HGA) and low frequency signals (LFS) during silent speech trials. Figure 17A shows the 10-fold cross-validation classification accuracy of NATO codewords tried in silence using HGA alone, LFS alone, and both HGA and LFS simultaneously. Classification accuracy using LFS alone is significantly higher than that using HGA alone, and classification accuracy using HGA and LFS results is significantly higher than that using either feature type alone (**P<0.001, two-tailed Wilcoxon rank-sum test with three-way Holm-Bonferroni correction). The chance accuracy is 3.7%. Each boxplot shows the quartiles of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. Each boxplot corresponds to n=10 cross-validation folds. Figure 17B shows electrode contributions from a classification model trained using only HGA features. The size and opacity of the plotted electrodes were scaled by their relative contribution, with larger, more opaque-appearing electrodes providing more significant features to the classification model. Figure 17C shows electrode contributions associated with HGA features from a classification model trained using the combined HGA+LFS feature set. Figure 17D shows electrode contributions from a classification model trained using only LFS features. Figure 17E shows electrode contributions associated with LFS features from a classification model trained using the combined HGA+LFS feature set. In Figures 17B-17E, the size and opacity of the plotted electrodes were scaled by their relative contribution, with larger, more opaque-appearing electrodes providing more significant features to the classification model. Figure 17F shows the minimum number of principal components (PCs) required to explain more than 80% of the variance in the spatial dimensions of each feature set across 100 or more bootstrap iterations. The number of PCs required differed significantly for each feature set (***P<0.0001, two-tailed Wilcoxon rank sum test with three-way Holm-Bonferroni correction; *P<0.01, two-tailed Wilcoxon rank sum test with three-way Holm-Bonferroni correction).Figure 17G. The minimum number of PCs required to explain more than 80% of the variance in the temporal dimension for each feature set across over 100 bootstrap iterations. In Figures 17F and 17G, the number of PCs required for each feature set is shown as a histogram, where the x-axis is the percentage of bootstrap iterations that required a particular number of PCs. Figure 17H. The effect of temporal smoothing on classification accuracy. Each point represents the median, and the error bars represent the 99% confidence interval around the bootstrapped estimate of the median. [Figure 17H]Characterization of high gamma activity (HGA) and low frequency signals (LFS) during silent speech trials. Figure 17A shows the 10-fold cross-validation classification accuracy of NATO codewords tried in silence using HGA alone, LFS alone, and both HGA and LFS simultaneously. Classification accuracy using LFS alone is significantly higher than that using HGA alone, and classification accuracy using HGA and LFS results is significantly higher than that using either feature type alone (**P<0.001, two-tailed Wilcoxon rank-sum test with three-way Holm-Bonferroni correction). The chance accuracy is 3.7%. Each boxplot shows the quartiles of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. Each boxplot corresponds to n=10 cross-validation folds. Figure 17B shows electrode contributions from a classification model trained using only HGA features. The size and opacity of the plotted electrodes were scaled by their relative contribution, with larger, more opaque-appearing electrodes providing more significant features to the classification model. Figure 17C shows electrode contributions associated with HGA features from a classification model trained using the combined HGA+LFS feature set. Figure 17D shows electrode contributions from a classification model trained using only LFS features. Figure 17E shows electrode contributions associated with LFS features from a classification model trained using the combined HGA+LFS feature set. In Figures 17B-17E, the size and opacity of the plotted electrodes were scaled by their relative contribution, with larger, more opaque-appearing electrodes providing more significant features to the classification model. Figure 17F shows the minimum number of principal components (PCs) required to explain more than 80% of the variance in the spatial dimensions of each feature set across 100 or more bootstrap iterations. The number of PCs required differed significantly for each feature set (***P<0.0001, two-tailed Wilcoxon rank sum test with three-way Holm-Bonferroni correction; *P<0.01, two-tailed Wilcoxon rank sum test with three-way Holm-Bonferroni correction).Figure 17G. The minimum number of PCs required to explain more than 80% of the variance in the temporal dimension for each feature set across over 100 bootstrap iterations. In Figures 17F and 17G, the number of PCs required for each feature set is shown as a histogram, where the x-axis is the percentage of bootstrap iterations that required a particular number of PCs. Figure 17H. The effect of temporal smoothing on classification accuracy. Each point represents the median, and the error bars represent the 99% confidence interval around the bootstrapped estimate of the median. [Figure 18A] Comparison of neural signals during silent speech trials of English letters and NATO code words. Figure 18A. Classification accuracy (n = 10 cross-validation folds) using a model trained with HGA+LFS features is significantly higher for NATO code words than for English letters (**P < 0.001, two-sided Wilcoxon rank-sum test). The dotted horizontal line represents chance accuracy. Figure 18B. Nearest class distance for the combined HGA+LFS feature set is significantly greater for NATO code words than for letters (box plots show values across n = 26 code words or letters; *P < 0.01, two-sided Wilcoxon rank-sum test). In Figures 18A and 18B, each box plot shows quartiles of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. Figure 18C. Nearest class distance is greater for the majority of code words than for the corresponding letters. In Figures 18B and 18C, the nearest class distance is calculated as the Frobenius norm between the test average HGA+LFS features. [Figure 18B]Comparison of neural signals during silent speech trials of English letters and NATO code words. Figure 18A. Classification accuracy (n = 10 cross-validation folds) using a model trained with HGA+LFS features is significantly higher for NATO code words than for English letters (**P < 0.001, two-sided Wilcoxon rank-sum test). The dotted horizontal line represents chance accuracy. Figure 18B. Nearest class distance for the combined HGA+LFS feature set is significantly greater for NATO code words than for letters (box plots show values across n = 26 code words or letters; *P < 0.01, two-sided Wilcoxon rank-sum test). In Figures 18A and 18B, each box plot shows quartiles of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. Figure 18C. Nearest class distance is greater for the majority of code words than for the corresponding letters. In Figures 18B and 18C, the nearest class distance is calculated as the Frobenius norm between the test average HGA+LFS features. [Figure 18C] Comparison of neural signals during silent speech trials of English letters and NATO code words. Figure 18A. Classification accuracy (n = 10 cross-validation folds) using a model trained with HGA+LFS features is significantly higher for NATO code words than for English letters (**P < 0.001, two-sided Wilcoxon rank-sum test). The dotted horizontal line represents chance accuracy. Figure 18B. Nearest class distance for the combined HGA+LFS feature set is significantly greater for NATO code words than for letters (box plots show values across n = 26 code words or letters; *P < 0.01, two-sided Wilcoxon rank-sum test). In Figures 18A and 18B, each box plot shows quartiles of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. Figure 18C. Nearest class distance is greater for the majority of code words than for the corresponding letters. In Figures 18B and 18C, the nearest class distance is calculated as the Frobenius norm between the test average HGA+LFS features. [Figure 19A]Differences in neural signals and classification performance between overt and silent speech trials. Figure 19A. MRI reconstruction of a participant's brain with the locations of implanted electrodes overlaid. The locations of the electrodes used in Figures 19B and 19C are bolded and numbered in the overlay. Figure 19B. High-gamma activity (HGA) event-related potentials between silent (orange) and overt (green) speech trials for the NATO code word "kilo." Figure 19C. High-gamma activity (HGA) event-related potentials between silent (orange) and overt (green) speech trials for the NATO code word "tango." Evoked responses in Figures 19B and C are aligned to the go cue, marked as a vertical dashed line at time 0. Each curve shows the mean ± standard error across n = 100 speech trials. Figure 19D. Codeword classification accuracy (across 10 cross-validation folds) with various model training schemes. All comparisons revealed significant differences between pairs of results (P<0.01, two-tailed Wilcoxon rank sum with 28-way Holm-Bonferroni correction), except for those marked "ns." Each box plot corresponds to n=10 cross-validation folds. Chance accuracy is 3.84%. [Figure 19B]Differences in neural signals and classification performance between overt and silent speech trials. Figure 19A. MRI reconstruction of a participant's brain with the locations of implanted electrodes overlaid. The locations of the electrodes used in Figures 19B and 19C are bolded and numbered in the overlay. Figure 19B. High-gamma activity (HGA) event-related potentials between silent (orange) and overt (green) speech trials for the NATO code word "kilo." Figure 19C. High-gamma activity (HGA) event-related potentials between silent (orange) and overt (green) speech trials for the NATO code word "tango." Evoked responses in Figures 19B and C are aligned to the go cue, marked as a vertical dashed line at time 0. Each curve shows the mean ± standard error across n = 100 speech trials. Figure 19D. Codeword classification accuracy (across 10 cross-validation folds) with various model training schemes. All comparisons revealed significant differences between pairs of results (P<0.01, two-tailed Wilcoxon rank sum with 28-way Holm-Bonferroni correction), except for those marked "ns." Each box plot corresponds to n=10 cross-validation folds. Chance accuracy is 3.84%. [Figure 19C]Differences in neural signals and classification performance between overt and silent speech trials. Figure 19A. MRI reconstruction of a participant's brain with the locations of implanted electrodes overlaid. The locations of the electrodes used in Figures 19B and 19C are bolded and numbered in the overlay. Figure 19B. High-gamma activity (HGA) event-related potentials between silent (orange) and overt (green) speech trials for the NATO code word "kilo." Figure 19C. High-gamma activity (HGA) event-related potentials between silent (orange) and overt (green) speech trials for the NATO code word "tango." Evoked responses in Figures 19B and C are aligned to the go cue, marked as a vertical dashed line at time 0. Each curve shows the mean ± standard error across n = 100 speech trials. Figure 19D. Codeword classification accuracy (across 10 cross-validation folds) with various model training schemes. All comparisons revealed significant differences between pairs of results (P<0.01, two-tailed Wilcoxon rank sum with 28-way Holm-Bonferroni correction), except for those marked "ns." Each box plot corresponds to n=10 cross-validation folds. Chance accuracy is 3.84%. [Figure 19D]Differences in neural signals and classification performance between overt and silent speech trials. Figure 19A. MRI reconstruction of a participant's brain with the locations of implanted electrodes overlaid. The locations of the electrodes used in Figures 19B and 19C are bolded and numbered in the overlay. Figure 19B. High-gamma activity (HGA) event-related potentials between silent (orange) and overt (green) speech trials for the NATO code word "kilo." Figure 19C. High-gamma activity (HGA) event-related potentials between silent (orange) and overt (green) speech trials for the NATO code word "tango." Evoked responses in Figures 19B and C are aligned to the go cue, marked as a vertical dashed line at time 0. Each curve shows the mean ± standard error across n = 100 speech trials. Figure 19D. Codeword classification accuracy (across 10 cross-validation folds) with various model training schemes. All comparisons revealed significant differences between pairs of results (P<0.01, two-tailed Wilcoxon rank sum with 28-way Holm-Bonferroni correction), except for those marked "ns." Each box plot corresponds to n=10 cross-validation folds. Chance accuracy is 3.84%. [Figure 20A] The spelling approach can be generalized to larger vocabularies and conversational settings. Figure 20A. Simulated letter error rates from a copy-typing task with different vocabularies, including the original vocabulary used during real-time decoding. Figure 20B. Word error rates from the corresponding simulation in Figure 20A. Figure 20C. Letter and word error rates across intentionally selected responses and messages decoded in real time during the conversational task condition. In Figures 20A-20C, each boxplot shows the quartile of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. Figure 20D. Example questions presented from a conversational task condition trial (left) with corresponding responses (right) decoded from the participant's brain activity. In the last example, the participant spelled out the intended message without being prompted by a question. [Figure 20B]The spelling approach can be generalized to larger vocabularies and conversational settings. Figure 20A. Simulated letter error rates from a copy-typing task with different vocabularies, including the original vocabulary used during real-time decoding. Figure 20B. Word error rates from the corresponding simulation in Figure 20A. Figure 20C. Letter and word error rates across intentionally selected responses and messages decoded in real time during the conversational task condition. In Figures 20A-20C, each boxplot shows the quartile of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. Figure 20D. Example questions presented from a conversational task condition trial (left) with corresponding responses (right) decoded from the participant's brain activity. In the last example, the participant spelled out the intended message without being prompted by a question. [Figure 20C] The spelling approach can be generalized to larger vocabularies and conversational settings. Figure 20A. Simulated letter error rates from a copy-typing task with different vocabularies, including the original vocabulary used during real-time decoding. Figure 20B. Word error rates from the corresponding simulation in Figure 20A. Figure 20C. Letter and word error rates across intentionally selected responses and messages decoded in real time during the conversational task condition. In Figures 20A-20C, each boxplot shows the quartile of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. Figure 20D. Example questions presented from a conversational task condition trial (left) with corresponding responses (right) decoded from the participant's brain activity. In the last example, the participant spelled out the intended message without being prompted by a question. [Figure 20D]The spelling approach can be generalized to larger vocabularies and conversational settings. Figure 20A. Simulated letter error rates from a copy-typing task with different vocabularies, including the original vocabulary used during real-time decoding. Figure 20B. Word error rates from the corresponding simulation in Figure 20A. Figure 20C. Letter and word error rates across intentionally selected responses and messages decoded in real time during the conversational task condition. In Figures 20A-20C, each boxplot shows the quartile of the data with whiskers extending to show the remainder of the distribution, excluding data points that are 1.5 times the interquartile range. Figure 20D. Example questions presented from a conversational task condition trial (left) with corresponding responses (right) decoded from the participant's brain activity. In the last example, the participant spelled out the intended message without being prompted by a question. [Figure 21]Data collection timeline. Each bar indicates the total number of trials collected on each day of recording. Participants and implantation dates are the same as in our previous study [2]. When multiple types of datasets were collected on a single day, the bars are colored by the proportion of each dataset collected. Each color represents a specific dataset (as specified in the legend). Datasets differed by task type (isolated target or real-time sentence spelling), utterance set (English letters, NATO codewords (including attempted hand grasps), copy-typing sentences, or conversational sentences), and, in the case of the real-time sentence spelling dataset, data purpose (for hyperparameter optimization or performance evaluation). All speech-related trials were associated with silent speech trials, except for datasets with "(overt)" in their legend label. Additionally, 3.06% of trials in this overt dataset were actually recorded during a version of the copy-typing sentence spelling task in which participants attempted overt production of codewords (see Section S3 for further details). Datasets were collected on an irregular schedule due to external and clinical time constraints that were unrelated to the neural implants. The gap between 55 and 88 weeks was due to clinical guidelines, particularly during the onset of the COVID-19 pandemic, which limited or prevented in-person recording sessions. [Figure 22]Real-time signal processing pipeline. A detachable data acquisition headstage (CerePlex E, Blackrock Microsystems) attached to the transcutaneous pedestal connector applied a hardware-based wideband Butterworth filter (0.3 Hz–7.5 kHz) to the ECoG signals, digitized them at 16 bits with 250-nV / bit resolution, and transmitted them to a Neuroport system (Blackrock Microsystems) through an additional connection at 30 kHz. The Neuroport system processed the signals using software-based line noise cancellation and an anti-aliasing low-pass filter (500 Hz). The processed signals were then streamed at 1 kHz to a separate computer for further real-time processing and analysis, where a common average reference was applied to each time sample of the ECoG data (across all electrode channels). The rereferenced signals were then processed in two parallel streams to extract high-gamma activity (HGA) and low-frequency signal (LFS) features. To calculate the HGA features, eight 390th-order bandpass finite impulse response (FIR) filters were applied to the rereferenced signal (filter center frequencies were within the high-gamma band at 72.0, 79.5, 87.8, 96.9, 107.0, 118.1, 130.4, and 144.0 Hz). A 170th-order FIR filter was then used to approximate the Hilbert transform for each channel and band. Specifically, for each channel and band, the real component of the analytic signal was set equal to the original signal delayed by 85 samples (half the filter order), and the imaginary component was set equal to the Hilbert transform of the original signal (approximated by this FIR filter)
[25] . The magnitude of each analytic signal at every fifth time sample was then calculated, generating a 200 Hz analytic amplitude signal. For each channel, the analytic amplitude values were averaged across the eight bands at each time point to obtain a single high-gamma analytic amplitude measure for that channel. To compute the LFS features, the re-referenced signal was downsampled to 200 Hz after applying a 130th-order anti-aliasing low-pass FIR filter with a cutoff frequency of 100 Hz.We then combined the time-synchronized values from the two feature streams (high-gamma analysis amplitude and downsampled signal) into a single feature stream. We then z-scored the values for each channel and feature type using Welford's method in a 30-second sliding window
[26] . Finally, we implemented a simple artifact removal technique to prevent samples with very large z-score magnitudes from interfering with the ongoing z-score statistics or downstream decoding process. We adapted this figure from previous work [2, 27], which implemented a similar preprocessing pipeline for computing high-gamma features. [Figure 23] Schematic of the speech detection model. To detect silent speech attempts from a participant's neural activity during real-time sentence spelling, z-scored low-frequency signals (LFS) and high-gamma activity (HGA) from each electrode are first processed sequentially by a stack of three long-short-term memory (LSTM) layers. Next, a single dense (fully connected) layer projects the final LSTM's latent dimensions into four possible classes: speech, speech preparation, pause, and movement. The stream of speech probabilities is then temporally smoothed, probability-thresholded, and time-thresholded to yield the onset and termination of complete speech events. When the participant silently attempts to speak something and the speech attempt is detected, the spelling system is engaged and a paced spelling procedure begins. The depicted neural features, predicted speech probability time series (top right), and detected speech events (bottom right) are actual neural data and detection results over a 5-second time window at the beginning of a trial of a real-time sentence copytyping task. This diagram is adapted from our previous work [2], which implemented a similar speech detection architecture. [Figure 24A]Impact of feature selection on codeword classification accuracy. Figure 24A. Using high gamma activity (HGA) and low frequency signal (LFS) together (combined HGA+LFS feature set) rather than HGA features alone improves classification accuracy for each codeword. Figure 24B. Using HGA+LFS rather than LFS alone improves classification accuracy for almost all codewords. In both Figures 24A and 24B, codewords are represented as lowercase letters, and Spearman rank correlations are shown. Associated p-values were calculated via permutation testing, where one group of observations (either HGA, LFS, or HGA+LFS codeword accuracy) was shuffled before recalculating the correlation between that group of observations and the other group. For each of the two comparisons, 2000 iterations were used during permutation testing. [Figure 24B] Impact of feature selection on codeword classification accuracy. Figure 24A. Using high gamma activity (HGA) and low frequency signal (LFS) together (combined HGA+LFS feature set) rather than HGA features alone improves classification accuracy for each codeword. Figure 24B. Using HGA+LFS rather than LFS alone improves classification accuracy for almost all codewords. In both Figures 24A and 24B, codewords are represented as lowercase letters, and Spearman rank correlations are shown. Associated p-values were calculated via permutation testing, where one group of observations (either HGA, LFS, or HGA+LFS codeword accuracy) was shuffled before recalculating the correlation between that group of observations and the other group. For each of the two comparisons, 2000 iterations were used during permutation testing. [Figure 25]Confusion matrix from isolated target trial classification. Confusion values calculated during offline classification of neural data recorded during isolated target trials (using both high-gamma activity and low-frequency signals) are shown for each NATO codeword and attempted hand grasp. Each row corresponds to a target codeword or attempted hand grasp, and the values within each column of that row correspond to the proportion of isolated target task trials that were correctly classified as a target (if the values fall along the diagonal) or incorrectly classified ("confused") as another potential target (if the values do not fall along the diagonal). The values for each row sum to 100%. In general, silent speech and hand grasp trials were reliably classified. [Figure 26A]Neural activation characteristics during overt and silent speech trials. Figure 26A. Each image shows an MRI reconstruction of a participant's brain superimposed with the electrode location and maximum neural activation for each electrode, speech trial type (overt or silent), and feature type (high gamma activity (HGA) or low frequency signal (LFS)), measured as the maximum peak codeword mean magnitude. To calculate these values, a trial-averaged neural feature time series was calculated for each codeword, electrode, speech trial type, and feature type using the isolated target dataset (a 2.5-second time window after the go cue was used for each trial). The peak magnitude (maximum absolute value) of each of these trial-averaged time series was then determined. The maximum peak codeword mean magnitude for each electrode, speech trial type, and feature type was then calculated as the maximum of these peak magnitudes across each combination of codewords. Two columns show the values for each type of speech trial (overt, then silent), and two columns show the values for each feature type (HGA, then LFS). Figure 26B. Standard deviation of peak codeword mean magnitude. Here, the standard deviation of peak mean magnitude across codewords for each electrode, speech trial type, and feature type (instead of the maximum value used in Figure 26A) is calculated and plotted to show how much magnitude varied across speech targets for that combination. For Figures 26A and 26B, the color of each plotted electrode indicates the true associated value for that electrode, and the size of each electrode indicates the associated value for that electrode compared to the values of other electrodes (for a given type of speech trial and feature type). [Figure 26B]Neural activation characteristics during overt and silent speech trials. Figure 26A. Each image shows an MRI reconstruction of a participant's brain superimposed with the electrode location and maximum neural activation for each electrode, speech trial type (overt or silent), and feature type (high gamma activity (HGA) or low frequency signal (LFS)), measured as the maximum peak codeword mean magnitude. To calculate these values, a trial-averaged neural feature time series was calculated for each codeword, electrode, speech trial type, and feature type using the isolated target dataset (a 2.5-second time window after the go cue was used for each trial). The peak magnitude (maximum absolute value) of each of these trial-averaged time series was then determined. The maximum peak codeword mean magnitude for each electrode, speech trial type, and feature type was then calculated as the maximum of these peak magnitudes across each combination of codewords. Two columns show the values for each type of speech trial (overt, then silent), and two columns show the values for each feature type (HGA, then LFS). Figure 26B. Standard deviation of peak codeword mean magnitude. Here, the standard deviation of peak mean magnitude across codewords for each electrode, speech trial type, and feature type (instead of the maximum value used in Figure 26A) is calculated and plotted to show how much magnitude varied across speech targets for that combination. For Figures 26A and 26B, the color of each plotted electrode indicates the true associated value for that electrode, and the size of each electrode indicates the associated value for that electrode compared to the values of other electrodes (for a given type of speech trial and feature type). DETAILED DESCRIPTION OF THE INVENTION
[0101] Methods, devices, and systems are provided for assisting a subject's communication. Specifically, methods, devices, and systems are provided for decoding words and sentences directly from an individual's neural activity. In the disclosed methods, cortical activity from brain regions involved in speech processing is recorded while the individual attempts to speak or spell out words of a sentence. A deep learning computational model is used to detect and classify words from the recorded brain activity. The decoding of speech from brain activity is aided by the use of a language model that predicts how a particular word sequence is likely to appear. Additionally, decoding of trial non-speech movements from neural activity can be used to further assist communication.
[0102] The methods, devices, and systems disclosed herein can be used to assist individuals with communication difficulties caused by conditions and diseases including, but not limited to, stroke, traumatic brain injury, brain tumor, amyotrophic lateral sclerosis, multiple sclerosis, Huntington's disease, Niemann-Pick disease, Friedreich's ataxia, Wilson's disease, cerebral palsy, Guillain-Barré syndrome, Tay-Sachs disease, encephalopathy, central pontine myelinolysis, and other conditions that cause muscle dysfunction or paralysis of the head, neck, or chest resulting in dysarthria. The methods disclosed herein can be used to restore communication to such individuals and improve their autonomy and quality of life.
[0103] Before describing exemplary embodiments of the invention, it is to be understood that the invention is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.
[0104] Where a range of values is provided, unless the context dictates otherwise, it is understood that each intervening value between the upper and lower limit of that range is also specifically disclosed, to the tenth of the unit of the lower limit. Each smaller range between any stated or intervening value in a stated range and any other stated or intervening value in that stated range is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included or excluded, and each of the ranges in which the smaller ranges include either, neither, or both limits is also encompassed within the invention, subject to any specifically excluded limits in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included within the invention.
[0105] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, some potential exemplary methods and materials may be described here. Any and all publications mentioned herein are incorporated by reference to disclose and describe the methods and / or materials for which the publications are cited. In case of conflict, it should be understood that the present disclosure supersedes any disclosure of the incorporated publication.
[0106] It should be noted that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, a reference to "an electrode" or "the electrode" includes a plurality of such electrodes, a reference to "a signal" or "the signal" includes a reference to one or more signals, etc.
[0107] It is further noted that the claims may be drafted to exclude elements that may be optional. Accordingly, this specification is intended to serve as a predicate basis for using exclusive terminology, such as "solely," "only," or "negative" limitations in connection with the recitation of claim elements.
[0108] The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein should be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed. To the extent that such publications provide definitions for terms that conflict with express or implied definitions in the present disclosure, the definitions in the present disclosure control.
[0109] As will be apparent to those skilled in the art upon reading this disclosure, each of the separate embodiments described and illustrated herein has distinct components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the invention. Any recited method may be carried out in the order of events recited or in any other order which is logically possible.
[0110] definition The term "communication disorders" is used herein to refer to a group of conditions that affect a subject's ability to speak, including, but not limited to, dysarthria, stroke, traumatic brain injury, brain tumor, amyotrophic lateral sclerosis, multiple sclerosis, Huntington's disease, Niemann-Pick disease, Friedreich's ataxia, Wilson's disease, cerebral palsy, Guillain-Barré syndrome, Tay-Sachs disease, encephalopathy, central pontine myelinolysis, and other conditions that cause dysfunction or paralysis of the muscles of the head, neck, or chest, resulting in dysarthria.
[0111] The term "communication" includes word-based communication, such as oral communication, including spoken words, spelling words, and generating text (e.g., controlling a personal device through speech attempts to generate emails or texts), as well as behavior-based communication, such as non-speech motor attempts. Attempted speech may include vocalized speech, which may or may not be intelligible, or unvocalized speech. Silent speech attempts are volitional attempts to articulate speech without vocalization. Silent spelling attempts are volitional attempts to spell alphabets or numbers without vocalization. Attempted non-speech movements may include imaginary movements without any detectable physical movement. Attempted non-speech movements may include, but are not limited to, imaginary head, arm, hand, foot, and leg movements. Attempted non-speech movements may be used to indicate the start or end of attempted speech or spelling, or to control an external device (e.g., to communicate with a personal device or software application, or to turn a device on or off). In the disclosed method, neural activity is recorded during communication attempts, regardless of whether the individual produces any vocal output or detectable movement.
[0112] The terms "subject," "individual," "patient," and "participant" are used interchangeably herein and refer to a patient with a communication disorder. The patient is preferably a human, e.g., a child, adolescent, adult, e.g., a young, middle-aged, or elderly human, who can benefit from the systems, devices, and methods disclosed herein for restoring communication. The patient may have been diagnosed as suffering from dysarthria.
[0113] The term "user" as used herein refers to a person who interacts with the devices and / or systems disclosed herein to perform one or more steps of the methods of the present disclosure. The user may be a patient receiving treatment. The user may also be a medical professional, such as the patient's physician.
[0114] method The present disclosure provides a method for assisting a subject's communication. A method is provided for decoding words and sentences directly from an individual's neural activity. In the disclosed method, cortical activity from brain regions involved in speech processing is recorded while the individual attempts to speak or spell words of a sentence. The attempt to speak or spell words may include or exclude vocalizations. That is, neural activity is recorded during the individual's attempt to speak or spell words, regardless of whether the individual produces any speech output. In some cases, the speech output may be unintelligible when the individual attempts to speak or spell words. A deep learning computational model is used to detect and classify words and / or spelled letters from the recorded brain activity. The decoding of speech from brain activity is aided by the use of a language model that predicts how specific word sequences will appear. The neurotechnology described herein can be used to restore communication to patients who have lost the ability to speak, potentially improving autonomy and quality of life. Various steps and aspects of the method are now described in more detail below.
[0115] The method includes positioning a neurorecording device comprising one or more electrodes at a location within a sensorimotor cortical region of the subject's brain to record electrical brain signal data associated with trial speech and / or trial spellings by the subject, and positioning an interface in communication with a computing device at a location on the subject's head. The electrical brain signal data associated with the trial speech and / or trial spellings by the subject is recorded using the neurorecording device, the interface receives the electrical brain signal data from the neurorecording device and transmits the electrical brain signal data to a processor, the processor being programmed to detect the trial speech and / or spellings by the subject and decode spelled letters, words, phrases, or sentences from the recorded electrical brain signal data.
[0116] The recording device may include a non-brain-invasive surface electrode or a brain-invasive deep electrode. Electrical signals may be recorded using a single electrode, an electrode pair, or an electrode array. In some embodiments, brain activity is recorded from two or more sites. In certain embodiments, brain electrical signal data is recorded from a sensorimotor cortical region of the brain involved in speech processing, such as the precentral gyrus, postcentral gyrus, posterior middle frontal gyrus, posterior superior frontal gyrus, or posterior inferior frontal gyrus region, or any combination thereof. In some embodiments, electrodes are positioned on the surface of the sensorimotor cortical region of the brain within the subdural space.
[0117] Positioning of electrodes to record brain activity in specific regions of the brain may be performed using standard surgical procedures for intracranial electrode placement. As used herein, the phrase "an electrode" or "the electrode" refers to a single electrode or multiple electrodes, such as an electrode array. As used herein, the term "contact," as used in the context of an electrode in contact with a brain region, refers to a physical association between the electrode and the region. In other words, an electrode in contact with a brain region is physically adjacent to the brain region. Electrodes in contact with a brain region can be used to detect electrical signals corresponding to neural activity associated with attempted speech and / or spelling. Electrodes used in the methods disclosed herein may be unipolar (cathode or anode) or bipolar (e.g., having an anode and a cathode).
[0118] In certain embodiments, one or more electrodes are used to record electrical signals for neural activity associated with attempted speech and / or spelling in one or more brain regions. The electrodes may be placed in regions of the sensorimotor cortex involved in speech processing, such as the precentral gyrus, postcentral gyrus, posterior middle frontal gyrus, posterior superior frontal gyrus, or posterior inferior frontal gyrus regions of the brain. In some cases, placing the electrodes may involve positioning the electrodes on the surface of a designated region of the brain. For example, the electrodes may be placed on the surface of the brain in the precentral gyrus, postcentral gyrus, posterior middle frontal gyrus, posterior superior frontal gyrus, or posterior inferior frontal gyrus regions, or any combination thereof. The electrodes may contact at least a portion of the brain's surface in the precentral gyrus, postcentral gyrus, posterior middle frontal gyrus, posterior superior frontal gyrus, or posterior inferior frontal gyrus regions. In some embodiments, the electrodes may contact substantially the entire surface area in the precentral gyrus, postcentral gyrus, posterior middle frontal gyrus, posterior superior frontal gyrus, or posterior inferior frontal gyrus regions. In some embodiments, the electrodes may additionally contact areas adjacent to the precentral gyrus, postcentral gyrus, posterior middle frontal gyrus, posterior superior frontal gyrus, or posterior inferior frontal gyrus regions.
[0119] In some embodiments, an electrode array disposed on a planar support substrate may be used to detect electrical signals of neural activity from one or more of the brain regions specified herein. The surface area of the electrode array may be determined by the desired contact area between the electrode array and the brain. Electrodes for implantation on the brain surface, such as surface electrodes or surface electrode arrays, can be obtained from commercial sources. Commercially available electrodes / electrode arrays may be modified to achieve the desired contact area. In some cases, non-brain-invasive electrodes (also referred to as surface electrodes) that may be used in the methods disclosed herein may be electrocorticography (ECoG) electrodes or electroencephalography (EEG) electrodes.
[0120] In some cases, positioning an electrode at a target region or site (e.g., a neural recording device electrode) may involve positioning a brain-penetrating electrode (also referred to as a depth electrode) within a specific region of the brain. For example, a depth electrode may be placed within a selected region of the sensorimotor cortex involved in speech processing (e.g., the precentral gyrus, postcentral gyrus, posterior middle frontal gyrus, posterior superior frontal gyrus, or posterior inferior frontal gyrus region). In some embodiments, the electrode may additionally contact a region adjacent to the selected region of the sensorimotor cortex involved in speech processing (e.g., adjacent to the precentral gyrus, postcentral gyrus, posterior middle frontal gyrus, posterior superior frontal gyrus, or posterior inferior frontal gyrus region). In some embodiments, an electrode array may be used to record electrical signals in a selected region of the sensorimotor cortex involved in speech processing (e.g., the precentral gyrus, postcentral gyrus, posterior middle frontal gyrus, posterior superior frontal gyrus, or posterior inferior frontal gyrus region), as designated herein.
[0121] The depth to which the electrodes are inserted into the brain can be determined by the desired level of contact between the electrode array and the brain and the type of neural population that the electrodes will have access to for recording electrical signals. Brain-penetrating electrode arrays can be obtained from commercial sources. Commercially available electrode arrays can be modified to achieve the desired insertion depth into brain tissue.
[0122] The exact number of electrodes included in an electrode array (e.g., for recording neural activity associated with trial speech) may vary. In certain embodiments, the electrode array may include two or more electrodes, such as 3 or more, 10 or more, 50 or more, 100 or more, 200 or more, 500 or more, or 4 or more, e.g., about 3-6 electrodes, about 6-12 electrodes, about 12-18 electrodes, about 18-24 electrodes, about 24-30 electrodes, about 30-48 electrodes, about 48-72 electrodes, about 72-96 electrodes, about 96-128 electrodes, about 128-196 electrodes, about 196-294 electrodes, or more. The electrodes may be arranged in a regular, repeating pattern (e.g., a grid, such as a grid with about 1 cm spacing between electrodes), or there may be no pattern. Electrodes can be used that fit the target area for optimal recording of electrical signals from neural activity associated with the subject's attempted speech and / or spelling. One such example is a single multi-contact electrode with eight contacts separated by 2.5 mm. Each contact would have a span of approximately 2 mm. Another example is an electrode with two 1 cm contacts with a 2 mm intervening gap. Yet another example of an electrode that can be used in the present method is a two- or three-pronged electrode to cover the target area. Each of these three-pronged electrodes has four 1- to 2-mm contacts, separated by a center-to-center distance of 2 to 2.5 mm and a span of 1.5 mm.
[0123] In some embodiments, a high-density ECoG electrode array is used to record electrical signals from neural activity associated with speech and / or spelling attempts by a subject. For example, the high-density ECoG electrode array may include at least 100 electrodes, at least 128 electrodes, at least 196 electrodes, at least 256 electrodes, at least 294 electrodes, at least 500 electrodes, or at least 1000 electrodes, or more. In some embodiments, the center-to-center electrode spacing in the high-density ECoG electrode array ranges from 250 mm to 4 mm, including any center-to-center electrode spacing within this range, such as 250 mm, 300 mm, 350 mm, 400 mm, 500 mm, 550 mm, 600 mm, 650 mm, 700 mm, 800 mm, 900 mm, 1 mm, 1.5 mm, 2 mm, 2.5 mm, 3 mm, 3.5 mm, or 4 mm. In some embodiments, a high-density ECoG microelectrode array is used. ECoG microelectrode arrays may include electrodes having diameters of 250 mm or less, 230 mm or less, or 200 mm or less, including electrodes having diameters ranging from 150 mm to 250 mm, including any diameter within this range, such as 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, or 250 mm. For descriptions of high-density ECoG electrode arrays and microelectrode arrays, see, e.g., Muller et al. (2015) Annu Int Conf IEEE Eng Med Biol Soc 2016:1528-1531; Chiang et al. (2020) J. Neural Eng. 17:046008; Escabi et al. (2014) J. Neurophysiol. 112(6):1566-1583, which are incorporated herein by reference.
[0124] The size of each electrode may also vary depending on factors such as the number of electrodes in the array, the location of the electrodes, the material, the age of the patient, and other factors. In certain embodiments, each electrode has a size (e.g., diameter) of about 5 mm or less, such as about 4 mm or less, including 4 mm to 0.25 mm, 3 mm to 0.25 mm, 2 mm to 0.25 mm, 1 mm to 0.25 mm, or about 3 mm, about 2 mm, about 1 mm, about 0.5 mm, or about 0.25 mm.
[0125] In certain embodiments, the method further includes mapping the subject's brain to optimize electrode positioning. The electrode positioning is optimized to detect brain activity characteristics associated with trial speech by the subject and achieve optimal decoding of the trial speech. For example, patterns of electrical signals within specific frequency ranges (e.g., alpha, delta, beta, gamma, and / or high gamma) may be used to detect and decode the trial speech and / or spelling of a word, phrase, or sentence intended by the subject. Thus, electrodes can be positioned to optimize detection and / or decoding of brain activity in specific frequency ranges to restore communication to a subject with a communication disorder.
[0126] In certain aspects, methods and systems of the present disclosure may include recording brain activity, e.g., electrical activity in the ventral sensorimotor cortex, where patterns of gamma frequency neural activity associated with words, phrases, and sentences of trial speech may be detected. In certain cases, electrical activity from multiple locations within the ventral sensorimotor cortex may be measured. In some embodiments, electrical activity in the high gamma frequency range (e.g., 70 Hz to 150 Hz) or low frequency range (e.g., 0.3 Hz to 100 Hz) may be measured from the precentral gyrus, postcentral gyrus, posterior middle frontal gyrus, posterior superior frontal gyrus, or posterior inferior frontal gyrus regions, or any combination thereof. In some embodiments, electrical activity in the high gamma frequency range (e.g., 70 Hz to 150 Hz) and low frequency range (e.g., 0.3 Hz to 100 Hz) may be measured from the precentral gyrus, postcentral gyrus, posterior middle frontal gyrus, posterior superior frontal gyrus, or posterior inferior frontal gyrus regions, or any combination thereof.
[0127] Detection of brain activity may be performed by any method known in the art. For example, functional brain imaging of neural activity may be performed by electrical methods such as electrocorticography (ECoG), electroencephalography (EEG), stereotactic intracranial electroencephalography (sEEG), magnetoencephalography (MEG), single-photon emission computed tomography (SPECT), as well as metabolic and blood flow studies such as functional magnetic resonance imaging (fMRI), positron emission tomography (PET), functional near-infrared spectroscopy (fNIRS), and time-domain functional near-infrared spectroscopy. In some embodiments, the precentral gyrus, postcentral gyrus, posterior middle frontal gyrus, posterior superior frontal gyrus, or posterior inferior frontal gyrus region is mapped to determine optimal positioning for electrodes for detecting neural activity associated with trial speech and / or trial spelling. One or more of these regions may be implanted with a neural recording device comprising electrodes for measuring electrical signals from neural activity associated with trial speech and / or trial spelling.
[0128] In some cases, electrical activity at one or more locations in the brain may be measured not only during the speech or spelling trial, but also during a period extending from just before the speech or spelling trial (i.e., the speech or spelling preparation period) to just after the speech or spelling trial (i.e., the pause period after the speech or spelling trial). Assessment of the accuracy of speech or spelling decoding from neural activity at a particular site may be determined by comparing the decoded word to the patient's intended word. For example, the patient may use an assistive typing device to communicate the correct intended word. Both the detection of the start and end of speech events and the word / letter classification accuracy from the decoding of neural activity may be assessed. False positives include detected speech events that are not associated with a true word or letter generation attempt, and false negatives include word / letter generation attempts that are not associated with a detected speech event. Lower error rates in detecting speech events and decoding words or spelled letters from neural activity indicate better performance. In some cases, the placement of the electrodes or the number of electrodes may be varied to improve detection of electrical signals and decoding of attempted speech and / or spelling by the subject.
[0129] Application of the present methods may include a preliminary step of selecting patients for implantation of a neurorecording device based on need as determined by a clinical assessment of the severity of the communication disorder and the desire for communication assistance, and may include cognitive, anatomical, behavioral, and / or neurophysiological assessments. Patients with communication difficulties may be implanted with a neurorecording device to assist with communication, as described herein.
[0130] An interface capable of communicating with a computing device is implanted intracranially or placed on the subject's head to provide an externally accessible platform from which brain electrical signals can be acquired from the neurorecording device and transmitted to a data processor for decoding. In some embodiments, the interface comprises a percutaneous pedestal connector secured within the subject's skull. The interface may be connected to a computing device, such as a computer or handheld computing device (e.g., a mobile phone or tablet), by, for example, a detachable digital connector and cable. Alternatively, the interface may be wirelessly connected to the computing device. In some embodiments, the interface comprises a first wireless communication unit that communicates with a computing device comprising a second wireless communication unit. In some embodiments, the first wireless communication unit transfers data from the interface to the computing device comprising the second wireless communication unit using a wireless communication protocol that uses an electromagnetic carrier wave (e.g., a radio wave, microwave, or infrared carrier wave) or ultrasound. Brain-computer interfaces are commercially available, including the Neuroport™ system from Blackrock Microsystems (Salt Lake City, Utah); see also, e.g., Weiss et al. (2019) Brain-Computer Interfaces 6:106-117, incorporated herein by reference.
[0131] The processor may be provided by a computer or handheld computing device (e.g., a mobile phone or tablet) programmed to decode the trial utterances and / or trial spellings from the recorded electrical brain signal data.
[0132] Analyzing the recorded electrical brain activity may include the use of algorithms or classifiers. In some embodiments, machine learning algorithms are used to automate speech detection, character classification (in the case of spelling trials), word classification, and sentence decoding from the analysis of recorded brain activity during speech or spelling trials. Machine learning algorithms may include supervised learning algorithms. Examples of supervised learning algorithms include average-one-dependence estimators (AODEs), artificial neural networks (e.g., artificial neural networks with stacks of long-short-term memory (LSTM) layers), Bayesian statistics (e.g., naive Bayes classifiers, Bayesian networks, Bayesian knowledge bases), case-based reasoning, decision trees, inductive logic programming, Gaussian process regression, group methods of data processing (GMDH), learning automata, learning vector quantization, minimum message length (decision trees, decision graphs, etc.), lazy learning, instance-based learning nearest neighbor algorithms, analogical modeling, probabilistic and approximately correct (PAC) learning, ripple-down rules, knowledge acquisition methodologies, symbolic machine learning algorithms, subsymbolic machine learning algorithms, support vector machines, random forests, ensembles of classifiers, bootstrap aggregating (bagging), and boosting. Supervised learning may also include regression analysis and ordinal classification such as information fuzzy networks (IFNs). Alternatively, supervised learning methods may include statistical classification such as AODE, linear classifiers (e.g., Fisher linear discriminant, logistic regression, naive Bayes classifier, perceptron, and support vector machine), quadratic classifiers, k-nearest neighbors, boosting, decision trees (e.g., C4.5, random forest), Bayesian networks, and hidden Markov models.
[0133] Machine learning algorithms may also include unsupervised learning algorithms. Examples of unsupervised learning algorithms may include artificial neural networks, data clustering, expectation maximization algorithms, self-organizing maps, radial basis function networks, vector quantization, generative terrain maps, information bottleneck methods, and IBSEAD. Unsupervised learning may also include association rule learning algorithms such as the Apriori algorithm, the Eclat algorithm, and the FP-growing algorithm. Hierarchical clustering, such as single-linkage clustering and conceptual clustering, may also be used. Alternatively, unsupervised learning may include divisive clustering, such as the K-means algorithm and fuzzy clustering.
[0134] In some cases, the machine learning algorithm includes a reinforcement learning algorithm. Examples of reinforcement learning algorithms include, but are not limited to, temporal difference learning, Q-learning, and learning automata. Alternatively, the machine learning algorithm may include data preprocessing.
[0135] In some cases, machine learning algorithms may use deep learning (e.g., deep neural networks, deep belief networks, graph neural networks, recurrent neural networks, and convolutional neural networks), which may be supervised, semi-supervised, or unsupervised.
[0136] In some embodiments, machine learning algorithms use artificial neural network (ANN) models for speech detection and word / character classification and natural language processing techniques such as, but not limited to, hidden Markov models (HMMs) or Viterbi decoding models for sentence decoding.
[0137] In some embodiments, the processor is programmed to use a speech detection model to determine the probability that a trial utterance or spelling is occurring at any time during the recording of neural activity and / or to detect the onset and end of a trial utterance or spelling during the recording of neural activity. Linear or nonlinear (e.g., artificial neural network (ANN)) models may be used to automate speech detection. In some embodiments, deep learning models are used for speech detection, specifically to automate the detection of the onset and end of a word production during a trial utterance by the subject or a letter production during a trial spelling by the subject. The processor may also be programmed to assign speech event labels for preparation, speech / spelling, and pauses to time points during the recording of electrical brain signal data. In some embodiments, recorded electrical brain signal data within a time window around the detected onset of a trial utterance / spelling (e.g., from 1 second before the detected onset of speech to 3 seconds after the detected onset of speech) is used for word classification or character classification.
[0138] Word classification can use machine learning algorithms to automate the identification of neural activity patterns of electrical signals in the recorded electrical brain signal data associated with trial word productions during trial speaking by the subject. Letter classification can use machine learning algorithms to automate the identification of neural activity patterns of electrical signals in the recorded electrical brain signal data associated with trial letter productions during trial spelling by the subject.
[0139] In certain embodiments, the subject is provided with a series of go cues instructing them when to begin their spelling trial of each letter of the word in the intended sentence. In some embodiments, the series of go cues are visually presented on a display. Each go cue may be preceded by a countdown to the presentation of the go cue, and a countdown to the next letter to be spelled is visually presented on the display and begins automatically after each go cue. For example, during the spelling procedure, the participant spells out the intended message through a series of letter decoding cycles. During each cycle, the participant is visually presented with a countdown, and finally, the go cue is presented. At the go cue, the participant silently attempts to speak the desired letter. In some embodiments, the series of go cues are provided with a set time interval between each go cue, which may be adjustable by the user. In certain embodiments, the processor is programmed to use electrical brain signal data recorded within a time window following the go cue.
[0140] In some embodiments, the processor is programmed to use a word classification model to decode words within a detected time window of neural activity (e.g., a time window identified by the speech detection model as occurring during the attempted speech or spelling). The word classification model is used to determine the probability that the subject intended a particular word in the attempted utterance across possible speech / text targets. For example, for each word in a vocabulary of possible words the user could utter, the word classification model determines the probability that neural activity was collected when the user attempted to utter that word. The word classification model may use a linear or nonlinear (e.g., ANN) model.
[0141] In some embodiments, the processor is programmed to use the character classification model to determine the probability that the subject intended a particular character during a spelling trial across all possible characters (i.e., alphabetic or numeric characters) of the language used by the subject. In particular embodiments, the processor is further programmed to constrain word classification from character sequences decoded from neural activity associated with the subject's spelling trial of a word to only words within the vocabulary of the language used by the subject.
[0142] In some embodiments, the processor is programmed to use a word sequence decoding model to decode sentences based on word sequence probabilities and determine the most likely word sequence associated with a speech event detected from the subject's corresponding neural activity during a trial utterance or spelling. The word sequence decoding model uses a sequence of probabilities from a classification model to construct a decoded sequence. This may involve using a language model to incorporate a priori character sequence or word sequence probabilities into the neural decoding pipeline. This may also involve a hidden Markov model (HMM) or Viterbi decoding model to process the incorporation of probabilities from the language model. This may use a linear or nonlinear (e.g., ANN) model. In some embodiments, the processor is also programmed to use a language model that provides the probability of a next word given a previous word or phrase in a word sequence to aid decoding by determining the probability of a predicted word sequence, with more frequently occurring words being assigned more weight than less frequently occurring words according to the language model. Additionally, decoded information from previously detected speech events may be used to aid decoding. See the Examples for a detailed discussion of the speech detection model, word classification model, and language model used to decode trial utterances from neural activity.
[0143] The subject can be instructed to limit trial utterances to words from a predefined vocabulary (i.e., word set). The number of words included is preferably large enough to create a variety of meaningful sentences, but small enough to allow sufficient neural-based classification performance. For word classification from neural activity, the subject is instructed to attempt to produce each word included in the word set to determine the pattern of electrical signals associated with each word. Exploratory preliminary assessment of the subject following device implantation can be used to evaluate word selection and word set size that can be easily decoded by the methods described herein and used to assist communication.
[0144] In some embodiments, a word set includes up to 50 words, up to 100 words, up to 200 words, up to 300 words, up to 400 words, or up to 500 words or more. For example, a word set may include 50 words, 55 words, 60 words, 65 words, 70 words, 75 words, 80 words, 85 words, 90 words, 95 words, 100 words, 125 words, 150 words, 175 words, 200 words, 225 words, 250 words, 275 words, 300 words, 325 words, 350 words, 375 words, 400 words, 500 words, 600 words, 700 words, 800 words, 900 words, 1000 words, or any number of words in between. In some embodiments, the word set includes am, are, bad, bring, clean, closer, comfortable, coming, computer, do, faith, family, feel, glasses, going, good, goodbye, have, hello, help, here, hope, how, hungry, I, is, it, like, music, my, need, no, not, nurse, okay, outside, please, right, success, tell, that, they, thirsty, tired, up, very, what, where, yes, and you.
[0145] In some embodiments, the subject's trial utterance may include any chosen sequence of words from the selected word set. In other embodiments, the subject's trial utterance is further restricted to a predefined sentence set that uses only words from the selected word set. The word set and sentence set may be selected to include sentences the subject can use to communicate with the caregiver regarding the task the caregiver wishes the caregiver to perform. For sentence classification from neural activity, the subject is instructed to attempt to generate each sentence included in the sentence set while the subject's neural activity is processed and decoded into text. A processor connected to the interface is programmed to calculate a probability that the word sequence is the intended sentence the subject intended to generate during the trial utterance. In some embodiments, the processor is programmed to calculate a probability that a number of possible sentences composed entirely of words from the specified word set are the intended sentence the subject intended to generate during the trial utterance. In some embodiments, the processor is programmed to retain the sentence composed entirely of words from the specified word set that the subject most likely intended to generate during the trial utterance, as well as other such sentences that are less likely. In some embodiments, the processor is programmed to maintain probabilities of the first, second, and third most likely sentences at any given time. As new word events are processed, the most likely sentence may change. For example, a second most likely sentence based on processing a word event may become the most likely sentence after one or more additional word events are processed.
[0146] In some embodiments, the sentence set includes up to 25 sentences, up to 50 sentences, up to 100 sentences, up to 200 sentences, up to 300 sentences, up to 400 sentences, or up to 500 sentences, or more. For example, the sentence set may include 50 sentences, 100 sentences, 200 sentences, 300 sentences, 400 sentences, 500 sentences, 600 sentences, 700 sentences, 800 sentences, 900 sentences, 1000 sentences, or any number of sentences in between.In some embodiments, the sentence set includes Are you going outside, Are you tired, Bring my glasses here, Bring my glasses please, Do not feel bad, Do you feel comfortable, Faith is good, Hello how are you, Here is my computer, How do you feel, How do you like my music, I am going outside, I am not going, I am not hungry, I am not okay, I am okay, I am outside, I am thirsty, I do not feel comfortable, I feel very comfortable, I feel very hungry, I hope it is clean, I like my nurse, I need my glasses, I need you, It is comfortable, It is good, It is okay, It is right here, My computer is clean, My family is here, My family is outside, My family is very comfortable, My glasses are clean, My glasses are comfortable, My nurse is outside, My nurse is right outside, No, Please bring my glasses here, Please clean it, Please tell my family, That is very clean, They are coming here, They are coming outside, They are going outside, They have faith, What do you do, Where is it, Yes, and You are not right.
[0147] In some embodiments, the subject's trial utterances include spelling out words of the intended message. The trial utterance targets may include the alphabet of any language (e.g., English) and / or code words representing letters of the alphabet (e.g., NATO code words such as alpha, bravo, etc.). Letter probabilities can be determined by classification of the utterance targets (which can use linear or nonlinear (e.g., ANN) models) and processed using sequence decoding techniques (e.g., language modeling, hidden Markov modeling, Viterbi decoding, etc.) to decode complete sentences from brain activity.
[0148] In certain embodiments, the method may further include decoding trial non-speech movements from the recorded neural activity. Non-speech movements may include, but are not limited to, imaginary head, arm, hand, foot, and leg movements. The non-speech movements can be used in any manner beneficial to the user. For example, decoding non-speech movements from the neural activity may be used to control a mouse cursor or otherwise interact with other devices, control error correction methods in a text decoding interface, or select high-level commands for controlling the system (such as "end of sentence" or "return to main menu" commands). A classification model can be used to identify motor commands (e.g., imaginary hand movements), which can be used to indicate to the system that the user is beginning or ending a trial speech or spelling out of the intended message.
[0149] Methods for assisting a subject in communicating through decoding neural activity associated with trial speech, trial spelling of words, or trial non-speech movements can be combined. These techniques are complementary. In some cases, decoding trial spellings may allow a larger vocabulary to be used than decoding trial speech. However, decoding trial utterances may be easier and more convenient for the subject because it allows for faster, more direct word decoding, which may be preferable for expressing frequently used words. To assist with decoding, trial non-speech movements can be used to signal that the subject is beginning or finishing spelling out the trial utterance or intended message.
[0150] Systems and computer-implemented methods for decoding trial speech, trial spelling, and / or trial non-speech movements from brain activity The present disclosure also provides systems that find use in practicing the subject methods. In some embodiments, the system includes: a) a neural recording device including electrodes adapted to be positioned at locations within a sensorimotor cortical region of a subject's brain to record electrical brain signal data associated with attempted speech and / or attempted spelling and / or attempted non-speech movements by the subject; b) a processor programmed to decode sentences from the recorded electrical brain signal data; c) an interface in communication with a computing device adapted to be positioned at locations on the subject's head, the interface receiving the electrical brain signal data from the neural recording device and transmitting the electrical brain signal data to the processor; and d) a display component for displaying the sentences decoded from the recorded electrical brain signal data.
[0151] For example, electrical activity in the high gamma frequency range (e.g., 70 Hz to 150 Hz) and / or low frequency range (e.g., 0.3 Hz to 100 Hz) from the precentral gyrus, postcentral gyrus, posterior middle anterior gyrus, posterior superior anterior gyrus, or posterior inferior anterior gyrus regions, or any combination thereof, can be recorded by a neurorecording device using this system, and the interface receives the brain electrical signal data from the neurorecording device and transmits the brain electrical signal data to a processor. The processor may execute programming for decoding letters, words, phrases, or sentences from the recorded brain electrical signal data using one or more algorithms, as described herein.
[0152] In some embodiments, a computer-implemented method is used to decode sentences from recorded electrical brain signal data associated with trial utterances by a subject. A processor can be programmed to perform the steps of the computer-implemented method, including: a) receiving recorded electrical brain signal data associated with trial utterances by the subject; b) analyzing the recorded electrical brain signal data using a speech detection model to calculate a probability that the trial utterance is occurring at any time and to detect the start and end of word productions during the trial utterances by the subject; c) analyzing the electrical brain signal data using a word classification model to identify patterns of electrical signals in the recorded electrical brain signal data associated with the trial word productions by the subject and to calculate predicted word probabilities; d) performing sentence decoding by using the calculated word probabilities from the word classification model in combination with predicted word sequence probabilities within the sentence using a language model that provides the probability of the next word given a previous word or phrase in the word sequence to calculate predicted word sequence probabilities, and determining the most likely word sequence within the sentence based on the predicted word probabilities determined using the word classification model and the language model; and e) displaying the sentence decoded from the recorded electrical brain signal data.
[0153] In some embodiments, a computer-implemented method is used to decode a sentence from recorded electrical brain signal data associated with a subject's trial spellings of letters of a word of the intended sentence, wherein a processor includes the steps of: a) receiving recorded electrical brain signal data associated with the subject's trial spellings of letters of the word of the intended sentence; b) analyzing the recorded electrical brain signal data using a speech detection model to calculate a probability that the trial spelling is occurring at any point in time and to detect the start and end of letter productions during the subject's trial spelling; c) analyzing the electrical brain signal data using a character classification model to identify patterns of electrical signals in the recorded electrical brain signal data associated with the subject's trial letter productions and to calculate a sequence of predicted letter probabilities; and d) generating a latent character classification model based on the sequence of predicted letter probabilities. The computer-implemented method can be programmed to perform the steps of a computer-implemented method, including: a) calculating potential sentence candidates and automatically inserting spaces into character sequences between predicted words in the sentence candidates, where decoded words in the character sequences are constrained to only be words in the vocabulary of the language used by the subject; b) analyzing the potential sentence candidates using a language model that provides the probability of the next word given the previous word or phrase in the word sequence to calculate a predicted word sequence probability, and determining the most likely sequence of words in the sentence; and c) displaying the sentence decoded from the recorded electrical brain signal data.
[0154] In some embodiments, a computer-implemented method is used to decode sentences from recorded electrical brain signal data associated with speech trials and spelling trials by a subject.
[0155] In certain embodiments, the system can be used not only to decode speech or spelling information from neural activity collected during speech or spelling trials, but also to decode trial non-speech movements from the recorded neural activity. Non-speech movements can include, but are not limited to, imaginary head, arm, hand, foot, and leg movements. The non-speech movements can be used in any manner beneficial to the user. For example, decoding non-speech movements from neural activity can be used to control a mouse cursor or otherwise interact with other devices, control error correction methods within a text decoding interface, or select high-level commands for controlling the system (such as "end of sentence" or "return to main menu" commands). A classification model can be used to identify motor commands (e.g., imaginary hand movements), which can be used to indicate to the system that the user is beginning or ending a speech or spelling trial of the intended message.
[0156] In some embodiments, the computer-implemented method further includes receiving and recording recorded electrical brain signal data associated with a subject's attempted non-speech movement, wherein the subject performs the trial non-speech movement to indicate the start or end of an attempted utterance or trial spelling of a word of an intended sentence or to control an external device; and analyzing the electrical brain signal data using a classification model that identifies patterns of electrical signals in the recorded electrical brain signal data associated with the attempted non-speech movement and calculates a probability that the subject attempted the non-speech movement.
[0157] In certain embodiments, the computer-implemented method further includes storing a user profile of the subject that includes information regarding patterns of electrical signals in the recorded electrical brain signal data that are associated with trial word productions by the subject.
[0158] In some embodiments, an artificial neural network (ANN) model is used for speech detection, and character / word classification and natural language processing techniques, such as, but not limited to, hidden Markov models (HMMs) or Viterbi decoding models, are used for sentence decoding.
[0159] In certain embodiments, the subject is restricted to a specified word set for the trial utterance. In some embodiments, the processor is further programmed to: calculate, for every word in the word set, a probability that the word in the word set is the intended word the subject was trying to produce during the trial utterance; and select the word in the word set that has the highest probability of being the intended word the subject was trying to produce during the trial utterance. In some embodiments, the subject's trial utterance may include any chosen sequence of words from the selected word set. In other embodiments, the subject is restricted to a specified sentence set for the trial utterance.
[0160] In some embodiments, the processor is further programmed to calculate the probability that the word sequence is the intended sentence the subject attempted to produce during the trial utterance. In some embodiments, the processor is programmed to calculate the probability that a number of possible sentences composed entirely of words from the specified word set are the intended sentence the subject attempted to produce during the trial utterance. In some embodiments, the processor is programmed to retain the sentence composed entirely of words from the specified word set that the subject most likely attempted to produce during the trial utterance, as well as one or more such sentences that are less likely. In some embodiments, the processor is programmed to track the probabilities of the first, second, and third most likely sentences at any given time. The most likely sentence may change as new word events are processed. For example, a second most likely sentence based on processing a previous word event may become the most likely sentence after one or more additional word events are processed.
[0161] In certain embodiments, the processor is further programmed to assign event labels for preparation, speech / spelling (complete words, letters, or any other speech target), non-speech movements, and pauses to time points during the recording of the electrical brain signal data. In some embodiments, the processor is further programmed to use electrical brain signal data recorded within a time window around the detected onset of a word or letter classification. For example, the processor may be programmed to use electrical brain signal data recorded from 1 second before to 3 seconds after the detected onset of a word or letter classification.
[0162] In certain embodiments, the processor is further programmed to assign a greater weight to more frequently occurring words than to less frequently occurring words according to the language model.
[0163] The recorded electrical brain signal data may be processed in various ways before decoding. For example, data processing may include, but is not limited to, real-time sample-by-sample processing of neural feature streams, using a common average reference across individual electrode channels, using finite impulse response (FIR) filters to perform digital signal filtering, performing a sliding window normalization procedure using, for example, Welford's method, automatic artifact removal, and parallelization and linear pipelining to improve computational efficiency. The neural feature processing may be performed in real time to extract one or more feature streams for use during speech / text decoding. For descriptions of data processing methods, see, e.g., Moses et al. (2018) J.Neural.Eng.15(3):036005, Moses et al. (2019) Nat.Commun.2019 10(1):3096, Moses et al. (2021) N.Engl.J.Med.385(3):217-227, Sun et al. (2020) J.Neural.Eng.17(6), and Makin et al. (2020) Nature Neuroscience 23:575-582, which are incorporated herein by reference in their entireties.
[0164] The methods described herein can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or control the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or any combination of these.
[0165] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, such as a stand-alone program or including modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subprograms, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communications network.
[0166] In a further aspect, a system for performing the computer-implemented method may include a computer including a processor, a storage component (i.e., memory), a display component, and other components typically present in a general-purpose computer, as described. The storage component stores information accessible by the processor, including instructions that can be executed by the processor and data that can be retrieved, manipulated, or stored by the processor.
[0167] The storage component includes instructions. For example, the storage component includes instructions for a computer-implemented method for decoding sentences from recorded electrical brain signal data associated with speech trial utterances and / or spelling trial utterances by a subject. A computer processor is coupled to the storage component and configured to execute the instructions stored in the storage component to receive the electrical brain signal data associated with speech trial utterances by the subject and analyze the data according to one or more algorithms, as described herein. A display component displays the sentences decoded from the recorded electrical brain signal data.
[0168] The storage component may be of any type capable of storing information accessible by the processor, such as a hard drive, memory card, ROM, RAM, DVD, CD-ROM, USB flash drive, writable memory, and read-only memory. The processor may be any well-known processor, such as a processor from Intel Corporation. Alternatively, the processor may be a dedicated controller, such as an ASIC or FPGA.
[0169] Instructions may be any set of instructions that are executed directly (such as machine code) or indirectly (such as a script) by a processor. In that regard, the terms "instructions," "steps," and "program" may be used interchangeably herein. Instructions may be stored in object code format for direct processing by a processor, or in any other computer language, including a script or collection of independent source code modules that are interpreted on demand or pre-compiled.
[0170] Data may be retrieved, stored, or modified by a processor in accordance with instructions. For example, the system is not limited by any particular data structure, but data may be stored in a computer register, a relational database, as a table with multiple different fields and records, an XML document, or a flat file. Data may also be formatted in any computer-readable format, such as, but not limited to, binary values, ASCII, or Unicode. Furthermore, data may include any information sufficient to identify related information, such as numbers, descriptive text, unique codes, pointers, references to data stored in other memory (including other network locations), or information used by a function to calculate related data.
[0171] In particular embodiments, the processor and storage component may comprise multiple processors and storage components that may or may not be stored within the same physical housing. For example, some of the instructions and data may be stored on a removable CD-ROM and others in a read-only computer chip. Some or all of the instructions and data may be stored in a location that is physically separate from the processor but still accessible by the processor. Similarly, a processor may comprise a collection of processors that may or may not operate in parallel.
[0172] The system also includes an interface capable of communicating with a computing device. The interface can be implanted intracranially or placed on the subject's head to provide an externally accessible platform from which brain electrical signals can be acquired from the neurorecording device and transmitted to the computing device for decoding. In some embodiments, the interface includes a percutaneous pedestal connector secured within the subject's skull. The interface can be connected to a computing device, such as a computer or handheld computing device (e.g., a mobile phone or tablet), by, for example, a detachable digital connector and cable. Alternatively, the interface can be wirelessly connected to the computing device. In some embodiments, the interface includes a first wireless communication unit that communicates with the computing device including the second wireless communication unit. In some embodiments, the first wireless communication unit transfers data from the interface to the computing device including the second wireless communication unit using a wireless communication protocol that uses an electromagnetic carrier wave (e.g., a radio wave, microwave, or infrared carrier wave) or ultrasound. Brain-computer interfaces are commercially available, including the Neuroport™ system from Blackrock Microsystems (Salt Lake City, UT); see also, e.g., Weiss et al. (2019) Brain-Computer Interfaces 6:106-117, incorporated herein by reference.
[0173] Components of systems for carrying out the methods of the present disclosure are further described in the Examples below.
[0174] kit Kits for carrying out the methods described herein are also provided. In some embodiments, the kit comprises software for carrying out a computer-implemented method for decoding sentences from recorded electrical brain signal data associated with speech trial and / or spelling trial by a subject, as described herein. In some embodiments, the kit comprises a system for assisting a subject's communication, as described herein. Such a system can comprise: a neural recording device comprising electrodes adapted to be positioned at locations within a sensorimotor cortical region of the subject to record electrical brain signal data associated with speech trial and / or spelling trial and / or non-speech movement by the subject; a processor programmed to decode sentences from the recorded electrical brain signal data in accordance with the computer-implemented methods described herein; an interface capable of communicating with the computing device, the interface adapted to be positioned at locations on the subject's head, the interface receiving the electrical brain signal data from the neural recording device and transmitting the electrical brain signal data to the processor; and a display component for displaying the sentences decoded from the recorded electrical brain signal data.
[0175] In addition, the kits may (in certain embodiments) further include instructions for practicing the subject methods. These instructions may be present in the subject kits in a variety of forms, one or more of which may be present in the kit. For example, the instructions may be present as printed information on a suitable medium or substrate, such as one or more sheets of paper having the information printed thereon, in the kit packaging, in a package insert, etc. Another form of these instructions is a computer-readable medium having the information recorded thereon, such as a diskette, compact disc (CD), flash drive, etc. Yet another form of these instructions may be present is a website address that may be used via the Internet to access the information at the removed site.
[0176] Utilities The disclosed methods, devices, and systems find application in assisting individuals with communication. Specifically, methods, devices, and systems are provided for decoding words and sentences directly from an individual's neural activity. In the disclosed methods, cortical activity from brain regions involved in speech processing is recorded while the individual attempts to speak or spell out words of an intended sentence. A deep learning computational model is used to detect and classify letters / words from the recorded brain activity. The decoding of speech from brain activity is aided by the use of a language model that predicts how a particular word sequence is likely to appear. Additionally, decoding of trial non-speech movements from neural activity can be used to further assist communication.
[0177] The methods, devices, and systems disclosed herein can be used to assist individuals with communication difficulties caused by conditions and diseases including, but not limited to, dysarthria, stroke, traumatic brain injury, brain tumor, amyotrophic lateral sclerosis, multiple sclerosis, Huntington's disease, Niemann-Pick disease, Friedreich's ataxia, Wilson's disease, cerebral palsy, Guillain-Barré syndrome, Tay-Sachs disease, encephalopathy, central pontine myelinolysis, and other conditions that cause dysfunction or paralysis of the muscles of the head, neck, or chest resulting in dysarthria. The methods disclosed herein can be used to restore communication and improve autonomy and quality of life for such individuals.
[0178] Examples of Non-Limiting Aspects of the Disclosure Aspects, including embodiments of the present subject matter described above, may be useful alone or in combination with one or more other aspects or embodiments. Without limiting the above description, certain non-limiting aspects of the present disclosure, numbered 1 through 159, are provided below. As will be apparent to one of skill in the art upon reading this disclosure, each individually numbered aspect may be used or combined with any of the preceding or following individually numbered aspects. This is intended to provide support for all such combinations of aspects, and is not limited to the combinations of aspects explicitly provided below. 1. A method for assisting a subject in communication, the method comprising: positioning a neural recording device comprising electrodes at a location within a sensorimotor cortical region of the subject's brain to record brain electrical signal data associated with a trial speech utterance by the subject; positioning an interface in communication with a computing device at a location on the subject's head, the interface being connected to a neural recording device; recording electrical brain signal data associated with trial speech by the subject using a neural recording device, wherein an interface receives the electrical brain signal data from the neural recording device and transmits the electrical brain signal data to a processor of a computing device; and using a processor to decode words, phrases, or sentences from the recorded electrical brain signal data. 2. The method of aspect 1, wherein the subject is having difficulty communicating due to dysarthria, stroke, traumatic brain injury, brain tumor, or amyotrophic lateral sclerosis. 3. The method of aspect 1 or 2, wherein the subject is paralyzed. 4. The method of any one of aspects 1-3, wherein the location of the neurorecording device is within the ventral sensorimotor cortex. 5. The method of any one of aspects 1-4, wherein the electrodes are positioned on the surface of or within the sensorimotor cortical area. 6. The method of embodiment 5, wherein the electrodes are positioned on the surface of the sensorimotor cortical region of the brain within the subdural space. 7. The method of any one of aspects 1-6, wherein the neural recording device comprises a brain-penetrating electrode array. 8. The method of any one of claims 1 to 7, wherein the neurorecording device comprises an electrocorticography (ECoG) electrode array. 9. The method of any one of aspects 1-8, wherein the electrode is a deep electrode or a surface electrode. 10. The method of any one of aspects 1-9, wherein the electrical signal data includes high gamma frequency components characteristic. 11. The method of aspect 10, wherein the electrical signal data comprises neural oscillations in the range of 70 Hz to 150 Hz. 12. The method of any one of aspects 1-11, wherein recording the electrical brain signal data comprises recording electrical brain signal data from a sensorimotor cortical region selected from the precentral gyrus, postcentral gyrus, posterior middle frontal gyrus, posterior superior frontal gyrus, or posterior inferior frontal gyrus region, or any combination thereof. 13. The method of any one of aspects 1-12, further comprising mapping the subject's brain to identify optimal locations for positioning electrodes to record brain electrical signals associated with trial speech by the subject. 14. The method of any one of aspects 1-13, wherein the interface comprises a percutaneous pedestal connector attached to the subject's skull. 15. The method of aspect 14, wherein the interface further comprises a headstage connected to the percutaneous pedestal connector. 16. The method of any one of aspects 1-15, wherein the processor is provided by a computer or a handheld device. 17. The method of aspect 16, wherein the handheld device is a mobile phone or a tablet. 18. The method of any one of aspects 1-17, wherein the processor is programmed to automate speech detection, word classification, and sentence decoding based on identifying neural activity patterns of electrical signals in the recorded electrical brain signal data that are associated with trial word productions. 19. The method of aspect 18, wherein the processor is programmed to use machine learning algorithms for speech detection, word classification, and sentence decoding. 20. The method of aspect 19, wherein an artificial neural network (ANN) model is used for speech detection and word classification, and a hidden Markov model (HMM), a Viterbi decoding model, or a natural language processing technique is used for sentence decoding. 21. The method of any one of aspects 1-20, wherein the processor is programmed to automate detection of the start and end of word production during a trial utterance by the subject. 22. The method of aspect 21, further comprising assigning speech event labels for preparation, speech, and pause to time points during recording of the electrical brain signal data. 23. The method of aspect 21 or 22, wherein the processor is programmed to use electrical brain signal data recorded within a time window around the detected onset of the word classification. 24. The method of any one of aspects 1-23, wherein the subject is restricted to a specified set of words for the trial utterance. 25. The method of aspect 24, wherein the processor is programmed to calculate a probability that a word in the word set is an intended word that the subject attempted to produce during the trial utterance. 26. The method of aspect 25, wherein the processor is programmed to calculate, for every word in the word set, a probability that the word in the word set is the intended word that the subject attempted to produce during the trial utterance. 27. The method of any one of aspects 24-26, wherein the word set includes am, are, bad, bring, clean, closer, comfortable, coming, computer, do, faith, family, feel, glasses, going, good, goodbye, have, hello, help, here, hope, how, hungry, I, is, it, like, music, my, need, no, not, nurse, okay, outside, please, right, success, tell, that, they, thirsty, tired, up, very, what, where, yes, and you. 28. The method of any one of aspects 1-27, wherein the subject can use words from the word set without restriction to create sentences. 29. The method of embodiment 28, wherein the processor is programmed to calculate a probability that the word sequence is an intended sentence that the subject attempted to produce during the trial utterance. 30. The method of any one of aspects 1-29, wherein the processor is programmed to use a language model that provides the probability of a next word given a previous word or phrase in a word sequence to aid decoding by determining predicted word sequence probabilities. 31. The method of aspect 30, wherein more frequently occurring words are assigned a greater weight than less frequently occurring words according to a language model. 32. The method of aspect 30 or 31, wherein the processor is programmed to use a Viterbi decoding model to determine the most likely word sequence within the subject's intended utterance, given electrical brain signal data associated with the trial utterance, predicted word probabilities from a word classification model using a machine learning algorithm, and word sequence probabilities using a language model. 33. Recording brain electrical signal data associated with a subject's trial non-speech movement, wherein the subject performs a trial non-speech movement to indicate the start or end of a trial speech utterance or to control an external device; 33. The method of any one of aspects 1-32, further comprising: analyzing the brain electrical signal data using a non-speech movement classification model that identifies patterns of electrical signals in the recorded brain electrical signal data that are associated with attempted non-speech movements and calculates a probability that the subject attempted the non-speech movement. 34. The method of aspect 33, wherein the attempted non-speech movement comprises an attempted movement of the head, arm, hand, foot, or leg. 35. The method of aspect 34, wherein the trial hand movement comprises an imaginary hand gesture or an imaginary hand grasp. 36. The method of any one of aspects 33-35, wherein the processor is further programmed to automate detection of an attempted non-speech movement of the subject signaling an end of the attempted speech by the subject based on identifying a neural activity pattern of electrical signals within the recorded electrical brain signal data that is associated with the attempted non-speech movement. 37. The method of embodiment 36, wherein the processor is further programmed to assign a trial non-speech movement event label to a time point during recording of the electrical brain signal data. 38. The method of any one of aspects 1-37, wherein the method further comprises evaluating the accuracy of the decoding. 39. A computer-implemented method for decoding sentences from recorded electrical brain signal data associated with trial utterances by a subject, the method comprising: a) receiving recorded electrical brain signal data associated with a speech trial by a subject; b) analyzing the recorded electrical brain signal data using a speech detection model to calculate the probability that a trial utterance is occurring at any time during the recording of the electrical brain signal data and to detect the start and end of word production during the trial utterance by the subject; c) analyzing the electrical brain signal data using a word classification model that identifies electrical signal patterns in the recorded electrical brain signal data associated with trial word productions by the subject and calculates predicted word probabilities; d) performing sentence decoding by using the calculated word probabilities from the word classification model in combination with predicted word sequence probabilities within the sentence using a language model that provides the probability of the next word given a previous word or phrase in the word sequence to calculate predicted word sequence probabilities, and determining the most likely word sequence within the sentence based on the predicted word probabilities determined using the word classification model and the language model; e) displaying the sentences decoded from the recorded electrical brain signal data. 40. The computer-implemented method of embodiment 39, wherein machine learning algorithms are used for speech detection, word classification, and sentence decoding. 41. The computer-implemented method of aspect 40, wherein an artificial neural network (ANN) model is used for speech detection and word classification, and a hidden Markov model (HMM), a Viterbi decoding model, or a natural language processing technique is used for sentence decoding. 42. The computer-implemented method of any one of aspects 39-41, wherein the subject is restricted to a specified set of words for the trial utterance. 43. The computer-implemented method of aspect 42, further comprising: calculating, for every word in the word set, a probability that the word in the word set is the intended word that the subject was attempting to produce during the trial utterance; and selecting the word in the word set that has the highest probability of being the intended word that the subject was attempting to produce during the trial utterance. 44. The computer-implemented method of any one of aspects 39-43, wherein the subject may use words from the word set without restriction to create sentences or is restricted to a specified sentence set for trial utterances. 45. The computer-implemented method of any one of aspects 39-44, further comprising calculating a probability that the word sequence is the intended sentence that the subject attempted to produce during the trial utterance. 46. The computer-implemented method of aspect 45, further comprising: maintaining a most likely sentence and one or more less likely sentences; and, after decoding each word, recalculating the probability that the word sequence is the intended sentence that the subject attempted to produce during the trial utterance. 47. The computer-implemented method of aspect 46, wherein the most likely sentence and the one or more less likely sentences are composed only of words from a word set used by the subject for the trial utterance. 48. The computer-implemented method of any one of aspects 39-47, further comprising assigning speech event labels for preparation, speech, and pause to time points during recording of the electrical brain signal data. 49. The computer-implemented method of aspect 48, wherein only electrical brain signal data recorded within a time window around the detected onset of a word classification is used. 50. The computer-implemented method of any one of aspects 39-49, wherein more frequently occurring words are assigned a higher weight than less frequently occurring words according to a language model. 51. The computer-implemented method of any one of aspects 39-50, further comprising storing a user profile of the subject including information regarding patterns of electrical signals in the recorded electrical brain signal data associated with trial word productions by the subject. 52. Receiving recorded electrical brain signal data associated with a subject's trial non-speech movement, wherein the subject performs a trial non-speech movement to indicate the start or end of a trial speech utterance or to control an external device; 52. The computer-implemented method of any one of aspects 39 to 51, further comprising: analyzing the brain electrical signal data using a classification model that identifies patterns of electrical signals in the recorded brain electrical signal data that are associated with attempted non-speech movements and calculates a probability that the subject attempted the non-speech movement. 53. The computer-implemented method of aspect 52, wherein the attempted non-speech movement includes an attempted movement of the head, arm, hand, foot, or leg. 54. The computer-implemented method of aspect 53, wherein the trial hand movement includes an imaginary hand gesture or an imaginary hand grasp. 55. The computer-implemented method of any one of aspects 52-54, further comprising assigning an event label of the trial non-speech movement to a time point during recording of the electrical brain signal data. 56. A non-transitory computer-readable medium comprising program instructions that, when executed by a processor in a computer, cause the processor to perform a method according to any one of aspects 39-55. 57. A kit comprising the non-transitory computer-readable medium of aspect 56 and instructions for decoding brain electrical signal data associated with trial utterances by a subject. 58. A system for supporting communication between subjects, the system comprising: a neural recording device comprising electrodes adapted to be positioned at locations within a sensorimotor cortical region of the subject's brain to record brain electrical signal data associated with attempted speech or attempted non-speech movements by the subject; A processor programmed to decode sentences from recorded electrical brain signal data according to the computer-implemented method of any one of aspects 39 to 55; an interface in communication with the computing device, the interface adapted to be positioned at a location on the subject's head, the interface receiving the brain electrical signal data from the neural recording device and transmitting the brain electrical signal data to the processor; a display component for displaying sentences decoded from the recorded electrical brain signal data. 59. The system of aspect 58, wherein the subject has difficulty communicating due to dysarthria, stroke, traumatic brain injury, brain tumor, or amyotrophic lateral sclerosis. 60. The system of aspect 58 or 59, wherein the location of the neural recording device is within the ventral sensorimotor cortex. 61. The system of any one of aspects 58-60, wherein the electrodes are adapted to be positioned on the surface of or within the sensorimotor cortical region. 62. The system of aspect 61, wherein the electrode is adapted to be positioned on the surface of a sensorimotor cortical region of the brain within the subdural space. 63. The system of any one of aspects 58-62, wherein the neural recording device comprises a brain-penetrating electrode array. 64. The system of any one of aspects 58-63, wherein the neurorecording device includes an electrocorticography (ECoG) electrode array. 65. The system of any one of aspects 58-64, wherein the electrode is a deep electrode or a surface electrode. 66. The system of any one of aspects 58-65, wherein the electrical signal data includes high gamma frequency components characteristic. 67. The system of aspect 66, wherein the electrical signal data comprises neural oscillations in the range of 70 Hz to 150 Hz. 68. The system of any one of aspects 58-67, wherein the interface comprises a percutaneous pedestal connector attached to the subject's skull. 69. The system of embodiment 68, wherein the interface further comprises a headstage connectable to the percutaneous pedestal connector. 70. The system of any one of aspects 58-69, wherein the processor is provided by a computer or a handheld device. 71. The system of aspect 70, wherein the handheld device is a mobile phone or a tablet. 72. The system of any one of aspects 58-71, wherein machine learning algorithms are used for speech detection, word classification, and sentence decoding. 73. The system of aspect 72, wherein an artificial neural network (ANN) model is used for speech detection and word classification, and a hidden Markov model (HMM), a Viterbi decoding model, or a natural language processing technique is used for sentence decoding. 74. The system of any one of aspects 58-73, wherein the processor is further programmed to assign speech event labels for preparation, speech, and pauses to time points during recording of the electrical brain signal data. 75. The system of embodiment 74, wherein the processor is further programmed to use electrical brain signal data recorded within a time window around the detected onset of the word classification. 76. The system of any one of aspects 58-75, wherein the subject is restricted to a specified set of words for the trial utterance. 77. The system of aspect 76, wherein the processor is further programmed to: calculate, for every word in the word set, a probability that the word in the word set is the intended word that the subject was attempting to produce during the trial utterance; and select the word in the word set that has the highest probability of being the intended word that the subject was attempting to produce during the trial utterance. 78. The system of aspect 76 or 77, wherein the word set includes am, are, bad, bring, clean, closer, comfortable, coming, computer, do, faith, family, feel, glasses, going, good, goodbye, have, hello, help, here, hope, how, hungry, I, is, it, like, music, my, need, no, not, nurse, okay, outside, please, right, success, tell, that, they, thirsty, tired, up, very, what, where, yes, and you. 79. The system of any one of aspects 76-78, wherein the subject can use any chosen sequence of words from the selected word set. 80. The system of aspect 79, wherein the processor is programmed to calculate a probability that the word sequence is an intended sentence that the subject attempted to produce during the trial utterance. 81. The system of aspect 80, wherein the processor is programmed to maintain a most likely sentence and one or more less likely sentences, and after decoding each word, recalculate the probability that the word sequence is the intended sentence that the subject attempted to produce during the trial utterance. 82. The system of aspect 81, wherein the most likely sentence and the one or more less likely sentences are composed only of words from a word set used by the subject for the trial utterance. 83. The system of any one of aspects 58 to 82, wherein the processor is further programmed to automate detection of the subject's attempted non-speech movements signaling the start or end of an attempted speech by the subject based on identifying neural activity patterns of electrical signals within the recorded brain electrical signal data that are associated with the attempted non-speech movements. 84. The system of aspect 83, wherein the processor is further programmed to assign an event label of the trial non-speech movement to a time point during recording of the electrical brain signal data. 85. A kit comprising a system described in any one of aspects 58 to 84 and instructions for using the system to record and decode brain electrical signal data associated with trial utterances by a subject. 86. A method for assisting a subject in communication, the method comprising: positioning a neural recording device comprising electrodes at locations within a sensorimotor cortical region of the subject's brain to record brain electrical signal data associated with the subject's attempted spelling of letters of a word of an intended sentence; positioning an interface in communication with a computing device at a location on the subject's head, the interface being connected to a neural recording device; recording electrical brain signal data associated with the spelling attempt by the subject using a neural recording device, wherein an interface receives the electrical brain signal data from the neural recording device and transmits the electrical brain signal data to a processor of a computing device; and decoding, using a processor, the spelled words of the intended sentence from the recorded electrical brain signal data. 87. The method of aspect 86, wherein the subject has difficulty communicating due to dysarthria, stroke, traumatic brain injury, brain tumor, or amyotrophic lateral sclerosis. 88. The method of aspect 86 or 87, wherein the subject is paralyzed. 89. The method of any one of aspects 86-88, wherein the location of the neural recording device is within the ventral sensorimotor cortex. 90. The method of any one of aspects 86-89, wherein the electrode is positioned on the surface of or within the sensorimotor cortical area. 91. The method of aspect 90, wherein the electrodes are positioned on the surface of the sensorimotor cortical region of the brain within the subdural space. 92. The method of any one of aspects 86-91, wherein the neural recording device comprises a brain-penetrating electrode array. 93. The method of any one of aspects 86-92, wherein the neurorecording device comprises an electrocorticography (ECoG) electrode array. 94. The method of any one of aspects 86-93, wherein the electrode is a deep electrode or a surface electrode. 95. The method of any one of aspects 86-94, wherein the electrical signal data includes high gamma frequency component features and low frequency component features. 96. The method of aspect 95, wherein the electrical signal data comprises neural oscillations in the high gamma frequency range of 70 Hz to 150 Hz and the low frequency range of 0.3 Hz to 100 Hz. 97. The method of any one of aspects 86-96, wherein recording the electrical brain signal data comprises recording electrical brain signal data from a sensorimotor cortical region selected from the precentral region, the postcentral region, the posterior middle frontal gyrus region, the posterior superior frontal gyrus region, or the posterior inferior frontal gyrus region, or any combination thereof. 98. The method of any one of aspects 86-97, further comprising mapping the subject's brain to identify optimal locations for positioning electrodes to record brain electrical signals associated with the subject's attempted spelling of a word or attempted non-speech movement speech. 99. The method of any one of aspects 86-98, wherein the interface comprises a percutaneous pedestal connector attached to the subject's skull. 100. The method of embodiment 99, wherein the interface further comprises a headstage connected to the percutaneous pedestal connector. 101. The method of any one of aspects 86-100, wherein the processor is provided by a computer or a handheld device. 102. The method of aspect 101, wherein the handheld device is a mobile phone or a tablet. 103. The method of any one of aspects 86-102, wherein the processor is programmed to automate the detection of the spelling attempts, letter classification, word classification, and sentence decoding based on identifying neural activity patterns of electrical signals in the recorded electrical brain signal data that are associated with the subject's spelling attempts of the words. 104. The method of aspect 103, wherein the processor is programmed to use machine learning algorithms for speech detection, character classification, word classification, and sentence decoding. 105. The method of aspect 104, wherein the processor is further programmed to constrain word classification from character sequences decoded from neural activity associated with attempted spellings of words by the subject to only words within the vocabulary of the language used by the subject. 106. The method of any one of aspects 86-105, wherein the processor is further programmed to assign speech event labels for preparation, attempt spelling, and pause to time points during recording of the electrical brain signal data. 107. The method of aspect 106, wherein the processor is programmed to use electrical brain signal data recorded within a time window around the detected onset of the subject's attempted spelling of the letter. 108. The method of any one of aspects 86-107, further comprising providing the subject with a series of go cues instructing the subject when to begin trial spelling of each letter of the word of the intended sentence. 109. The method of aspect 108, wherein the series of go cues are visually provided on a display. 110. The method of embodiment 109, wherein each go cue is preceded by a countdown to the presentation of the go cue, and a countdown of the next letter to be spelled is provided visually on the display and begins automatically after each go cue. 111. The method of any one of aspects 108-110, wherein a series of go cues are provided with set time intervals between each go cue. 112. The method of aspect 111, wherein the subject can control the set time interval between each go cue. 113. The method of any one of aspects 108-112, wherein the processor is programmed to use electrical brain signal data recorded within a time window following a go cue. 114. The method of any one of aspects 86-113, wherein the processor is programmed to calculate a probability that a decoded word sequence from the decoded character sequence is the intended sentence that the subject attempted to produce during the subject's trial spelling of the characters of the words of the intended sentence. 115. The method of any one of aspects 86-114, wherein the processor is programmed to use a language model that provides the probability of a next word given a previous word or phrase in a word sequence to aid decoding by determining predicted word sequence probabilities. 116. The method of aspect 115, wherein more frequently occurring words are assigned a greater weight than less frequently occurring words according to the language model. 117. The method of any one of aspects 86-116, wherein the processor is further programmed to use the sequence of predicted character probabilities to calculate potential sentence candidates and automatically insert spaces in the character sequence between predicted words in the sentence candidates. 118. Recording brain electrical signal data associated with a subject's trial non-speech movements, wherein the subject performs trial non-speech movements to indicate the start or end of a trial spelling of a word of an intended sentence or to control an external device; A computer-implemented method described in any one of aspects 86 to 117, further comprising: analyzing the brain electrical signal data using a classification model that identifies patterns of electrical signals in the recorded brain electrical signal data that are associated with attempted non-speech movements and calculates a probability that the subject attempted the non-speech movement. 119. The method of aspect 118, wherein the attempted non-speech movement includes an attempted movement of the head, arm, hand, foot, or leg. 120. The method of aspect 119, wherein the trial hand movement comprises an imaginary hand gesture or an imaginary hand grasp. 121. The computer-implemented method of any one of aspects 118-120, further comprising assigning an event label of the trial non-speech movement to a time point during recording of the electrical brain signal data. 122. The method of any one of aspects 86-121, further comprising evaluating the accuracy of the decoding. 123. Recording electrical brain signal data associated with trial speech by the subject using a neural recording device, wherein an interface receives the electrical brain signal data from the neural recording device and transmits the electrical brain signal data to a processor of a computing device; The method of any one of aspects 86 to 122, further comprising: using a processor to decode words, phrases, or sentences from the recorded electrical brain signal data associated with the trial utterances by the subject. 124. A computer-implemented method for decoding a sentence from recorded electrical brain signal data associated with attempted spellings of letters of words of an intended sentence by a subject, the method comprising: a) receiving recorded electrical brain signal data associated with a subject's attempted spelling of letters of a word of an intended sentence; b) analyzing the recorded electrical brain signal data using a speech detection model to calculate the probability that a spelling trial is occurring at any time during the recording of the electrical signal data and to detect the start and end of letter production during the subject's spelling trial; c) analyzing the electrical brain signal data using a character classification model that identifies electrical signal patterns in the recorded electrical brain signal data associated with trial letter productions by the subject and calculates a series of predicted character probabilities; d) calculating potential sentence candidates based on the sequence of predicted character probabilities and automatically inserting spaces into the character sequence between predicted words in the sentence candidate, where decoded words in the character sequence are constrained to only be words in the vocabulary of the language used by the subject; e) analyzing potential sentence candidates using a language model that provides the probability of a next word given a previous word or phrase in a word sequence to calculate a predicted word sequence probability, and determining the most likely word sequence in the sentence; f) displaying the sentences decoded from the recorded electrical brain signal data. 125. The computer-implemented method of aspect 124, wherein the recorded electrical brain signal data is used only within a time window around the detected onset of the subject's attempted spelling of the letter. 126. The computer-implemented method of aspect 124 or 125, further comprising displaying to the subject a series of go cues instructing the subject when to begin a spelling trial of each letter of the word of the intended sentence. 127. The computer-implemented method of aspect 126, wherein each go cue is preceded by a countdown to the presentation of the go cue being displayed, and wherein a countdown to the next letter to be spelled is automatically initiated after each go cue. 128. The computer-implemented method of aspect 126 or 127, wherein a series of go cues are provided with a set time interval between each go cue. 129. The computer-implemented method of aspect 128, wherein the subject can control the set time interval between each go cue. 130. The computer-implemented method of any one of aspects 122-127, wherein electrical brain signal data recorded within a time window following a go cue is used for character classification. 131. Receiving recorded electrical brain signal data associated with a subject's trial non-speech movement, wherein the subject performs a trial non-speech movement to indicate the start or end of a trial spelling of a word of an intended sentence or to control an external device; A computer-implemented method according to any one of aspects 124 to 130, further comprising: analyzing the brain electrical signal data using a classification model that identifies patterns of electrical signals in the recorded brain electrical signal data that are associated with attempted non-speech movements and calculates a probability that the subject attempted the non-speech movement. 132. The method of aspect 131, wherein the attempted non-speech movement includes an attempted movement of the head, arm, hand, foot, or leg. 133. The method of aspect 132, wherein the trial hand movement comprises an imaginary hand gesture or an imaginary hand grasp. 134. The computer-implemented method of any one of aspects 124-133, wherein a machine learning algorithm is used for detection of spelling attempts or non-speech movements attempts or character classification. 135. The computer-implemented method of any one of aspects 124-134, further comprising assigning a higher weight to more frequently occurring words than to less frequently occurring words according to a language model. 136. The computer-implemented method of any one of aspects 124-135, further comprising storing a user profile of the subject, the user profile including information regarding patterns of electrical signals in the recorded electrical brain signal data associated with letter productions during spelling attempts by the subject. 137. The computer-implemented method of any one of aspects 124-136, wherein the electrical signal data includes high gamma frequency component features and low frequency component features. 138. The computer-implemented method of aspect 137, wherein the electrical signal data comprises neural oscillations in a high gamma frequency range of 70 Hz to 150 Hz and a low frequency range of 0.3 Hz to 100 Hz. 139. The computer-implemented method of any one of aspects 124-138, further comprising evaluating the accuracy of the decoding. 140. The method further comprising decoding sentences from recorded electrical brain signal data associated with trial utterances by the subject, wherein the computer: a) receiving recorded electrical brain signal data associated with a speech trial by a subject; b) analyzing the recorded electrical brain signal data using a speech detection model to calculate the probability that a speech trial is occurring at any point in time and to detect the start and end of word production during the speech trial by the subject; c) analyzing the electrical brain signal data using a word classification model that identifies electrical signal patterns in the recorded electrical brain signal data associated with trial word productions by the subject and calculates predicted word probabilities; d) performing sentence decoding by using the calculated word probabilities from the word classification model in combination with predicted word sequence probabilities within the sentence using a language model that provides the probability of the next word given a previous word or phrase in the word sequence to calculate predicted word sequence probabilities, and determining the most likely word sequence within the sentence based on the predicted word probabilities determined using the word classification model and the language model; e) displaying the sentences decoded from the recorded electrical brain signal data. 141. The computer-implemented method of aspect 140, wherein machine learning algorithms are used for speech detection and word classification, and sentence decoding. 142. The computer-implemented method of aspect 141, wherein an artificial neural network (ANN) model is used for speech detection and word classification, and a hidden Markov model (HMM), a Viterbi decoding model, or a natural language processing technique is used for sentence decoding. 143. A non-transitory computer-readable medium comprising program instructions that, when executed by a processor in a computer, cause the processor to perform a method according to any one of aspects 124-142. 144. A kit comprising the non-transitory computer-readable medium of aspect 143 and instructions for decoding electrical brain signal data associated with attempted spellings of letters of words in an intended sentence by a subject. 145. A system for supporting communication between subjects, the system comprising: a neural recording device comprising electrodes adapted to be positioned at locations within a sensorimotor cortical region of the subject's brain to record brain electrical signal data associated with attempted speech, attempted spelling of letters of a word of an intended sentence, or attempted non-speech movements by the subject, or a combination thereof; A processor programmed to decode sentences from recorded electrical brain signal data according to the computer-implemented method of any one of aspects 124 to 142; an interface in communication with the computing device, the interface adapted to be positioned at a location on the subject's head, the interface receiving the brain electrical signal data from the neural recording device and transmitting the brain electrical signal data to the processor; a display component for displaying sentences decoded from the recorded electrical brain signal data. 146. The system of aspect 145, wherein the subject has difficulty communicating due to dysarthria, stroke, traumatic brain injury, brain tumor, or amyotrophic lateral sclerosis. 147. The system of aspect 145 or 146, wherein the location of the neural recording device is within the ventral sensorimotor cortex. 148. The system of any one of aspects 145-147, wherein the electrodes are adapted to be positioned on the surface of or within the sensorimotor cortical region. 149. The system of aspect 148, wherein the electrode is adapted to be positioned on the surface of a sensorimotor cortical region of the brain within the subdural space. 150. The system of any one of aspects 145-149, wherein the neural recording device comprises a brain-penetrating electrode array. 151. The system of any one of aspects 145-150, wherein the neurorecording device includes an electrocorticography (ECoG) electrode array. 152. The system of any one of aspects 145-151, wherein the electrode is a deep electrode or a surface electrode. 153. The system of any one of aspects 145-152, wherein the electrical signal data includes high gamma frequency component features and low frequency component features. 154. The system of aspect 153, wherein the electrical signal data comprises neural oscillations in the high gamma frequency range of 70 Hz to 150 Hz and the low frequency range of 0.3 Hz to 100 Hz. 155. The system of any one of aspects 145-154, wherein the interface comprises a percutaneous pedestal connector attached to the subject's skull. 156. The system of aspect 155, wherein the interface further comprises a headstage connectable to the percutaneous pedestal connector. 157. The system of any one of aspects 145-156, wherein the processor is provided by a computer or a handheld device. 158. The system of aspect 157, wherein the handheld device is a mobile phone or a tablet. 159. A kit comprising the system of any one of aspects 145-158 and instructions for using the system to record and decode electrical brain signal data associated with speech trials, word spelling trials, or non-speech trial movements by a subject, or a combination thereof. [Example]
[0179] As can be understood from the disclosure provided above, the present disclosure has a wide variety of applications. Thus, the following examples are presented to provide those of ordinary skill in the art with a complete disclosure and description of how to make and use the present invention, and are not intended to limit the scope of what the inventors regard as their invention, nor are they intended to represent that the following experiments are all or the only experiments performed. Efforts have been made to ensure accuracy with respect to numbers used (e.g., amounts, dimensions, etc.), but some experimental error and deviation should be accounted for. Those of ordinary skill in the art will readily recognize a variety of non-critical parameters that can be changed or modified to yield essentially similar results.
[0180] Example 1: Speech neuroprosthesis for decoding words in humans with severe paralysis Introduction Dysarthria is the loss of the ability to speak. Dysarthria can result from a variety of conditions, including stroke, traumatic brain injury, and amyotrophic lateral sclerosis.[1] For paralyzed individuals with severe motor disabilities, dysarthria can interfere with communication with family, friends, and caregivers, reducing self-reported quality of life.[2]
[0181] Advances are being made in typing-based brain-computer interfaces that allow speech-impaired individuals to spell out intended messages using cursor control [3-7]. However, letter-by-letter selection interfaces driven by neural signal recordings can be relatively slow and tedious. A more efficient and natural approach may be to directly decode entire words from brain regions that control speech. Over the past decade, our understanding of how the speech motor cortex coordinates rapid articulation of the vocal tract has expanded [8-13]. In parallel, technological efforts have leveraged these discoveries to demonstrate that speech can be decoded from brain activity in people without speech disorders [14-17].
[0182] However, it is unclear whether speech decoding approaches work in paralyzed individuals who are unable to speak. Neural activity cannot be accurately aligned with intended speech due to the lack of speech output, presenting an obstacle to training computational models
[18] . Additionally, it is unclear whether the neural signals underlying speech control remain normal in individuals who have not spoken for years or even decades. In previous studies, humans with locked-in syndrome produced vowels and phonemes through an audiovisual interface using an implanted two-channel microelectrode device [19, 20]. However, it remains unclear whether it is possible to reliably decode entire words from neural activity in humans with dysarthria.
[0183] In this study, we demonstrate real-time word and sentence decoding from neural activity in a human suffering from severe paralysis and dysarthria resulting from a remote brainstem stroke (Figure 1). Our findings represent proof-of-concept for long-term communication restoration through a direct speech brain-computer interface.
[0184] method Test Overview This study was conducted as part of the BRAVO trial (BCI Restoration of Arm and Voice function, clinicaltrials.gov registration number NCT03698149), a single-center clinical trial designed to evaluate the potential of electrocorticography (ECoG, a method of recording neural activity directly from the brain's surface) and custom decoding techniques for long-term communication and motor recovery. The ECoG device used in this study received investigational device exemption approval from the U.S. Food and Drug Administration. At the time of writing, only one clinical trial participant ("Bravo-1," the present study participant) has had an ECoG device implanted.
[0185] participants The participant was a right-handed male who was 36 years old at the start of the study. At age 20, he suffered an extensive bilateral pontine stroke associated with a right vertebral artery dissection, resulting in severe spastic quadriplegia and dysarthria (diagnosed by a speech-language pathologist and neurologist, Figure 5). He is cognitively normal (assessed with the Mini-Mental State Examination). He can produce grunts and groans but cannot produce intelligible speech. He typically communicates using an assistive computer-based typing interface controlled by residual head movements, with a typing speed of approximately 5 correct words or 18 correct letters per minute (Supplementary Methods S1).
[0186] implant device The neural implant used to acquire brain signals from participants was a customized hybrid of a high-density ECoG electrode array (PMT Corporation, MN, USA) and a pedestal connector (Blackrock Microsystems, UT, USA). The ECoG array consisted of 128 flat, disk-shaped electrodes with a center-to-center spacing of 4 mm. During surgical implantation, the speech sensorimotor cortex was exposed through a craniotomy, and the array was placed on the surface of the brain within the subdural space. The dura was sutured closed, and a skull flap was replaced. A percutaneous pedestal connector was placed at a separate site and secured to the skull with small titanium screws. This pedestal connector is an externally accessible platform that can acquire brain signals and transmit them to a computer via a detachable digital connector and cable (Figure 1). Participants underwent surgical implantation of the device in early 2019. The surgery was successful, and they recovered uneventfully. Electrode coverage allowed sampling from multiple cortical regions involved in speech processing, including the left precentral gyrus, postcentral gyrus, posterior middle frontal gyrus, and parts of the posterior inferior frontal gyrus [ 8 , 10 - 12 ].
[0187] Neural data acquisition and real-time processing Using a digital signal processing unit and peripheral hardware (NeuroPort system, Blackrock Microsystems), signals from all 128 channels of the implant device were acquired and transmitted to a separate computer running custom software for real-time analysis (Supplementary Methods S2, Figures 6 and 7) [16, 21]. On this computer, high-gamma activity (neural oscillations within the 70–150 Hz frequency range) for each channel was measured and used during all subsequent analyses and real-time decoding.
[0188] Task Design Participants engaged in two tasks: an isolated word task and a sentence task (Supplementary Methods S3). In each trial of each task, participants were visually presented with a text target and then attempted to produce (speak aloud) that target.
[0189] In the isolated word task, participants attempted to produce individual words from a set of 50 English words. This word set included common English words that could be used to create a variety of sentences, including care-related words and words requested by the participant. On each trial, participants were presented with one of these 50 words and, after a short delay, attempted to produce that word when a visual go-cue was presented.
[0190] In the isolated word task, participants attempted to generate word sequences from a set of 50 English sentences, each consisting of only words from a 50-word set (Supplementary Methods S4 and S5). In each trial, participants were presented with a target sentence and attempted to generate the words from that sentence (in order) as quickly as they were comfortable. Throughout the trial, the word sequence decoded from neural activity was updated in real time and displayed as feedback to the participant.
[0191] modeling We used neural activity collected during the task to train, optimize, and evaluate custom models (Supplementary Methods S6 and S7, Figure 8, Supplementary Table S1). Specifically, we created speech detection and word classification models, both of which leverage deep learning techniques to make predictions from neural activity. We used a decoding pipeline containing these two models, a language model, and a Viterbi decoder (Figure 1) to decode sentences from participants' neural activity in real time during the sentence task.
[0192] The speech detector processed each time point of neural activity during the task and detected the start and end of trial word-generation events in real time (Supplementary Methods S8, Figure 9). We fitted this model using only neural data and task timing information from the isolated-word task.
[0193] For each detected event, the word classifier predicted a set of word probabilities by processing neural activity spanning from 1 second before to 3 seconds after the detected event (Supplementary Methods S9, Figure 10). The predicted probability associated with each word in the 50-word set quantified how likely it was that participants attempted to produce that word during the detected event. We fitted this model using neural data from the isolated word task.
[0194] In English, certain word sequences are more likely than others. We exploit this underlying structure by using a language model that yields the probability of the next word given the previous words in the sequence [22, 23] (Supplementary Methods S10). The model was trained on a collection of sentences consisting only of words from a 50-word set obtained using a custom task on a crowdsourcing platform (Supplementary Methods S4).
[0195] We used a custom Viterbi decoder as the final component of the decoding pipeline, a type of model that determines the most likely word sequence given predicted word probabilities from a word classifier and word sequence probabilities from a language model
[24] (Supplementary Methods S11, Figure 11). By incorporating a language model, the Viterbi decoder was able to decode more plausible sentences than those resulting from simply concatenating predicted words from the word classifier.
[0196] evaluation To evaluate the performance of our decoding pipeline, we analyzed the decoded sentences in real time using two metrics: word error rate and words per minute (Supplementary Methods S12). The word error rate of a decoded sentence is defined as the edit distance (number of word errors in that sentence) divided by the number of words in the target sentence. The words per minute metric measures the number of words decoded per minute of neural data. We also measured the latency of our system during real-time decoding.
[0197] To further characterize the detection and classification of word production trials from participants' neural activity, we processed isolated word data using a speech detector and a word classifier in offline analysis (see Supplementary Methods S13). To assess how performance was affected by the amount of training data, we measured classification accuracy using predicted word probabilities from the word classifier while varying the number of trials used during training. Here, classification accuracy is equal to the proportion of predictions in which the word classifier correctly assigned the highest probability to the target word. We also measured the contribution of each electrode to detection and classification by measuring the influence of each channel of neural activity on the model's predictions [17, 25].
[0198] To investigate the clinical viability of our approach for long-term application, we used isolated word data to evaluate the stability of the acquired ECoG signals over time (Supplementary Methods S14). First, we determined whether the magnitude of the neural responses collected during word generation trials changed over the course of the 81-week study period. We also evaluated whether detection and classification performance remained stable throughout the study period by training and testing models using neural data sampled from four different date ranges ("early," "middle," "late," and "late") and comparing the resulting classification accuracy and electrode contributions.
[0199] statistical analysis The statistical tests used in this study are listed with their corresponding significance claims, and a detailed description of the tests is provided in Supplementary Methods S15. Briefly, we used the Wilcoxon signed-rank test to compare decoding performance to chance, evaluate the impact on language model performance (using a word error rate metric), linear mixed-effects modeling to assess signal stability, Fisher's exact test and McNemar's exact test to compare classification accuracy across different date ranges, and the Wilcoxon signed-rank test to compare electrode contributions across different date ranges. An alpha level of 0.01 was used for all tests. When neural data used in individual statistical tests of the same type were not independent of each other, a Holm-Bonferroni correction was used to account for multiple comparisons.
[0200] result Sentence Decoding During real-time sentence decoding, the median decoded word error rate across sentence blocks (each block containing 10 trials) was 60.5% without language modeling and 25.6% with language modeling (Figure 2A). The lowest word error rate observed for a single test block was 6.98% (with language modeling). Word error rates were significantly better than chance and significantly reduced when incorporating a language model (P < 0.001, one-sided Wilcoxon signed-rank test, three-way Holm-Bonferroni correction). Across all 150 trials, the median decoding speed was 15.2 words per minute when including all decoded words and 12.5 words per minute when including only correctly decoded words (Figure 2B). In 92.0% of trials, the number of detected words equaled the number of words in the target sentence (Figure 2C). The detected sentence length was at least one word too short in 2.67% of trials and at least one word too long in 5.33% of trials. Across all 15 sentence blocks, five speech events were incorrectly detected before the first trial in the block and were excluded from real-time decoding and analysis (all other detected speech events were included). For nearly all target sentences, the mean edit distance decreased when using the language model (Figure 2D). Furthermore, more than half of the sentences were decoded without error (80 out of 150 trials using language modeling, indicated by an edit distance of zero). The use of the language model during decoding improved performance by correcting for grammatically and semantically implausible word predictions (Figure 2E). The mean latency associated with real-time word prediction was estimated to be 4.0 seconds (with a standard deviation of 0.91 seconds).
[0201] Word Detection and Classification During offline analysis of isolated word generation trials using the detected time window of cortical activity, classification accuracy increased with increasing amounts of training data (up to 47.1% when all available data was used; Figure 3A). Performance improved more rapidly during the first 4 hours of training data, and less rapidly during the following 5 hours, but did not plateau. Of 9,000 word generation trials in the isolated word data, 98% were successfully detected (191 trials were not associated with a detected event), and 968 detected events were spurious (not associated with a trial; see Figures 12 and 13 for additional isolated word analysis results). Electrodes contributing to word classification performance were primarily localized in the most ventral aspect of the ventral sensorimotor cortex (vSMC), with electrodes in the dorsal aspect of the vSMC contributing to both speech detection and word classification performance (Figure 3B). Overall, electrode contributions were more distributed toward speech detection than toward word classification, with over 50% of the total contributions coming from the top 37 electrodes for word classifiers and the top 50 electrodes for speech detectors. Word confusion analysis revealed consistent classification accuracy across the majority of word targets (Figure 3C; mean classification accuracy along the diagonal of the row-normalized confusion matrix was 47.1% and standard deviation was 14.5%).
[0202] Long term signal stability We observed relatively stable single-trial neural activity patterns during word-generation trials throughout the 81-week study period (Figure 4A). Across all electrodes and isolated-word trials, there was a slight overall negative effect of time since implantation on the magnitude of neural responses during speech trials (slope = -0.00021, SE = 0.000011, P < 0.001, linear mixed-effects modeling, 129-way Holm-Bonferroni correction, Figure 14). However, individual electrode modeling revealed significant effects at only four of the 128 electrodes (one positive, three negative, P < 0.01, linear mixed-effects modeling, 129-way Holm-Bonferroni correction).
[0203] By training and testing speech detectors and word classifiers on subsets of isolated-word data from separate date ranges, we found that classification accuracy was lowest for the earliest subset and relatively consistent across the remaining subsets (P = 0.0015 for the "early" vs. "late" comparison, P = > 0.01 for all other comparisons, two-tailed Fisher's exact test, 10-way Holm-Bonferroni correction, Figure 4B). When evaluating data in the two most recent subsets, classification accuracy was significantly higher when training models on data from within the same subset, as opposed to data from other subsets (P < 0.001 for the "late" and "latest" subsets, P > 0.01 for the other subsets, two-tailed McNemar's exact test, 10-way Holm-Bonferroni correction). There was no significant change in electrode contribution across the four subsets (all P > 0.32, two-tailed Wilcoxon signed-rank test, uncorrected).
[0204] Consideration Using high-resolution recordings of cortical activity in severely paralyzed individuals, we demonstrated that complete words and sentences could be decoded in real time. Our deep learning models were able to detect and classify word-generation attempts from neural activity, and these models, along with language modeling techniques, were used to decode a variety of meaningful sentences. Signals recorded from the neural interface exhibited stability throughout the study period, enabling successful decoding even up to 90 weeks after surgical implantation. Collectively, these results have immediate practical implications for paralyzed individuals who may benefit from speech neuroprosthetic technologies.
[0205] Previous demonstrations of decoding words and sentences from neural activity were performed in participants who retained normal speech and did not require assistive technology for communication [14-17]. When decoding speech from non-speech-deprived individuals, the lack of precise temporal alignment between intended speech and neural activity poses significant challenges during model training. Here, we managed this temporal alignment issue with detection techniques [16, 26, 27] and classifiers that leverage advances in machine learning, such as model ensembles and data augmentation (described in Supplementary Methods S9), to increase tolerance to small temporal variations [28, 29]. Additionally, our decoding model exploits neural activity patterns in the ventral sensorimotor cortex, consistent with previous studies suggesting this region is involved in normal speech production [8, 11, 12]. These results demonstrate the persistence of functional cortical speech representations more than 15 years after dysarthria, similar to previous findings of limb-related cortical motor representations in tetraplegics after several years of motor decline
[30] .
[0206] Despite imperfect word classification performance, the incorporation of language modeling techniques enabled perfect decoding in over half of the sentence trials. This improvement was driven by leveraging additional probabilistic information from the word classifier (beyond the most likely word identifier for each detected word production trial), allowing the decoder to correct previous errors given new input. These results demonstrate the benefits of integrating linguistic information when decoding speech from neural recordings. Speech decoding approaches are typically usable with word error rates below 30%
[31] , suggesting that our approach can be readily applied in clinical settings.
[0207] A fundamental consideration when designing a long-term brain-computer interface (BCI) is the choice of neural recording modality (e.g., invasive vs. non-invasive) and its impact on the resolution, spatial coverage, and stability of the acquired neural signals. Previous motor control BCI studies have demonstrated that electrocorticography (ECoG, the recording modality used in this study) has relatively high signal stability over long evaluation periods compared to other recording modalities [4, 32-34]. However, these decoding efforts were constrained by limited channel count and spatial coverage. By leveraging the wide spatial coverage and high spatial resolution of our high-density ECoG device, words were reliably decoded, with relatively stable cortical activity observed throughout the study (only three electrodes significantly decreased the magnitude of the neural response over time). Offline classification performance improved and then largely stabilized after the first few weeks of the study. This may potentially be explained by the settling of brain tissue during early healing after implantation [35, 36]. Consistent with a recent cursor control study using this implanted device and study participants
[37] , our results show that an ECoG-based BCI can maintain consistent speech decoding performance over several months with occasional model recalibration. Overall, our findings join the demonstration of long-term survivability, safety, and signal stability of ECoG-based interfaces for responsive neurostimulation in epilepsy [35, 36] and long-term BCI controls [34, 37], extending these attributes to include speech BCIs with high-density ECoG.
[0208] Speech is typically the fastest, most natural, and most efficient method of communication for healthy humans.
[38] While our current decoding speed is much slower than natural speech rates, which often exceed 130 words per minute, [38, 39] these results demonstrate the early feasibility of direct speech decoding from cortical signals in paralyzed individuals with dysarthria. From this proof-of-principle, our invention allows us to develop and evaluate novel decoders to enable the generation of a wider variety of sentences with larger vocabularies. Ultimately, through future research to improve decoding accuracy, flexibility, and speed, our goal is to realize the full communicative potential of speech-based neuroprosthetic prostheses for people with severe communication disorders.
[0209] References: 1. Beukelman DR, Fager S, Ball L, Dietz A. AAC for adults with acquired neurological conditions: A review. Augmentative and Alternative Communication 2007;23(3):230-42. 2.Felgoise SH, Zaccheo V, Duff J, Simmons Z.Verbal communication impacts quality of life in patients with amyotrophic lateral sclerosis.Amyotrophic lateral sclerosis and Frontotemporal Degeneration 2016;17(3-4):179-83. 3.Sellers EW, Ryan DB, Hauser CK.Noninvasive brain-computer interface enables communication after brainstem stroke.Science translational medicine 2014;6(257):257re7. 4.Vansteensel MJ,Pels EGM,Bleichner MG,et al.Fully Implanted Brain-Computer Interface in a Locked-In Patient with ALS.New England Journal of Medicine 2016;375(21):2060-6. 5.Pandarinath C,Nuyujukian P,Blabe CH,et al.High performance communication by people with paralysis using an intracortical brain-computer interface.ELife 2017;6:1-27. 6.Brumberg JS,Pitt KM,Mantie-Kozlowski A,Burnison JD.Brain-Computer Interfaces for Augmentative and Alternative Communication:A Tutorial.Am J Speech Lang Pathol 2018;27(1):1-12. 7.Linse K,Aust E,Joos M,Hermann A,Oliver DJ.Communication Matters-Pitfalls and Promise of Hightech Communication Devices in Palliative Care of Severely Physically Disabled Patients With Amyotrophic Lateral Sclerosis.2018;9(July):1-18. 8.Bouchard KE,Mesgarani N,Johnson K,Chang EF.Functional organization of human sensorimotor cortex for speech articulation.Nature 2013;495(7441):327-32. 9.Lotte F,Brumberg JS,Brunner P,et al.Electrocorticographic representations of segmental features in continuous speech.Frontiers in Human Neuroscience 2015;09(February):1-13. 10.Guenther FH,Hickok G.Neural Models of Motor Speech Control.In:Neurobiology of Language.Elsevier;2016.p.725-40. 11.Mugler EM,Tate MC,Livescu K,Templer JW,Goldrick MA,Slutzky MW.Differential Representation of Articulatory Gestures and Phonemes in Precentral and Inferior Frontal Gyri.The Journal of Neuroscience 2018;4653:1206-18. 12.Chartier J,Anumanchipalli GK,Johnson K,Chang EF.Encoding of Articulatory Kinematic Trajectories in Human Speech Sensorimotor Cortex.Neuron 2018;98(5):10421054.e4. 13.Salari E,Freudenburg ZV,Branco MP,Aarnoutse EJ,Vansteensel MJ,Ramsey NF.Classification of Articulator Movements and Movement Direction from Sensorimotor Cortex Activity.Sci Rep 2019;9(1):14165. 14.Herff C,Heger D,de Pesters A,et al.Brain-to-text:decoding spoken phrases from phone representations in the brain.Frontiers in Neuroscience 2015;9(June):1-11. 15.Anumanchipalli GK,Chartier J,Chang EF.Speech synthesis from neural decoding of spoken sentences.Nature 2019;568(7753):493-8. 16.Moses DA,Leonard MK,Makin JG,Chang EF.Real-time decoding of question-and-answer speech dialogue using human cortical activity.Nat Commun 2019;10(1):3096. 17.Makin JG,Moses DA,Chang EF.Machine translation of cortical activity to text with an encoder-decoder framework.Nat Neurosci 2020;23(4):575-82. 18.Martin S,Iturrate I,Millan J del R,Knight RT,Pasley BN.Decoding Inner Speech Using Electrocorticography:Progress and Challenges Toward a Speech Prosthesis.Front Neurosci 2018;12:422. 19.Guenther FH,Brumberg JS,Wright EJ,et al.A Wireless Brain-Machine Interface for Real-Time Speech Synthesis.PLoS ONE 2009;4(12):e8218. 20.Brumberg JS,Wright EJ,Andreasen DS,Guenther FH,Kennedy PR.Classification of intended phoneme production from chronic intracortical microelectrode recordings in speech-motor cortex.Front Neurosci 2011;5:65. 21.Moses DA,Leonard MK,Chang EF.Real-time classification of auditory sentences using evoked cortical activity in humans.J Neural Eng 2018;15(3):036005. 22.Kneser R,Ney H.Improved backing-off for M-gram language modeling.In:1995 International Conference on Acoustics,Speech,and Signal Processing.Detroit,MI,USA:IEEE;1995.p.181-4. 23.Chen SF,GoodmanJ.An empirical study of smoothing techniques for language modeling.Computer Speech & Language 1999;13(4):359-93. 24.Viterbi AJ.Error Bounds for Convolutional Codes and an Asymptotically Optimum Decoding Algorithm.IEEE Transactions on Information Theory 1967;13(2):260-9. 25.Simonyan K,Vedaldi A,Zisserman A.Deep Inside Convolutional Networks:Visualising Image Classification Models and Saliency Maps.In:Bengio Y,LeCun Y,editors.Workshop at the International Conference on Learning Representations.Banff,Canada:2014. 26.Kanas VG,Mporas I,Benz HL,Sgarbas KN,Bezerianos A,Crone NE.Real-time voice activity detection for ECoG-based speech brain machine interfaces.In:19th International Conference on Digital Signal Processing.2014.p.862-5. 27.Dash D,Ferrari P,Dutta S,WangJ.NeuroVAD:Real-Time Voice Activity Detection from Non-Invasive Neuromagnetic Signals.Sensors 2020;20(8):2248. 28.Sollich P,Krogh A.Learning with ensembles:How overfitting can be useful.In:Touretzky DS,Mozer MC,Hasselmo ME,editors.Advances in Neural Information Processing Systems 8.MIT Press;1996.p.190-196. 29.Krizhevsky A,Sutskever I,Hinton GE.ImageNet Classification with Deep Convolutional Neural Networks.In:Pereira F,Burges CJC,Bottou L,Weinberger KQ,editors.Advances in Neural Information Processing Systems 25.Curran Associates,Inc.;2012.p.1097-1105. 30.Shoham S,Halgren E,Maynard EM,Normann RA.Motor-cortical activity in tetraplegics.Nature 2001;413(6858):793-793. 31.Watanabe S,Delcroix M,Metze F,Hershey JR.New era for robust speech recognition: exploiting deep learning.Berlin,Germany:Springer-Verlag; 2017. 32.Chao ZC,Nagasaka Y,Fujii N.Long-term asynchronous decoding of arm motion using electrocorticographic signals in monkey.FrontNeuroeng 2010;3:3. 33.Freudenburg ZV,Branco MP,Leinders S,et al.Sensorimotor ECoG Signal Features for BCI Control:A Comparison Between People With Locked-In Syndrome and Able-Bodied Controls.Front Neurosci 2019;13:1058. 34.Pels EGM,Aarnoutse EJ,Leinders S,et al.Stability of a chronic implanted brain-computer interface in late-stage amyotrophic lateral sclerosis.Clinical Neurophysiology 2019;130(10):1798-803. 35.Rao VR,Leonard MK,Kleen JK,Lucas BA,Mirro EA,Chang EF.Chronic ambulatory electrocorticography from human speech cortex.NeuroImage 2017;153:273-82. 36.Sun FT,Arcot Desai S,Tcheng TK,Morrell MJ.Changes in the electrocorticogram after implantation of intracranial electrodes in humans:The implant effect.Clinical Neurophysiology 2018;129(3):676-86. 37.Silversmith DB,Abiri R,Hardy NF,et al.Plug-and-play control of a brain-computer interface through neural map stabilization.Nat Biotechnol 2020. 38.Hauptmann AG,Rudnicky AI.A comparison of speech and typed input.In:Proceedings of the workshop on Speech and Natural Language-HLT ’90.Hidden Valley,Pennsylvania:Association for Computational Linguistics;1990.p.219-24. 39.Waller A.Telling tales:unlocking the potential of AAC technologies.International Journal of Language & Communication Disorders 2019;54(2):159-69.
[0210] Example 2: Supplementary Methods for Word Decoding Method S1. Participant's Assistive Typing Device Assistive Typing Device Description Participants often communicate with other humans using a commercially available touchscreen typing interface (Tobii Dynavox), controlled by a long (approximately 18 inches) plastic stylus attached to a baseball cap using residual head and neck movements. The device displays letters, words, and other options (such as punctuation) that participants can select with the stylus, allowing them to construct a text string. After creating the desired text string, participants can use their stylus to press an icon that synthesizes the text string into an audible speech waveform. This process of spelling out a desired message and having the device synthesize it is a typical way participants communicate with caregivers and visitors.
[0211] Typing speed assessment task design To compare the neural-based decoding speed achieved with our system, we measured participants' typing speed while they used a typing interface in a custom task. In each trial of this task, a word or sentence was presented on the screen, and participants typed the word or sentence using the typing interface. Participants were instructed not to use any word suggestion or auto-completion options in their interface, but were allowed to use correction features (e.g., backspace or undo options). We measured the time from when the target word or sentence first appeared on the screen to when participants typed the final letter of the target. This duration and the target word or utterance were then used to measure the number of words per minute and the number of correct letters per minute for each trial.
[0212] A total of 35 trials (25 words and 10 sentences) were used. Although punctuation was included when presented to participants, participants were instructed not to enter punctuation during the task. The target words and sentences were as follows: 1. Thirsty 2.I 3. Tired 4.Are 5.Up 6.How 7.Outside 8.You 9. Bad 10.Clean 11.Have 12.Tell 13. Hello 14.Going 15.Right 16. Closer 17.What 18.Success 19.It 20. Family 21.That 22.Help 23.Do 24.Am 25.Okay 26.It is good. 27.I am thirsty. 28.They are coming here. 29.Are you going outside? 30.I am outside. 31.Faith is good. 32.My family is here. 33.Please tell my family. 34.My glasses are comfortable. 35.They are coming outside.
[0213] Typing speed results and discussion Across all trials of this typing task, participants' mean ± standard deviation typing speed was 5.03 ± 3.24 correct words per minute or 17.9 ± 3.47 correct letters per minute. Although these typing speeds are slower than the real-time decoding speed of our approach, the unlimited vocabulary size of the typing interface is a significant advantage over our approach. Given the number of correct letters per minute that participants can achieve using the typing interface, replacing the letters in the interface with 50 words from this task could potentially yield higher decoding speeds and accuracy than those achieved with our approach. However, this typing interface is less natural and appears to require more physical effort than speech attempts, suggesting that the typing interface may be more fatiguing than our approach.
[0214] Method S2. Neural data acquisition and real-time processing Initial data acquisition and preprocessing steps The implanted electrocorticography (ECoG) array (PMT Corporation) contained electrodes arranged in a 16 x 8 grid configuration with 4 mm center-to-center spacing. The rectangular ECoG array was 6.7 cm long, 3.5 cm wide, and 0.51 mm thick, and the electrode contacts were disk-shaped with a 2 mm contact diameter. To process and record neural data, signals were acquired from the ECoG array and processed in several steps involving multiple hardware devices (Figures 6 and 7). First, a headstage (Detachable Digital Link, Blackrock Microsystems) connected to a transcutaneous pedestal connector (Blackrock Microsystems) acquired potentials from the implanted electrode array. The pedestal has a male connector, and the headstage has a female connector. The headstage performed bandpass filtering of the signals using a hardware-based Butterworth filter with a frequency range of 0.3 Hz to 7.5 kHz. The digitized signal (16 bits, 250 nV per bit resolution) was then sent via an HDMI cable to a digital hub (Blackrock Microsystems), which then transmitted the data via a fiber optic cable to a Neuroport system (Blackrock Microsystems). During early recording sessions, before the digital headstage was approved for use in human studies, a human patient cable (Blackrock Microsystems) was used to connect the pedestal to a front-end amplifier (Blackrock Microsystems). The amplifier amplifies and digitizes the signal before transmitting it via fiber optic to the Neuroport system. This Neuroport system sampled all 128 channels of ECoG data at 30 kHz, applied software-based line noise cancellation, and performed anti-aliasing low-pass filtering at 500 Hz. The processed signal was then streamed to another real-time processing machine (Colfax International) at 1 kHz.The Neuroport system also acquired, streamed, and stored synchronous recordings of relevant sounds at 30 kHz (microphone input and speaker output from a real-time processing computer).
[0215] Further preprocessing and feature extraction A real-time processing computer, a Linux machine (64-bit Ubuntu 18.04, 48 Intel Xeon Gold 6146 3.20 GHz processors, 500 GB RAM), analyzed and processed the incoming neural data, executed the tasks, performed real-time decoding, and stored task data and metadata on disk using a custom software package called Real-Time Neural Speech Recognition (rtNSR) [1, 2]. This software performed the following preprocessing steps in real time on all acquired neural signals:
[0216] A common average criterion was applied to each time sample of acquired ECoG data (across all electrodes), which is a standard technique for reducing shared noise in multi-channel data [3, 4].
[0217] Eight bandpass finite impulse response (FIR) filters with logarithmically increasing center frequencies in the high-gamma band (72.0, 79.5, 87.8, 96.9, 107.0, 118.1, 130.4, and 144.0 Hz, rounded to one decimal place) were applied. Each of these 390th-order filters was designed using the Parks-McClellan algorithm [5].
[0218] Analytical amplitude values were calculated for each band and channel using a 170th-order FIR filter designed by the Parks-McClellan algorithm to approximate the Hilbert transform. For each band and channel, an analytic signal was estimated using the original signal (delayed 85 samples, half the filter order,) as the real component and the Hilbert transform of the original signal (approximated by this FIR filter) as the imaginary component [6]. Analytical amplitude values were then obtained by calculating the magnitude of each of these analytic signals. This analytical amplitude calculation was applied only to every fourth sample of the band-passed signal, and the analytical amplitude was reduced to 200 Hz.
[0219] A single high-gamma analytical amplitude measure was calculated for each channel by averaging the analytical amplitude values across the eight bands.
[0220] High-gamma analysis amplitude values were z-scored for each channel using the method of Welford with a 30-second sliding window [7].
[0221] These high-gamma analysis amplitude z-score time series (sampled at 200 Hz) were used in all analyses and during online decoding.
[0222] Portability and cost of hardware infrastructure In this study, the hardware used was fairly large but still portable, with most hardware components mounted on a mobile rack approximately 76 cm long and 76 cm wide. All data collection and online decoding tasks were performed in the participant's bedroom or a small office near the participant's residence. While all use of the hardware was supervised throughout the clinical trial, the hardware and software setup procedures required to begin recording were simple. After a few hours of training, caregivers, with appropriate regulatory approval, can prepare the system of the present invention for use by participants without our direct supervision. To set up the system for use, caregivers perform the following steps: 1. Remove and clean the percutaneous connector cap. This protects the external electrical contacts of the percutaneous connector while the system is not in use. 2. Clean the transcutaneous connector, digital link, and scalp area around the transcutaneous connector 3. Connect the digital link to the transcutaneous connector 4. Turn on your computer and launch the software 5. Ensure the screen is properly positioned in front of the participant for use Then, to disengage from the system, the caregiver takes the following steps: 1. Close the software and turn off your computer 2. Disconnect the digital link from the transcutaneous connector 3. Clean the transcutaneous connector, digital link, and scalp area around the transcutaneous connector 4. Put the percutaneous connector cap back onto the percutaneous connector
[0223] The entire hardware infrastructure was quite expensive, primarily due to the relatively high cost of the new Neuroport system (compared to the costs of the other hardware devices used in this study). However, recent work has demonstrated that relatively inexpensive and portable brain-computer interface systems can be deployed without significant degradation in system performance (compared to typical systems, including Blackrock Microsystems devices, such as the system used in this study) [8]. This study's demonstration suggests that future iterations of the inventive hardware infrastructure can be made cheaper and more portable without sacrificing decoding performance.
[0224] Computational Modeling Infrastructure The collected data from the real-time processing computer was uploaded to our laboratory's computation and storage server infrastructure, where multiple NVIDIA V100 GPUs were used to reduce computation time and to adapt and optimize the decoding model. The finalized model was then downloaded to the real-time processing computer for online decoding.
[0225] Method S3. Task design All data were collected in a series of "blocks." Each block lasted approximately 5 or 6 minutes and consisted of multiple trials. There were two types of tasks: isolated word tasks and sentence tasks.
[0226] Isolated word task In the isolated word task, participants attempted to produce individual words from a set of 50 words while we simultaneously recorded their cortical activity for offline processing. The word set was chosen based on the following criteria: 1. You can easily create a variety of sentences using these words. 2. Be able to easily communicate basic care needs using that vocabulary. 3. Participants were interested in including those words. We iterated several versions of the 50-word set using feedback that participants provided to us through commercially available assistive communication technology. 4. The desired word count was large enough to allow for meaningful sentence variety, yet small enough to allow adequate neural categorization performance. This latter criterion was informed by exploratory preliminary assessments with participants after device implantation (before collecting any of the data analyzed in this study). The list of words included in this 50-word set is provided at the end of this section.
[0227] To keep task blocks short in duration, this word set was arbitrarily divided into three disjoint subsets: two subsets containing 20 words each, and the third subset containing the remaining 10 words. During each block of the task, participants attempted to generate each word in one of these subsets twice, resulting in a total of either 40 or 20 word generation attempts per block (depending on the size of the word subset). In the three blocks of the third, smaller subset, participants attempted to generate each of the 10 words in that subset four times (instead of the usual two).
[0228] Each trial in this task block began with a blank screen against a black background. After 1 second (or 1.5 seconds in a very small number of blocks), one word from the current word subset appeared on the screen in white text surrounded by four period characters on either side (e.g., if the current word was "Hello," the text "...Hello..." appeared). For the next 2 seconds, the outer periods on either side (the first and last characters of the displayed text string) disappeared every 500 milliseconds, visually representing a countdown. Once the final period on either side of the word disappeared, the text turned green and remained on the screen for 4 seconds. This color transition from white to green represented the go cue for each trial, and participants were instructed to attempt to generate the word as soon as the text turned green. The task then continued with the next trial. Word presentation order was randomized within each task block. Participants chose this countdown-style task paradigm from a range of possible paradigm options we presented during the pre-operative interview, claiming that using consistent countdown timing allowed them to better align their production attempts with the go cue on each trial.
[0229] Statement Task In the sentence task, participants attempted to generate sentences from a set of 50 sentences while we processed their neural activity and decoded them into text. These sentences consisted exclusively of words from a set of 50 words. These 50 sentences were selected semi-randomly from a corpus of possible sentences (see Methods S5). A list of the sentences included in this set of 50 sentences is provided at the end of this section. To keep the duration of task blocks short, we arbitrarily divided this sentence set into five disjoint subsets, each containing 10 sentences. During each block of this task, participants attempted to generate each sentence included in one of these subsets once, resulting in a total of 10 sentences per block.
[0230] Each trial in this task block began with a blank screen horizontally divided into upper and lower halves, both with a black background. After 2 seconds, one of the sentences in the current sentence subset was presented in white text in the upper half of the screen. Participants were instructed to attempt to generate a word in the sentence as quickly as they felt comfortable as soon as the text appeared on the screen. While the target sentence was presented to the participant, their cortical activity was processed in real time by a speech detection model. Each time a trial word production was detected from the acquired neural signals, a set of cyclic ellipses (a text string that cycled between one, two, and three period characters per second) was added to the lower half of the screen as feedback indicating that a speech event had been detected. Word classification, linguistic, and Viterbi decoding models were then used to decode the most likely word associated with the currently detected speech event given the corresponding neural activity and the decoded information from any previous detected events within the current trial. Each time a new word was decoded, it replaced the associated cyclic ellipses text string in the lower half of the screen, providing further feedback to the participant. The Viterbi decoding model, which maintained the most likely word sequence on a trial given the observed neural activity, often updated its predictions of previous speech events given new speech events, changing previously decoded words in the feedback text string as new information became available. After a predetermined time had elapsed from the detected onset of the most recent speech event, the sentence target text changed from white to blue, indicating that the decoding portion of the trial had ended and the decoded sentence was finalized for that trial. This predetermined time was either 9 or 11 seconds, depending on the block type (see next paragraph). After 3 seconds, the task continued with the next trial.
[0231] Two types of blocks of sentence tasks were collected: optimization blocks and test blocks. The differences between these two types of blocks are as follows: 1. The optimization block was used to perform hyperparameter optimization, and the test block was used to evaluate the performance of the decoding system. 2. The intermediate (non-optimized) model was used when collecting the optimization blocks, and the finalized (optimized) model was used when collecting the test blocks. 3. Detected speech attempts and decoded word sequences were always provided to participants as feedback during this task, but during the collection of optimization blocks, participants were instructed not to repeat words if a speech event was missed or to use the feedback to modify which word they attempted to produce. We included these instructions to protect the integrity of the data for use in the hyperparameter optimization procedure (if participants modified their own data behavior due to imperfect speech detection, a mismatch between the prompted word sequence and the word sequence they actually attempted could interfere with the optimization procedure). However, during test blocks, participants were encouraged to take the feedback into account when attempting to produce the target sentence. For example, if a trial word production was not detected, participants could repeat the production attempt before moving on to the next word. 4. During the optimization block, the predetermined time (see previous paragraph) that controlled when the decoded word sequence in each trial was finalized was set to 9 seconds. During the test block, this task parameter was set to 11 seconds to allow participants extra time to incorporate the feedback provided by the decoding pipeline.
[0232] We also collected a conversational variation of the sentence task to demonstrate that the decoding approach could be used in a more open-ended environment, allowing participants to generate custom responses to questions from 50 words. In this variation of the task, instead of being prompted with a target sentence to attempt to repeat, participants were prompted with a question or statement that mimicked their conversation partner and instructed to attempt to generate a response to the prompt. Other than this change in the conversational prompt and task instructions to participants, this variation of the task was identical to the regular version. We did not perform any analyses using the data collected from this variation of the sentence task; it was used for demonstration purposes only. This variation of the task is shown in Figure 1 of the main text.
[0233] Word and Sentence List The 50-word set used in this study was as follows: 1.Am 2.Are 3. Bad 4.Bring 5.Clean 6.Closer 7.Comfortable 8.Coming 9.Computer 10.Do 11. Faith 12.Family 13. Feel 14. Glasses 15.Going 16.Good...
Claims
【Claim 1】 An apparatus as described in the body of the specification and shown in the drawings.