Methods and systems for translation of neural activity into embodied digital-avatar animation
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-06
- Publication Date
- 2026-03-11
AI Technical Summary
Individuals with anarthria and paralysis face challenges in communicating effectively due to limitations in decoding brain activity into naturalistic speech and gestures, with existing methods only enabling text output with limited speed and vocabulary.
The development of methods and systems that utilize neural interfaces and machine learning techniques to decode brain activity from the sensorimotor cortex, enabling the production of speech sounds, facial movements, and non-speech communicative gestures, using self-supervised learning to discretize actions into meaningful outputs and synthesize them into audio and visual stimuli.
Enables full, embodied communication for individuals with severe paralysis by restoring the ability to produce speech and perform gestures, enhancing communication beyond text-based outputs with naturalistic and expressive capabilities.
Smart Images

Figure US2024032886_12122024_PF_FP_ABST
Abstract
Description
[0001] Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 METHODS AND SYSTEMS FOR TRANSLATION OF NEURAL ACTIVITY INTO EMBODIED DIGITAL-AVATAR ANIMATION STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH This invention was made with government support under grant no. NIH U01 DC018671- 01A1, awarded by the National Institutes of Health (NIH). The government has certain rights in the invention. CROSS-REFERENCE TO RELATED APPLICATION Pursuant to 35 U.S.C. § 119 (e), this application claims priority to the filing date of United States Provisional Patent Application Serial No. 63,471,485 filed June 6, 2023, the disclosure of which is herein incorporated by reference in its entirety. INTRODUCTION Speech is the ability to express thoughts and ideas through spoken words. Anarthria, or the loss of the ability to articulate speech, can result from a variety of conditions, including stroke, traumatic brain injury, and amyotrophic lateral sclerosis. For paralyzed individuals with severe movement impairment, anarthria hinders communication with family, friends, and caregivers, reducing self-reported quality of life. Speech neuroprostheses have the potential to restore communication to people living with paralysis, but naturalistic speed and expressivity remain elusive. Previous demonstrations have shown that it is possible to decode speech from the brain activity of a person with paralysis, but only in the form of text and with limited speed and vocabulary. While text outputs are good for basic communication and messaging, speaking has rich prosody, expressiveness, and identity that can enhance embodied communication beyond what can be conveyed in text alone. SUMMARY Thus, there remains a need for better methods and systems for restoring the ability to communicate to patients with anarthria and paralysis. This invention provides such new and useful methods and systems, addressing the limitations mentioned above. To accomplish this, the invention leverages recent advances in neural interfaces and machine learning techniques which Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 enable models to be trained from highly stable neural recordings using only task go cues for data segmentation, i.e., without any other alignment of neural activity and output features. The methods and systems of the invention, e.g., as described in greater detail below, find use in assisting individuals with communication. In particular, methods, devices, and systems are provided that facilitate full, embodied communication to people living with severe paralysis by restoring the ability to produce speech sounds and facial movements related to speaking, as well as by restoring the ability to perform non-speech communicative gestures. Methods of assisting individuals with communication are provided. In the disclosed methods, cortical activity from a region of the brain associated with movement, speech production, and / or language perception is recorded while an individual attempts to perform an action (e.g., to say words, express an emotion, perform a movement, etc.). Deep learning computational models are used to detect and decode the attempted action from the recorded brain activity. Decoding of actions from brain activity is aided by the use of self-supervised machine learning techniques, which discretize each action into one or more action representations that serve as an effective intermediary for decoding neural activity patterns and features into meaningful action outputs. In addition, methods for synthesizing decoded actions into audio and / or visual stimuli are provided, allowing for more naturalistic and expressive communication for individuals who are unable to speak or experience other mobility limitations that inhibit full embodied communication. In one aspect, methods of producing speech audio directly from the neural activity of an individual are provided. Aspects of the methods include: positioning a neural recording device comprising an electrode at a location in a sensorimotor cortex region of the brain of the subject to record brain electrical signal data associated with attempted speech by the subject; positioning an interface in communication with a computing device at a location on the head of the subject, wherein the interface is connected to the neural recording device; recording the brain electrical signal data associated with attempted speech by the subject using the neural recording device, wherein the interface receives the brain electrical signal data from the neural recording device and transmits the brain electrical signal data to a processor of the computing device; and decoding one or more speech sounds from the recorded brain electrical signal data using the processor, wherein the processor is programmed to use a machine learning model for the decoding. In certain embodiments, the subject has difficulty speaking because of anarthria, a stroke, a traumatic brain injury, a brain tumor, or amyotrophic lateral sclerosis. In some embodiments, the Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 subject has a speech intelligibility of 10% or less for prompted words. In some embodiments, the subject is paralyzed. In certain embodiments, neural activity from the sensorimotor cortex region of the subject’s brain is recorded using a neural recording device comprising an electrode. In some embodiments, the neural recording device includes an electrocorticography (ECoG) electrode array, such as a high density ECoG electrode array. For example, the high density ECoG electrode array may include 250 electrodes or more. In some embodiments, the electrodes may be non-penetrating surface electrodes. In certain embodiments, the neural recording device is positioned on the pial surface of the sensorimotor cortex such that the device covers a middle portion of the superior and / or middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus. In some embodiments, the neural recording device is implanted over the subject’s lateral cortex and is centered on the central sulcus. In some embodiments, the neural recording device acquires brain electrical signal data from a sensorimotor cortex region selected from the precentral gyrus, postcentral gyrus, superior temporal gyrus, middle temporal gyrus, or any combination thereof. In certain embodiments, neural signals acquired by the neural recording device are processed to extract high-gamma activity (HGA) and / or low-frequency signals (LFS). In some embodiments, the HGA electrical signal data may include neural oscillations in a range from 70 Hz to 150 Hz. In some embodiments, the LFS electrical signal data may include neural oscillations in a range from 0.3 Hz to 17 Hz. In certain embodiments, the one or more speech sounds form a word or a sentence. In some embodiments, the subject is limited to a specified word set for the attempted speech. In some embodiments, the word set is chosen in order to enable the subject to express basic concepts and communicate caregiving needs. In some embodiments, the subject may switch between two or more word sets depending on context. For example, the subject may select a first word set to communicate with a caregiver and a second word set to discuss a baseball game with a friend or family member. In some embodiments, the word set includes 100 words or more, or 350 words or more, or 1000 words or more. In certain embodiments, the method further includes providing a series of go cues to the subject indicating when the subject should initiate attempted speech. In some embodiments, the series of go cues are provided visually on a display. In some embodiments, each go cue is preceded Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 by a countdown to the presentation of the go cue, wherein the countdown for the next spoken word or sentence is provided visually on the display and automatically started after each go cue. In some embodiments, the series of go cues are provided with a set interval of time between each go cue. In some embodiments, the subject can control the set interval of time between each go cue. In some embodiments, the processor is programmed to use the recorded brain electrical signal data within a time window following the go cue. In certain embodiments, the machine learning model used to decode the one or more speech sounds from the recorded brain electrical signal data includes a neural network. In some embodiments, the neural network includes one or more convolutional layers and is bidirectional. In some embodiments, the neural network is a recurrent neural network (RNN) such as an RNN including gated recurrent units (GRUs). In certain embodiments, brain electrical signal data is decoded into one or more speech sounds using intermediate representations. In some embodiments, the intermediate representations are a set of discrete speech units. In some embodiments, the discrete speech units are derived using an encoding machine learning model, such as an encoding machine learning model including a neural network. In some embodiments, the neural network includes a transformer encoder. For example, the encoding machine learning model may be a Hidden-Unit Bidirectional Encoder Representations from Transformers (HuBERT) model. In some embodiments, the discrete speech units are derived during self-supervised training of the machine learning model. In certain embodiments, the methods of producing speech audio directly from neural activity further includes training the machine learning model used to decode the one or more speech sounds from the recorded brain electrical signal data. In some embodiments, the training includes: obtaining reference electronic speech waveforms for a plurality of phrases; encoding each reference speech waveform into a temporal sequence of discrete speech units using the encoding model; recording the brain electrical signal data associated with attempted speech by the subject for each one of the plurality of phrases; training the decoding machine learning model to predict the most likely discrete speech unit associated with a segment of electrical signal data using the speech units derived from the reference speech waveforms and the brain electrical signal data associated with attempted speech for each one of the plurality of phrases. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 In certain embodiments, the reference electronic speech waveforms are obtained from a recruited speaker. In some embodiments, the reference electronic speech waveforms are obtained using a text-to-speech algorithm. In certain embodiments, the decoding machine learning model is trained to learn mappings between neural activity patterns of electrical signals in the brain electrical signal data and the discrete speech units. In certain embodiments, the decoding machine learning model is trained using a CTC loss function. In some embodiments, the training may further include validating the machine learning model or machine learning models. In some embodiments, the methods may further include testing the trained machine learning model. In certain embodiments, each speech sound is decoded from one or more discrete speech units using a speech synthesizer. In some embodiments, the speech synthesizer includes a machine learning model. In some embodiments, the synthesizing machine learning model is configured to generate a Mel spectrogram from one or more discrete speech units. In some embodiments, the speech synthesizer further includes a vocoder configured to synthesize an electronic speech waveform from the spectrogram. In certain embodiments, the electronic speech waveform decoded from recorded brain electrical signal data and synthesized using the speech synthesizer is transformed into a personalized electronic speech waveform. For example, the electronic speech waveform may be transformed such that it resembles the subject’s own voice. In some embodiments, the electronic speech waveform is transformed using a machine learning model, such as the YourTTS zero-shot voice conversion model. In certain embodiments, the electronic speech waveform decoded from recorded brain electrical signal data and synthesized using the speech synthesizer is converted into an audible speech waveform using a loudspeaker. In some embodiments, an electronic output device or system is controlled using the electronic speech waveform. For example, an animated avatar may be controlled to perform orofacial movements based on the electronic speech waveform. In another aspect, the methods of producing speech audio directly from neural activity are provided as computer implemented methods. Aspects of the computer implemented methods include: receiving the brain electrical signal data associated with attempted speech by the subject using the neural recording device; and decoding one or more speech sounds from the recorded brain electrical signal data using a machine learning model. In some embodiments, the computer Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 implemented methods further include implementing any of the embodiments of the methods of producing speech audio directly from recorded brain electrical signals described herein using a computer. In certain embodiments, the computer implemented method further includes storing a user profile for the subject including information regarding the patterns of electrical signals in the recorded brain electrical signal data associated with attempted speech by the subject. In another aspect, a non-transitory computer-readable medium is provided, the non- transitory computer-readable medium including program instructions that, when executed by a processor in a computer, causes the processor to perform the computer implemented methods described herein. In another aspect, a kit is provided, the kit including the non-transitory computer-readable medium and instructions for decoding brain electrical signal data associated with attempted speech by a subject. In another aspect, a system for producing speech audio directly from neural activity is provided, the system including: a neural recording device including an electrode adapted for positioning at a location in a sensorimotor cortex region of the brain of the subject to record brain electrical signal data associated with attempted speech by the subject; a processor programmed to use a machine learning model to decode one or more speech sounds from the recorded brain electrical signal data according to a computer implemented method described herein; an interface in communication with a computing device, said interface adapted for positioning at a location on the head of the subject, wherein the interface receives the brain electrical signal data from the neural recording device and transmits the brain electrical signal data to the processor; and an audio speaker for playing the one or more speech sounds from the recorded brain electrical signal data. In certain embodiments, the neural recording device includes an ECoG electrode array, such as a high density ECoG electrode array. For example, the high density ECoG electrode array may include 250 electrodes or more. In some embodiments, the electrodes may be non-penetrating surface electrodes. In certain embodiments, the interface includes a percutaneous pedestal connector attached to the subject's cranium. In certain embodiments, the interface further includes a headstage that is connectable to the percutaneous pedestal connector. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 In certain embodiments, the processor is provided by a computer or handheld device (e.g., a cell phone or tablet). In another aspect, a kit including a system described herein and instructions for using the system for recording and decoding brain electrical signal data associated with an attempted action by a subject. In one aspect, methods of controlling an electronic output device or system to perform one or more actions using brain electrical signals are provided. Aspects of the methods include: positioning a neural recording device including an electrode at a location in a sensorimotor cortex region of the brain of the subject to record brain electrical signal data associated with an attempted action by the subject; positioning an interface in communication with a computing device at a location on the head of the subject, wherein the interface is connected to the neural recording device; recording the brain electrical signal data associated with the attempted action by the subject using the neural recording device, wherein the interface receives the brain electrical signal data from the neural recording device and transmits the brain electrical signal data to a processor of the computing device; decoding one or more electronic output device or system actions from the recorded brain electrical signal data, wherein the processor is programmed to use a machine learning model for the decoding; and controlling the electronic output device or system to perform the one or more decoded electronic output device or system actions. In certain embodiments, the subject has difficulty speaking because of anarthria, a stroke, a traumatic brain injury, a brain tumor, or amyotrophic lateral sclerosis. In some embodiments, the subject is paralyzed. In some embodiments, the subject is quadriplegic and / or experiences partial or total facial paralysis. In certain embodiments, neural activity from the sensorimotor cortex region of the subject’s brain is recorded using a neural recording device including an electrode. In some embodiments, the neural recording device includes an ECoG electrode array, such as a high density ECoG electrode array. For example, the high density ECoG electrode array may include 250 electrodes or more. In some embodiments, the electrodes may be non-penetrating surface electrodes. In certain embodiments, the neural recording device is positioned on the pial surface of the sensorimotor cortex such that the device covers a middle portion of the superior and / or middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus. In some embodiments, the neural recording device is implanted over the subject’s lateral cortex and is centered on the central Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 sulcus. In some embodiments, the neural recording device acquires brain electrical signal data from a sensorimotor cortex region selected from the precentral gyrus, postcentral gyrus, superior temporal gyrus, middle temporal gyrus, or any combination thereof. In certain embodiments, neural signals acquired by the neural recording device are processed to extract HGA and / or LFS. In some embodiments, the HGA electrical signal data may include neural oscillations in a range from 70 Hz to 150 Hz. In some embodiments, the LFS electrical signal data may include neural oscillations in a range from 0.3 Hz to 17 Hz. In certain embodiments, the interface includes a percutaneous pedestal connector attached to the subject's cranium. In certain embodiments, the interface further includes a headstage that is connectable to the percutaneous pedestal connector. In certain embodiments, the processor is provided by a computer or handheld device (e.g., a cell phone or tablet). In certain embodiments, the attempted action performed by the subject is the movement of one or more body parts of the subject. In some embodiments, the one or more body parts is one or more limbs of the subject, such as the subject’s arms or fingers. In some embodiments, the movement of the one or more body parts is directed to perform a task. In some embodiments, the attempted action performed by the subject is attempted speech. In certain embodiments, the methods further include mapping the actions an electronic output device or system is capable of performing to attempted actions performed by the subject. In some embodiments, the attempted action performed by the subject is different than the one or more electronic device or system actions. In some embodiments, the attempted action performed by the subject is a hand gesture and the action of the electronic output device or system is powering on / off or opening / closing. In some embodiments, the attempted action performed by the subject corresponds to the one or more electronic device or system actions. In certain embodiments, the subject is limited to a specified action set for the attempted action. In some embodiments, the action set is chosen in order to enable the subject to control the electronic output device or system for a specific purpose. In some embodiments, the subject may switch between two or more action sets depending on context. For example, the electronic device or system may include a humanoid avatar and the subject may select a first action set to control the avatar for the purpose of communicating with a caregiver and a second action set to control Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 the avatar for the purpose of playing a baseball video game. In some embodiments, the action set includes 100 actions or more, or 350 actions or more, or 1000 actions or more. In certain embodiments, the method further includes providing a series of go cues to the subject indicating when the subject should initiate an attempted action. In some embodiments, the series of go cues are provided visually on a display. In some embodiments, each go cue is preceded by a countdown to the presentation of the go cue, wherein the countdown for the action is provided visually on the display and automatically started after each go cue. In some embodiments, the series of go cues are provided with a set interval of time between each go cue. In some embodiments, the subject can control the set interval of time between each go cue. In some embodiments, the processor is programmed to use the recorded brain electrical signal data within a time window following the go cue. In certain embodiments, the machine learning model used to decode the one or more electronic output device or system actions from the recorded brain electrical signal data includes a neural network. In some embodiments, the neural network includes one or more convolutional layers and is bidirectional. In some embodiments, the neural network is an RNN such as an RNN including GRUs. In certain embodiments, brain electrical signal data is decoded into one or more device or system actions using intermediate representations. In some embodiments, the intermediate representations are a set of discrete action representations. In some embodiments, the discrete action representations are derived using an encoding machine learning model, such as an encoding machine learning model including a neural network. In some embodiments, the neural network includes one or more components of a variational autoencoder and / or may include one or more rectified linear unit (ReLU) activations. For example, the encoding machine learning model may be a vector-quantized variational autoencoder (VQ-VAE) model. In some embodiments, the discrete action representations are derived during self-supervised training of the machine learning model. In certain embodiments, the electronic output device or system is configured to produce or display the same action that is attempted by the subject. In some embodiments, the electronic device or system is a prosthetic limb. In some embodiments, the electronic device or system includes a humanoid avatar. For example, the electronic output device or system may include a visual display and / or a loudspeaker configured to present a humanoid avatar. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 In certain embodiments, the methods of controlling an electronic output device or system to perform one or more actions using recorded neural activity further includes training the machine learning model used to decode the attempted action from the recorded brain electrical signal data. In some embodiments, the training includes: obtaining reference output device or system actions for a plurality of actions; encoding each reference action into a temporal sequence of discrete action representations using the encoding machine learning model; recording the brain electrical signal data associated with attempted action by the subject for each one of the plurality of actions; training the decoding machine learning model to predict the most likely discrete action representation associated with a segment of electrical signal data using the action representations derived from the reference output device or system actions and the brain electrical signal data associated with the attempted action for each one of the plurality of actions. In some embodiments, device or system actions include avatar animations. In some embodiments, the avatar animations are of the avatar performing the same action as the action attempted by the subject. In certain embodiments, the reference avatar animations are obtained from an avatar- animation system. In some embodiments, the reference avatar animations are of the avatar performing orofacial movements for speech. In some embodiments, the speech orofacial movements include one or more of: a tongue tip raise, tongue retraction, tongue body raise, tongue advance, lip rounding, pinching nostril flare, upper lip pull, lower lip tuck, lower lip push, lower lip pull, lip flare, jaw opening, lip compression, and / or lip adduction. In certain embodiments, the reference avatar animations are of the avatar performing non- speech communicative gestures. In some embodiments, the non-speech communicative gestures include the abduction, adduction, flexion, extension, and / or circumduction of one or more body parts. In some embodiments, the non-speech communicative gestures include emotional expressions using facial muscles. For example, the emotional expressions may include happy, sad, and / or surprised expressions. In certain embodiments, the decoding machine learning model is trained to learn mappings between neural activity patterns of electrical signals in the brain electrical signal data and the discrete action representations. In certain embodiments, the decoding machine learning model is trained using a CTC loss function. In some embodiments, the training may further include validating the machine learning model or machine learning models. In some embodiments, the methods may further include testing the trained machine learning model. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 In certain embodiments, each electronic output device or system action is decoded from one or more discrete action representations using a decoder. In some embodiments, the decoder includes a machine learning model. In some embodiments, the decoder machine learning model is trained with the encoding machine learning model or using the encoding machine learning model. In some embodiments, the encoding machine learning model and the decoder machine learning model are separate components of the same machine learning model. For example, a VQ-VAE may be used to discretize or encode each reference electronic output device or system action into one or more discrete action representations and the VQ-VAE’s decoder may be used to decode electronic output device or system actions from discrete action representations decoded from recorded brain electrical signal data. In certain embodiments, the attempted actions performed by the subject and the actions of the device or system include both speech associated actions and non-speech communicative gestures. In some embodiments, the decoding machine learning model is trained to discriminate between actions performed by different regions of the body and / or between speech associated actions and non-speech communicative gestures. In certain embodiments, the electronic output device or system is controlled using actions decoded from multiple machine learning models. In some embodiments, the actions decoded from the multiple machine learning models occur concurrently and the electronic output device or system is controlled to perform the actions simultaneously. In some embodiments, separate decoding machine learning models are trained for speech and orofacial movements for speech. In some embodiments, a decoded non-speech communicative gesture action affects how the electronic output device or system performs a decoded speech associated action. For example, a decoded emotional expression may affect the inflection of decoded speech performed by an avatar. In certain embodiments, the electronic output device or system action decoded from recorded brain electrical signal data is used by the subject to communicate with one or more individuals in person. In some embodiments, the electronic device or system action decoded from recorded brain electrical signal data is used by the subject to communicate with one or more individuals virtually. For example, an animated avatar may be controlled to communicate with one or more individuals in an interactive virtual environment. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 In certain embodiments, an animated avatar controlled using brain electrical signals is used by the subject to play a video game. In some embodiments, the avatar is used by the subject for physical therapy. In some embodiments, avatar is used by the subject to control one or more electronic devices in the subject’s environment. In another aspect, the methods of controlling an electronic output device or system to perform one or more actions using brain electrical signals are provided as computer implemented methods. Aspects of the computer implemented methods include: receiving the brain electrical signal data associated with an attempted action by the subject using the neural recording device; decoding one or more electronic output device or system actions from the recorded brain electrical signal data using a machine learning model; and controlling the electronic output device or system to perform the one or more decoded electronic output device or system actions. In some embodiments, the computer implemented methods further include implementing any of the embodiments of the methods of controlling an electronic output device or system using brain electrical signals described herein using a computer. In certain embodiments, the computer implemented method further includes storing a user profile for the subject including information regarding the patterns of electrical signals in the recorded brain electrical signal data associated with an attempted action by the subject. In another aspect, a non-transitory computer-readable medium is provided, the non- transitory computer-readable medium including program instructions that, when executed by a processor in a computer, causes the processor to perform the computer implemented methods described herein. In another aspect, a kit is provided, the kit including the non-transitory computer-readable medium and instructions for decoding brain electrical signal data associated with an attempted action by a subject. In another aspect, a system for controlling an avatar to perform one or more actions using brain electrical signals is provided, the system including: a neural recording device including an electrode adapted for positioning at a location in a sensorimotor cortex region of the brain of the subject to record brain electrical signal data associated with an attempted action by the subject; a processor programmed to use a machine learning model to decode an avatar animation from the recorded brain electrical signal data according to a computer implemented method described herein; an interface in communication with a computing device, said interface adapted for Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 positioning at a location on the head of the subject, wherein the interface receives the brain electrical signal data from the neural recording device and transmits the brain electrical signal data to the processor; and a display component for displaying the avatar animation from the recorded brain electrical signal data. In certain embodiments, the neural recording device includes an ECoG electrode array, such as a high density ECoG electrode array. For example, the high density ECoG electrode array may include 250 electrodes or more. In some embodiments, the electrodes may be non-penetrating surface electrodes. In certain embodiments, the interface includes a percutaneous pedestal connector attached to the subject's cranium. In certain embodiments, the interface further includes a headstage that is connectable to the percutaneous pedestal connector. In certain embodiments, the processor is provided by a computer or handheld device (e.g., a cell phone or tablet). In certain embodiments, the display component includes a computer monitor, a television, and / or a visual projection device. In some embodiments, the visual display includes a virtual reality headset, goggles, or contacts. In some embodiments, the visual display includes an augmented reality headset, goggles, or contacts. In another aspect, a kit including a system described herein and instructions for using the system for recording and decoding brain electrical signal data associated with an attempted action by a subject. BRIEF DESCRIPTION OF THE FIGURES FIG. 1 illustrates an overview of a multimodal speech decoding pipeline in a participant with vocal-tract paralysis in accordance with an embodiment of the invention. FIGS. 2A to 2C depict multimodal speech decoding in a participant with vocal-tract paralysis in accordance with an embodiment of the invention. (A) a sagittal MRI showing brainstem atrophy (in the bilateral pons; red arrow) resulting from stroke. (B) an MRI reconstruction of the participant’s brain overlaid with the locations of implanted electrodes. (C) an example of ECoG features resulting from attempted orofacial movements. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 FIGS.3A to 3E depict high-performance text decoding from neural activity in accordance with an embodiment of the invention. (A) is a schematic diagram of the text-decoding algorithm. (B) depicts median phone error rates with the 1024-word-General sentence set. (C) depicts word error rates for chance and real-time results. (D) depicts character error rates for chance and real- time results. FIGS. 4A to 4B depict results of high-performance text decoding from neural activity in accordance with an embodiment of the invention. (A) depicts offline evaluation of error rates as a function of number of recording days and data quantity. (B) depicts real-time classification accuracy during attempts to silently say 26 NATO code words across many recording days. FIG. 5 is a schematic diagram of the speech-synthesis decoding algorithm in accordance with an embodiment of the invention. FIG. 6 depicts intelligible speech synthesis from neural activity in accordance with an embodiment of the invention. The top provides three example decoded spectrograms and waveforms from the 529-phrase-AAC sentence set. The bottom provides the corresponding reference spectrograms and waveforms representing the decoding targets. FIGS. 7A to 7C depict results of intelligible speech synthesis from neural activity in accordance with an embodiment of the invention. (A) depicts Mel-cepstral distortions for the decoded waveforms. (B) shows perceptual word error rates from untrained human evaluators via a transcription task. (C) shows perceptual character error rates from the same human-evaluation results as (B). FIG. 8 is a schematic diagram of the avatar decoding algorithm in accordance with an embodiment of the invention. FIGS. 9A to 9B depict results of direct decoding of orofacial articulatory gestures from neural activity to drive an avatar in accordance with an embodiment of the invention. (A) binary perceptual accuracies from human evaluators on avatar animations generated from neural activity. (B) correlations for jaw, lip, and mouth-width movements between decoded avatar renderings and videos of real human speakers on the 1024-word-General sentence set. FIGS. 10A to 10B depict results of direct decoding of orofacial articulatory gestures from neural activity to drive an avatar in accordance with an embodiment of the invention. (A) top: snapshots of avatar animations of 6 non-speech articulatory movements in the articulatory- movement task; bottom: confusion matrix depicting classification accuracy across the movements. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 (B) top: snapshots of avatar animations of 3 non-speech emotional expressions in the emotional- expression task; bottom: confusion matrix depicting classification accuracy across 3 intensity levels (high, medium, and low) of the 3 expressions, ordered via hierarchical agglomerative clustering on the confusion values. FIGS. 11A to 11B depict articulatory encodings driving speech decoding. (A) is a mid- sagittal schematic of the vocal tract with phone place of articulation (POA) features labeled. (B) bottom-right: visualization of the locations of electrodes with the greatest encoding weights for labial, front-tongue, and vocalic phones on the electrocorticography array, the electrodes that most strongly encoded finger flexion during the NATO-motor task are also included. FIGS. 12A to 12D depict articulatory encodings driving speech decoding. (A) illustrates phone-encoding vectors for each electrode computed by a temporal receptive-field model on neural activity recorded during attempts to silently say sentences from the 1024-word-General set, organized by unsupervised hierarchical clustering. (B) provides Z-scored POA encodings for each electrode, computed by averaging across positive phone encodings within each POA category. (C) and (D) provide projection of consonant (C) and vowel (D) phone encodings into a two- dimensional space via multidimensional scaling (MDS). FIGS. 13A to 13C provide electrode-tuning comparisons between front-tongue phone encoding and tongue-raising attempts (A), labial phone encoding and lip-puckering attempts (B), and tongue-raising and lip-rounding attempts (C). FIGS. 14A to 14C depict effects of anatomical coverage, electrode density, and feature sets on decoding performance. (A) depicts MRI reconstruction of the participant’s brain overlaid with the locations of implanted electrodes. (B) and (C) provide the effect of excluding each region during training and testing on text-decoding word error rates (B) and NATO code-word classification accuracies (C). FIGS. 15A to 15C depict effects of anatomical coverage, electrode density, and feature sets on decoding performance. (A) is a visualization of the checkerboard-downsampling procedure used to simulate a low-density electrocorticography array with 127 electrodes instead of 253. (B) and (C) provide the effect of modulating electrode density and feature set on text-decoding word error rates (B) and NATO code-word classification accuracies (C). Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 FIG.16 provides results of simulated text decoding with a larger vocabulary in accordance with an embodiment of the invention. Text decoding results were simulated using a 42,391-word vocabulary on the blocks used for real-time evaluation with the 1024-word-General set. FIG.17 provides results of simulated text decoding on the 50-phrase-AAC sentence set in accordance with an embodiment of the invention. Text decoding results were simulated on the real-time blocks used for evaluation with the synthesis models. FIG. 18 provides results of simulated text decoding on the 529-phrase-AAC sentence set in accordance with an embodiment of the invention. Text decoding results were simulated on the real-time blocks used for evaluation of the 529-phrase-AAC sentence set with the synthesis models. FIG. 19 provides Mel-cepstral distortions using a personalized voice tailored to the participant in accordance with an embodiment of the invention. The Mel-cepstral distortion was calculated between decoded speech with the participant’s personalized voice and reference waveforms for the 529-phrase-AAC, 50-phrase-AAC, and 1024-word-General set. FIG. 20 provides examples of directly decoded articulatory gestures (colored) compared with reference articulatory gestures (black) in accordance with an embodiment of the invention. Examples were taken from the 50-phrase-AAC sentence set. FIG. 21 provides correlations of directly decoded avatar articulatory gestures with reference articulatory gestures in accordance with an embodiment of the invention. FIG. 22 provides correlations between avatar articulatory gestures with reference articulatory gestures using the acoustic approach in accordance with an embodiment of the invention. FIG. 23 illustrates binary perceptual accuracy from human evaluation of silent videos extracted from the audio-visual synthesis task in accordance with an embodiment of the invention. FIG. 24 provides correlations of facial landmark trajectories within healthy speakers and between healthy speakers and the avatar using the acoustic approach for avatar decoding the 1024- word-General sentence set or in accordance with an embodiment of the invention. FIG. 25 provides classification results of emotional expressions. Full 15-fold cross validation classification accuracy for emotional expressions across different subsets of intensities in the emotional-expression task are depicted. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 FIG.26 illustrates cross-phone place-of-articulation encoding. For each electrode included in FIG. 11B the relationship between encoding of phone place of articulation (POA) categories was visualized. FIGS. 27A to 27D demonstrate spatial distribution of electrode tuning to articulatory features. Shown are normalized [0-1] encoding weights across electrodes for (A) hand finger flexion, (B) labial phones, (C) front tongue phones, and (D) vocalic phones. FIGS. 28A to 28B demonstrate that attempted finger flexion and speech are largely encoded orthogonally. (A) for each electrode, the normalized [0,1] encoding in response to attempted production of NATO code-words is plotted against attempted finger flexion in the NATO-motor task. (B) confusion matrix from the NATO-motor task, showing minimal confusion between hand and speech targets. FIG. 29 illustrates a virtual environment for avatar decoding in accordance with an embodiment of the invention. FIGS.30A to 30C provide examples of dlib facial-landmark detection in accordance with an embodiment of the invention. (A) an example of detected facial landmarks overlaid on a frame selected from a video rendering of an avatar during a single trial and using the direct approach to avatar decoding. (B) an example set of plotted detected facial landmark key points from a video rendering of an avatar during a single trial and using the direct approach to avatar decoding. (C) an example set of plotted facial landmark key points that shows exemplar facial landmark detection of a human face. FIG. 31 depicts the distribution of phone-encoding r-values across electrodes. Shown are encoding r-values across electrodes from the linear encoding model trained to predict each electrode’s high gamma activity from phoneme emission probabilities in accordance with an embodiment of the invention. FIG.32 provides a flow diagram depicting a method of controlling a fully embodied virtual avatar using recorded brain electrical signal data in accordance with an embodiment of the invention. FIG. 33 provides a flow diagram depicting a method of controlling and displaying a fully embodied virtual avatar using recorded brain electrical signal data in accordance with an embodiment of the invention. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 FIG.34 provides a flow diagram depicting a method for training a machine learning model to predict avatar movements using recorded neural activity in accordance with an embodiment of the invention. FIG. 35 illustrates an overview of a non-speech communicative gesture neural-decoding pipeline in accordance with embodiments of the invention. FIG. 36 provides a block diagram of a pipeline for concurrent multi-effector neural- decoding in accordance with embodiments of the invention. FIG. 37 provides a block diagram of a pipeline for non-verbal linguistic neural-decoding in accordance with embodiments of the invention. FIGS. 38A to 38B depict an overview of a naturalistic streaming silent-speech neuroprosthesis in accordance with an embodiment of the invention. (A) an overview of the streaming speech-synthesis and text-decoding pipeline. (B) an example of an online waveform (top) and spectrogram (bottom) from the 1024-word-General set. FIGS. 39A to 39F provide results of online continuously streaming synchronized speech synthesis and text decoding from neural activity in accordance with an embodiment of the invention. (A) latency for speech-synthesis and text-decoding. (B) synchronization time between the speech-synthesis onset and text-decoding onset. (C) synthesized words per minute compared to delayed synthesis. (D) phoneme error rates. (E) word error rates. (F) character error rates. FIGS. 40A to 40E illustrate offline long-form continuous speech decoding with implicit speech detection. (A) Top: a heatmap of log-scaled high-gamma activity (HGA) from the top 20 most speech-responsive electrodes during silent speech attempts of an entire block (5.9 minutes) of 1024-word-General sentences. Bottom: a continuously synthesized speech waveform from the aforementioned neural activity. (B) latency between the detected onset of silently attempted speech to synthesized speech output onset and latency between the detected offset of silently attempted speech to synthesized speech output offset. (C) phoneme error rates. (D) word error rates. (E) character error rates. FIGS.41A to 41C illustrate speech synthesis generalization across silent-speech interfaces in accordance with an embodiment of the invention. (A) phoneme error rates. (B) word error rates. (C) character error rates. FIGS.42A to 42D demonstrate model-generated auditory feedback does not interfere with articulatory-driven speech decoding. (A) placement of the electrodes on the speech sensorimotor Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 cortex. (B) contribution maps calculated from two conditions: blocks with auditory feedback during online speech-synthesis demonstrations (left) and blocks without decoder feedback (right). (C) contribution comparison for each channel, colored by anatomical region. (D) for both speech and text, there is no significant difference in decoding performance between conditions. FIG.43A to 43C provides decoding accuracy results for real-time text-to-speech decoding using the 1024-word-General sentence set in accordance with an embodiment of the invention. FIG. 44 provides latency results from go-cue to speech-decoding in accordance with an embodiment of the invention. FIGS. 45A to 45C provide results characterizing the latency of the speech synthesis and text decoding system in accordance with an embodiment of the invention. (A) latency per time step for the neural encoder, speech joiner and beam search, speech synthesizer, text joiner and beam search, and complete system. (B) latency by module. (C) success rate averaged across time for all trials. FIGS. 46A to 46C provide region-exclusion analysis for 1024-word-General decoding in accordance with an embodiment of the invention. (A) phoneme error rates. (B) word error rates. (C) character error rates. FIGS. 47A to 47C provide decoding accuracy results by the length of training data for models in accordance with embodiments of the invention. (A) phoneme error rates. (B) word error rates. (C) character error rates. FIGS. 48A to 48C provide real-time decoding performance of speech synthesis using predicted transcripts gathered from perceptual evaluations or automatic speech recognition using in accordance with embodiments of the invention. (A) phoneme error rates. (B) word error rates. (C) character error rates. FIGS.49A to 49L demonstrate reading and listening do not interfere with speech decoding using methods in accordance with embodiments of the invention. (A) chronically implanted ECoG grids in two participants. (B) a speech decoding system. (C) the decoding framework. (D) the false positive rate (FPR) of the full speech decoding system. (E) the true positive rate (TPR) of the full speech decoding system. (F) electrode contributions for the speech-detection model for Bravo-1 and Bravo-3. (G) the false positive and negative rates as a function of the probability threshold used in the speech verification classifier. (H) accuracy of the 10-word classifier during baseline and listening blocks across 6 pseudo-blocks. (I) electrode contributions for the 10-word classifier Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 for Bravo-1 and Bravo-3. (J) scatter plot comparing electrode contributions for the 10-word classifier during online evaluations with the listening distractor versus without. (K) FPR of the decoding system when using the full system and a speech-detection model only trained on speech. (L) the number of false positives using the full system and a speech-only speech detection model during long periods of listening and reading. FIGS. 50A to 50M depict shared and distinct cortical activations for reading, listening, and attempted speech. (A) electrode heat maps for responsiveness in each task measured by non- parametric tests of pre-trial vs post go-cue mean high-gamma amplitude (HGA). (B) electrodes that have task modulation for reading, listening, and attempted speech visualized alongside electrodes that encode movements of the vocal-tract articulators during continuous speech and hand in Bravo-3. (C)-(E) Two-sided Wilcoxon rank-sum test for each electrode are plotted for reading and speech (C) speech and listening (D), and listening and reading (E). (F) example evoked response potentials (ERPs) for an electrode in Bravo-3 that has significant task modulation across attempted speech, reading, and listening. (G) the mean high-gamma amplitude, theta power, and beta power during reading, listening, and attempted speech across electrodes with significant task modulation for listening, reading, and attempted speech. (H) classification accuracy for the speech-verification model. (I) electrode contributions to the full speech-verification model in Bravo-3. (J) a confusion matrix for predictions from the full speech-verification model in Bravo- 3. (K)-(M) the same as (H)-(J) for Bravo-1. FIGS. 51A to 51D illustrate distinct representations for attempted speech, listening, and reading on speech cortex. (A) classification accuracy by anatomical region on the isolated-word set during reading, listening, and attempted speech. (B) for the Bravo-3 temporal-lobe attempted speech and listening classification models, accuracy is shown for evaluating models across tasks. (C) electrode contributions for Bravo-3 temporal-lobe attempted speech and listening classification models. (D) scatterplot of electrode contributions from (C) across temporal-lobe electrodes. FIGS. 52A to 52E demonstrate the role for shared electrodes in speech-motor planning. (A) example evoked response potentials (ERPs) from two electrodes in Bravo-3 that are task- modulated during reading. (B) average HGA for each reading-responsive electrode from the isolated-word set and false fonts. (C) average HGA for each reading-responsive electrode from the isolated-word set and sentences. (D) electrode contributions for attempted-speech models, Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 trained on different windows of time around the go-cue. (E) average electrode contribution for postcentral, precentral, and temporal-lobe electrodes in each decoding window. FIGS. 53A to 53D illustrate the effect of speech decoding system ablations on false positive and negative performance. (A) the false positive rate (FPR) of the system with only a speech-detection model trained only on attempted speech, the full model, and the full speech decoding system during baseline, listening, and reading blocks. (B) the true positive rate (TPR) for each of the system ablations in (A) during baseline, listening, and reading blocks. (C) the absolute number of false positives for each of the system ablations during “long” distractor blocks of listening and reading, where there were no speech attempts and only the distractor. (D) the number of time points (at 200 Hz) that had a speech probability greater than the probability threshold of 0.5 for each of the system ablations and “long” distractor blocks. FIG.54 depicts speech-verification model performance for different time windows around detected events for models in accordance with embodiments of the invention. FIGS.55A to 55C illustrate temporal dynamics of shared activity during attempted speech, listening, and reading. (A) mean evoked response potentials (ERPs) for attempted-speech, listening, and reading. (B) the maximum HGA evoked by attempted speech, reading, and listening for each tri-function electrode. (C) the evoked HGA 500ms before the onset / go-cue for attempted speech, reading, and listening for each tri-function electrode. FIG. 56 provides electrodes with comparable HGA for reading words and false fonts localized around the frontal eye fields in accordance with embodiments of the invention. FIG. 57 provides anatomical characteristics of electrodes important for decoding attempted speech pre go-cue and after the go-cue in accordance with embodiments of the invention. FIG. 58 illustrates a sample artifact that occurred during three trials of online evaluation with Bravo-1in accordance with embodiments of the invention. FIGS. 59A to 59G demonstrate the implementation of a bilingual speech neuroprosthesis in accordance with embodiments of the invention. (A) schematic diagram of the bilingual decoding system. (B) word error rates with the phrase test set, calculated using shuffled neural data, neural decoding from the RNN without language modeling, and the full online system with language modeling. (C) language classification accuracy for chance, neural-only, and online results. (D) the decoding rate compared to the participant’s communication speed with his alternative Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 augmentative communication (AAC) strategy. (E) the language-classification accuracy as a function of word position in a phrase. (F) phrase likelihood scores from GPT2 (large language model) for trials where the language is correctly and incorrectly classified. (G) word error rates, as in (B) when the target language is manually set rather than freely decoded. FIGS. 60A to 60G provide offline characterizations of the bilingual classification algorithms in accordance with embodiments of the invention. (A) 10-fold cross-validation classification accuracy for English words, Spanish words, and across both languages. (B) classification of words in English, Spanish, and across both languages for 48 days without retraining or recalibration of the system. (C) classification performance before (n=5 days) and after (n=5 days) a 30-day break in recording without retraining. (D) electrode contributions for models trained only on English or Spanish words, separated by neural-feature type (HGA and LFS). (E) relationship between 128 HGA (left) and LFS (right) electrode contributions for Spanish and English models. (F) selected portion of the confusion matrix between bilingual words, highlighting confusability. (G) multiple regression models were fit to predict confusability between a pair of words from their acoustic similarity, semantic similarity, and whether the words are in the same language. FIGS. 61A to 61J demonstrate shared articulatory representations in speech-motor cortex across languages. (A) large stimulus set of unique words and phrases used to cover a larger articulatory space in each language, relative to the vocabulary used for core phrase-decoding (FIG. 59). (B) standard deviation of the average high-gamma amplitude (HGA) from 0 to 2 seconds (relative to the visual go cue) for English and Spanish phrases across each electrode. (C) sample evoked response potentials (ERPs) to English and Spanish phrases for two electrodes noted in (B). (D) relationship between the maximum HGA for each electrode during English and Spanish phrases. (E) relationship between the HGA standard deviation (as in (B)) for English and Spanish phrases for each electrode. (F) relationship of the correlation of ERPs within a language to the correlation of ERPs between languages for each electrode. (G) 10-fold cross-validation (CV) classification accuracy for classifying each phrase as English or Spanish. (H) stimulus set designed to probe a shared articulatory, syllabic representation between languages. (I) sample ERPs from an electrode indicated in (B) for different syllables in the same language and a shared syllable in different languages. (J) 10-fold CV syllable-classification accuracy across training and testing paradigms. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 FIGS. 62A to 62F demonstrate rapid transfer learning between languages in accordance with embodiments of the invention. (A) schematic depiction of the paradigm used to evaluate transfer learning between languages. (B) learning curves for fine-tuning and evaluating on a new Spanish vocabulary. (C) learning curves for fine-tuning and evaluating on a new English vocabulary. (D) schematic depiction of the paradigm used to evaluate the effect of acoustic similarity between the train and fine-tune set on transfer learning efficacy. (E) mel-cepstral distortion (MCD, as in FIG. 60G) between each word in the acoustically similar or different training sets with the corresponding word in the fine-tune / test set. (F) learning curves for fine- tuning and evaluating on the “Fine-tune and test set,” defined in (D), with transfer learning from the acoustically similar and different models. FIG. 63 illustrates the effect of the amount of pretrain data on transfer learning efficacy for models in accordance with embodiments of the invention. FIG.64 illustrates the effect of pre-training with silently attempted speech data for models in accordance with embodiments of the invention. FIG. 65 demonstrates the performance of an attempted speech model using windows of various input lengths in accordance with embodiments of the invention. FIG.66 provides a comparison of all online-evaluation sentences and the AAC evaluation subset in accordance with embodiments of the invention. DETAILED DESCRIPTION Methods of assisting individuals with communication are provided. In the disclosed methods, cortical activity from a region of the brain associated with movement, speech production, and / or language perception is recorded while an individual attempts to perform an action (e.g., to say words, express an emotion, perform a movement, etc.). Deep learning computational models are used to detect and decode the attempted action from the recorded brain activity. Decoding of actions from brain activity may be aided by the use of self-supervised machine learning techniques, which discretize each action into one or more action representations that serve as an effective intermediary for decoding neural activity patterns and features into meaningful action outputs. In addition, methods for synthesizing decoded actions into audio and / or visual stimuli are provided, allowing for more naturalistic and expressive communication for individuals who are unable to speak or experience other mobility limitations that inhibit full embodied communication. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 Before the present invention is described in greater detail, it is to be understood that this invention is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims. Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention. Certain ranges are presented herein with numerical values being preceded by the term “about.” The term “about” is used herein to provide literal support for the exact number that it precedes, as well as a number that is near to or approximately the number that the term precedes. In determining whether a number is near to or approximately a specifically recited number, the near or approximating unrecited number may be a number which, in the context in which it is presented, provides the substantial equivalent of the specifically recited number. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, representative illustrative methods and materials are now described. All publications and patents cited in this specification are herein incorporated by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference and are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed. It is noted that, as used herein and in the appended claims, the singular forms “a”, “an”, and “the” include plural referents unless the context clearly dictates otherwise. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely,” “only” and the like in connection with the recitation of claim elements, or use of a “negative” limitation. As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present invention. Any recited method can be carried out in the order of events recited or in any other order which is logically possible. While the apparatus and method has or will be described for the sake of grammatical fluidity with functional explanations, it is to be expressly understood that the claims, unless expressly formulated under 35 U.S.C. §112, are not to be construed as necessarily limited in any way by the construction of “means” or “steps” limitations, but are to be accorded the full scope of the meaning and equivalents of the definition provided by the claims under the judicial doctrine of equivalents, and in the case where the claims are expressly formulated under 35 U.S.C. §112 are to be accorded full statutory equivalents under 35 U.S.C. §112. DEFINITIONSThe term “communication” includes word-based communication such as verbal communication including spoken speech, spelling of words, and production of text (e.g., controlling a personal device to generate email or text via attempts to speak) as well as action- based communication such as through attempted non-speech motor movement. Attempted speech may include vocalized speech, which may or may not be intelligible, or non-vocalized speech. Silent-speech attempts are volitional attempts to articulate speech without vocalizing. Attempted non-speech motor movement may include imagined movement without any detectable physical movement. Attempted non-speech motor movements may include, without limitation, imagined head, arm, hand, foot, and leg movements. Attempted non-speech motor Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 movements may be used to indicate the initiation or termination of attempted speech or an attempted action or to control an external device (e.g., for communication with a personal device or software applications or to turn on or off a device). In the disclosed methods, neural activity is recorded during attempted actions whether or not the individual produces any vocal output or detectable motor movement. The term “communication disorders” is used herein to refer to a group of conditions that affect the ability of a subject to communicate (e.g., using word-based and / or action-based communication). Communication disorders include, without limitation, anarthria, strokes, traumatic brain injuries, brain tumors, amyotrophic lateral sclerosis, multiple sclerosis, Huntington's disease, Niemann-Pick disease, Friedreich's ataxia, Wilson's disease, cerebral palsy, Guillain-Barré syndrome, Tay-Sachs disease, encephalopathy, central pontine myelinolysis, and other conditions causing dysfunction or paralysis of the muscles of the head, neck, or chest, or arms resulting in anarthria. The terms “subject”, “individual”, “patient”, and “participant” are used interchangeably herein and refer to a patient having a communication disorder. The patient is preferably human, e.g., a child, an adolescent, an adult, such as a young, middle-aged, or elderly human who may benefit from the systems, devices, and methods disclosed herein for restoring communication. The patient may have been diagnosed as having anarthria or may be paralyzed. The term “user” as used herein refers to a person that interacts with a device and / system disclosed herein for performing one or more steps of the presently disclosed methods. The user may be the patient receiving treatment. The user may be a health care practitioner, such as the patient’s physician. METHODS As summarized above, methods of assisting individuals with communication are provided. In the disclosed methods, cortical activity from a region of the brain associated with movement, speech production, and / or language perception is recorded while an individual attempts to perform an action (e.g., to say words, express an emotion, perform a movement, etc.). Deep learning computational models are used to detect and decode the attempted action from the recorded brain activity. In some embodiments, decoding of actions from brain activity is aided by the use of self- supervised machine learning techniques, which discretize each action into one or more action Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 representations that serve as an effective intermediary for decoding neural activity patterns and features into meaningful action outputs. In addition, methods for synthesizing decoded actions into audio and / or visual stimuli are provided, allowing for more naturalistic and expressive communication for individuals who are unable to speak or experience other mobility limitations that inhibit full embodied communication. Attempted Actions and Decoded Actions As described above, embodiments of the methods include recording cortical activity from a region of the brain associated with movement, speech production, and / or language perception while a subject attempts to perform an action. The action attempted by the subject may be an action the subject is unable to perform without aid, or an action the subject has difficulty performing without aid. Attempted actions may include imagined movement of one or more body parts without any detectable physical movement or imagined speech without any discernable noise. Neural activity is recorded during attempted actions whether or not the individual produces any vocal output or detectable motor movement. In some embodiments, the attempted action includes attempted speech. In these instances, the subject may have difficulty articulating intelligible words. In some embodiments, the subject has a has a low intelligibility for speaking prompted words or sentences (e.g., as measured by speech perception testing (SPT)). For example, the subject may have a speech intelligibility of 87% or less for prompted words, or 78% or less, or 67% or less, or 50% or less, or 10% or less, or 5% or less, or 0%. In some cases, the subject may have a speech intelligibility of 95% or less for prompted sentences, or 89% or less, or 50% or less, or 10% or less, or 5% or less, or 0%. In some embodiments, the subject may have difficulty articulating intelligible words, or may have anarthria, as the result of experiencing a disease or condition. In these instances, the disease or condition may be, but is not limited to, any condition or disease causing dysfunction or paralysis of the muscles of the head, neck, or chest, or arms resulting in anarthria. For example, the disease or condition may be, or may result from, strokes, traumatic brain injuries, brain tumors, amyotrophic lateral sclerosis, multiple sclerosis, Huntington's disease, Niemann-Pick disease, Friedreich's ataxia, Wilson's disease, cerebral palsy, Guillain-Barré syndrome, Tay-Sachs disease, encephalopathy, central pontine myelinolysis, etc. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 In some embodiments, the attempted action includes an attempted movement. In some embodiments, the attempted movement may be an orofacial movement for speech or a non-speech related movement. In these instances, the subject may have difficulty performing the movement. In some embodiments, the subject may have no ability to perform the movement as, e.g., the body part used to perform the movement is missing or is completely paralyzed. In some embodiments, the subject may have difficulty performing the movement, or may be completely unable to perform the movement, as the result of experiencing a disease or condition. In these instances, the disease or condition may be, but is not limited to, any condition or disease reducing the mobility of one or more body parts of the subject. For example, the disease or condition may be, or may result from, strokes, traumatic brain or spinal injuries, amputations, birth defects, cerebral palsy, Friedreich's ataxia, Guillain-Barré syndrome, Lyme disease, spina bifida, arthritis, tendonitis, a tendon or myotendinous tear, a hernia, old age, chronic health problems, etc. In some instances, the subject may experience diplegia, hemiplegia, monoplegia, paraplegia, or quadriplegia. As described above, the action attempted by the subject may be an action the subject is unable to perform without aid or assistance, or an action the subject has difficulty performing without aid or assistance. In embodiments where the attempted action includes speech, the speech may include a single speech sound (i.e., a single phoneme), a single word, or a single sentence. In other cases, the speech may include multiple speech sounds, words, or sentences. For example, the speech may include a phrase made up of 2 or more words, or 5 or more words, or 10 or more words, or 20 or more, or 50 or more, or 100 or more. The words and / or phrases may be in any language. For example, the words may be in English, Spanish, Mandarin, Dutch, Swahili, Hindi, etc. In some embodiments, spoken phrases or sentences may include words in multiple different languages. In these cases, the words of the phrases / sentences may switch between languages in any manner, e.g., mirroring the nature of a bilingual conversation. In embodiments where the attempted action includes a movement, the attempted action may include a single movement. For example, the attempted action may consist of the abduction, adduction, flexion, extension, and / or circumduction of a body part or a single orofacial movement. In other cases, the attempted action may include multiple movements. For example, the attempted action may include making a facial expression, shooting a basketball, opening a door, walking, hand gestures including multiple finger flexions, a shoulder shrug, a dance move, a hug, etc. In these instances, the attempted action may Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 include 2 or more movements, or 3 or more movements, or 5 or more, or 10 or more, or 15 or more, or 20 or more. In some embodiments, the attempted action may be selected from a limited group or set of actions. The number of actions included is preferably large enough to create a meaningful variety of actions but small enough to enable satisfactory neural-based classification performance. In embodiments where the attempted action includes speech, the set of actions may be a set of words or phrases. In other words, the attempted speech may include one or more words from a set of words or one or more phrases from a set of phrases. In some instances, the set of words may include 50 words or more, or 100 words or more, or 300 words or more, or 500 words or more, or 1,000 words or more, or 5,000 words or more, or all the words of a specific languages (e.g., English, Spanish, Mandarin, etc.). In some cases, the set of phrases may include 50 phrases or more, or 100 phrases or more, or 300 phrases or more, or 500 phrases or more, or 1,000 phrases or more. In some embodiments, the word or phrase set is adapted or configured for a specific use. For example, the word or phrase set may include words or phrases useful for expressing basic emotions, communicating caregiving needs, discussing a specific interest or hobby (e.g., a sport, a fandom, a videogame, music, etc.), discussing a profession (e.g., accounting, a field of scientific research, law, finance, etc.), or interacting with other individuals in a specific context (e.g., shopping at a store, interacting in a specific videogame or metaverse, etc.). In some cases, the word or phrase set may be adapted or configured for a specific individual based on the linguistic ability of the individual. For example, word or phrase sets including words from multiple different languages may be created for individuals having the ability to communicate in multiple different languages (such as, e.g., a word set including both English and Spanish words for a bilingual speaker of both languages). In some instances, the word or phrase set may be adapted or configured based on the region an individual is from and / or based on the vocabulary of an individual (i.e., the number of words known to the individual). In some embodiments, the subject is limited to 2 or more word or phrase sets, such as 5 or more, or 10 or more, or 20 or more, or 50 or more, or 100 or more, or 1,000 or more. In some cases, additional word or phrase sets may be created for the subject as needed. In some instances, the subject is able to switch between word or phrase sets as desired. For example, the subject may switch from a word set adapted for communicating caregiving needs to a word set adapted for playing a specific video game when the subject begins playing the video game. In some cases, a Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 specific non-speech related movement may be used to indicate a switch between word or phrase sets, followed by, e.g., the attempted speech of a word associated with a specific word or phrase set to be switched to. In some instances, a switch between word sets may be automatically initiated based on the context of a conversation. As discussed above, the attempted action may be selected from a limited group or set of actions. In embodiments where the attempted action includes a movement, the set of actions may be a set of movements. In other words, the attempted action may include one or more movements from a set of movements (i.e., each action in the action set may include one or more movements). In some instances, the set of movements may include 3 movements or more, or 6 movements or more, or 9 movements or more, or 20 movements or more, or 50 movements or more, or 100 movements or more, or 1,000 movements or more. In some embodiments, the movement set is adapted or configured for a specific use. For example, the movement set may include movements useful for expressing basic emotions, communicating caregiving needs, performing tasks associated with a specific interest or hobby (e.g., a sport, a videogame, music, etc.), performing tasks associated with a profession (e.g., accounting, lab work, law, finance, etc.), or interacting with other individuals in a specific context (e.g., shopping at a store, interacting in a specific videogame or metaverse, etc.). In some cases, the movement sent may be adapted or configured to communicate a non-verbal language. For example, the movement set may be configured to communicate via American Sign Language (ASL) and, as such, may include movements associated with specific words or phrases in ASL. In these instances, the movement set may be adapted to include specific word or phrase associated movements in a similar manner as the word or phrase sets for attempted speech discussed above (i.e., based on the linguistic abilities of an individual, specific conversational contexts, etc.). In some embodiments, the subject is limited to 2 or more movement sets, such as 5 or more, or 10 or more, or 20 or more, or 50 or more, or 100 or more, or 1,000 or more. In some cases, additional movement sets may be created for the subject as needed. In some instances, the subject is able to switch between movement sets as desired. For example, the subject may switch from a movement set adapted for communicating caregiving needs to a movement set adapted for playing a specific video game when the subject begins playing the video game. In some cases, a specific movement that is not included in any movement set, or not included in the presently selected movement set, may be used to indicate a switch between movement sets, followed by, e.g., the Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 attempted speech of a word associated with a specific movement set to be switched to. In some instances, a switch between movement sets may be automatically initiated based on the context of a conversation. In some embodiments, multiple different attempted action sets (i.e., word or phrase sets and / or movement sets) may be combined, e.g., such that one or more body parts of a subject may simultaneously be limited via a plurality of different attempted action sets. In some embodiments, a single part of the body may be limited via multiple different attempted action sets. For example, the hands of an individual may simultaneously be limited to attempting to communicate via ASL and attempting to interact with a specific virtual or real-world environment. In some instances, different parts of the body may be limited via different attempted action sets. For example, the legs of an individual may be limited to attempting to walk and / or run while the facial features of an individual may be limited to attempting to display specific emotions. In some embodiments, the attempted action is performed in order to synthesize audio, generate text, or control an electronic output device or system to perform one or more actions. By electronic output device or system is meant any device or system capable of being controlled using electrical signals. For example, the device may be a prosthetic limb having an electric motor, a visual display (e.g., a computer monitor, a television, a virtual or augmented reality headset, etc.), a program or software configured to control one or more visual displays, a robot, a speaker, etc. In some embodiments, the electronic output device or system includes a computer system and the attempted action performed by the subject controls a cursor (via, e.g., a virtual mouse), controls a keyboard (e.g., a virtual keyboard), and / or enters computer commands. The output device or system may receive electrical control signals through a variety of means. In some embodiments, the output device or system may receive electrical control signals directly, e.g., through a wire. In other embodiments, the output device or system may receive electrical control signals by converting an electromagnetic or ultrasound wave into an electronic control signal. For example, the output device or system may be controlled using Bluetooth®, Wi-Fi, cell phone towers and / or satellites (e.g., via Global System for Mobile communications (GSM) or Starlink), etc. In some cases, the electrical control signals are digital signals. In some embodiments, the electronic output device or system is controlled to perform a different action than the action attempted by the subject. For example, the attempted action performed by the subject may be a hand gesture and the action of the electronic output device or Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 system may be to power on or off. In some embodiments, attempted communicatory actions beyond or different from attempted speech (e.g., facial movements, non-facial movements, silence, hand gestures conveying sign language, etc.) are decoded into speech audio or text. In this way, any method of communication (e.g., overt, covert, silent, etc.) capable of being employed by a subject (e.g., a human) is encompassed within the invention. In some instances, attempted actions performed by the subject that are capable of being decoded quickly and / or are easily distinguishable from one another (i.e., via brain electrical signals using, e.g., the methods as described in greater detail below) may be used to control the one or more actions performed by the electronic output device or system. For example, hand movements that are easily distinguishable from attempted speech, and from other hand movements in a movement set (i.e., as described above), may be used to control a mobility scooter or a prosthetic arm. In some embodiments, the electronic output device or system is controlled to perform the same or a similar action as the action attempted by the subject. For example, the attempted action may be waving, and a prosthetic limb may be controlled to wave. In some embodiments, multiple simultaneous actions may be used to control the electronic output device or system. In these cases, some, all, or none of the simultaneously attempted actions may be the same or similar to the action the electronic output device or system is controlled to perform. For example, the attempted speech of a subject may be used to control the audible words (e.g., sound frequences) produced by a speaker while finger movements may be used to control the volume of the speaker. In some embodiments, the subject may be able to control multiple different electronic output devices or systems simultaneously. For example, the subject may be able to simultaneously generate noise via a speaker by attempting to speak and move a mobility scooter by attempting to point in a direction. As discussed above, the attempted action may be performed in order to synthesize audio or generate text. In embodiments where the attempted action includes speech, the attempted speech may be performed in order to generate text or synthesize speech audio such as, e.g., electronic speech audio (i.e., speech audio in a computer readable form). As discussed above, the attempted action may be performed in order to control an electronic output device or system to perform one or more actions. In some embodiments, the electronic output device or system includes an avatar such as, e.g., a humanoid avatar. In some instances, the avatar may be a physical avatar such as, e.g., an animatronic robot. In some Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 embodiments, the avatar is a virtual avatar. In these instances, the output device or system may include, but is not limited to, a visual display configured to present the avatar (and, e.g., actions performed by the avatar or text generated from attempted actions), a loudspeaker configured to produce sounds made by the avatar (such as, e.g., electronic speech audio synthesized using attempted speech by the subject, as described in greater detail below), and / or an avatar-animation system that includes a computer program for designing and / or animating the avatar. In these embodiments, the avatar may be controlled to perform the same action as the action attempted by the subject, or the same action as at least one of the simultaneously attempted actions of the subject. In some embodiments, the actions performed by the avatar include speech, orofacial movements for speech, and / or non-speech movements (e.g., non-speech communicative gestures). The speech may include any of the words or phrases as discussed above. The orofacial movements may include, but are not limited to, one or more of: a tongue tip raise, tongue retraction, tongue body raise, tongue advance, lip rounding, pinching nostril flare, upper lip pull, lower lip tuck, lower lip push, lower lip pull, lip flare, jaw opening, lip compression, and / or lip adduction. The non-speech movements may include, but are not limited to, the abduction, adduction, flexion, extension, and / or circumduction of one or more body parts. In some instances, the non-speech movements may include emotional expressions using facial muscles, such as, e.g., happy, sad, and / or surprised expressions. In these instances, the emotional expressions may include different levels of emotion or expression. For example, the emotional expressions may include extremely happy, very happy, and / or somewhat happy expressions. As discussed above, embodiments of the methods include recording cortical activity from a region of the brain associated with movement, speech production, and / or language perception while a subject attempts to perform an action. In some embodiments, the attempted action includes attempted speech. In these instances, the subject may have difficulty articulating intelligible words. In some embodiments, the attempted action includes an attempted movement. In these instances, the subject may have difficulty performing the movement. In some embodiments, the attempted action may be selected from a limited group or set of actions. In embodiments where the attempted action includes speech, the set of actions may be a set of words or phrases. In some instances, the set of words or phrases may include 1000 words or more. In embodiments where the attempted action includes a movement, the set of actions may be a set of movements. In some instances, the set of movements may include 9 movements or more. In some embodiments, the subject is limited Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 to 2 or more action sets. In some instances, the subject is able to switch between action sets as desired. In embodiments where the attempted action includes speech, the attempted speech may be performed in order to synthesize speech audio such as, e.g., electronic speech audio. In some embodiments, the attempted action may be performed in order to control an electronic output device or system to perform one or more actions. In some embodiments, the electronic output device or system includes an animated humanoid avatar. In some embodiments, the actions performed by the avatar include speech, orofacial movements for speech, and / or non-speech movements. In some instances, actions performed by the avatar are the same actions as the actions attempted by the subject. Brain electrical signal data associated with the attempted movement by the subject may be recorded and used to synthesize speech audio or control the humanoid avatar, as is discussed in greater detail below. Recording Brain Electrical Signals Embodiments of the methods include positioning a neural recording device including an electrode at a location in a region of the brain of the subject to record brain electrical signal data associated with an attempted action by the subject, e.g., as described above. Embodiments of the methods further include positioning an interface in communication with a computing device at a location on the head of the subject, wherein the interface is connected to the neural recording device. Embodiments of the methods further include recording the brain electrical signal data associated with the attempted action by the subject using the neural recording device, wherein the interface receives the brain electrical signal data from the neural recording device and transmits the brain electrical signal data to a processor of the computing device. As discussed above, embodiments of the methods include positioning a neural recording device including one or more electrodes at a location of the brain of the subject such as, e.g., in a sensorimotor cortex region, to record brain electrical signal data associated with an attempted action by the subject; and positioning an interface in communication with a computing device at a location on the head of the subject. Brain electrical signal data associated with attempted speech and / or an attempted movement by the subject is recorded using the neural recording device, wherein the interface receives the brain electrical signal data from the neural recording device and transmits the brain electrical signal data to a processor programmed to detect an attempted action by the subject and decode the detected action from the recorded brain electrical Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 signal data. In embodiments where the attempted action includes one or more movements, what the movement is (e.g., a finger flexion) and / or specifics of the movement (e.g., the intensity and duration of the movement) may be decoded from the recorded brain electrical signal data. In embodiments where the attempted action includes speech, one or more speech sounds (e.g., words or phrases) may be decoded from the recorded brain electrical signal data. The recording device may include non-brain penetrating surface electrodes and / or brain- penetrating depth electrodes. In some embodiments, the recording device may include non-brain penetrating surface electrodes. The electrical signals may be recorded using a single electrode, electrode pairs, or an electrode array. In some embodiments, the brain activity is recorded from more than one site. In certain embodiments, brain electrical signal data is recorded from a sensorimotor cortex region of the brain involved in speech processing such as the precentral gyrus, postcentral gyrus, posterior middle frontal gyrus, posterior superior frontal gyrus, or posterior inferior frontal gyrus region, or any combination thereof. In certain embodiments, brain electrical signal data is recorded from a sensorimotor cortex region of the brain involved in movement, such as a region associated with movement of a specific muscle group or area of the body. In some embodiments, the electrode is positioned on the pial surface of the sensorimotor cortex region of the brain. Positioning an electrode for recording brain activity at specified region(s) of the brain may be carried out using standard surgical procedures for placement of intra-cranial electrodes. As used herein, the phrases “an electrode” or “the electrode” refer to a single electrode or multiple electrodes such as an electrode array. As used herein, the term “contact” as used in the context of an electrode in contact with a region of the brain refers to a physical association between the electrode and the region. In other words, an electrode that is in contact with a region of the brain is physically touching the region of the brain. An electrode in contact with a region of the brain can be used to detect electrical signals corresponding to neural activity associated with attempted speech and / or an attempted movement. Electrodes used in the methods disclosed herein may be monopolar (cathode or anode) or bipolar (e.g., having an anode and a cathode). In certain embodiments, one or more electrodes are used to record electrical signals for neural activity associated with attempted speech and / or attempted movement in one or more brain regions. An electrode may be placed, for example, in a region of the sensorimotor cortex involved in speech processing such as the superior gyrus, the middle temporal gyrus, the Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 precentral gyrus, and / or the postcentral gyrus regions of the brain. In certain cases, placing the electrode may involve positioning the electrode on the surface of the specified region(s) of the brain. For example, electrodes may be placed on the surface of the brain at the superior gyrus, the middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus, or any combination thereof. The electrode may contact at least a portion of the surface of the brain at the superior gyrus, the middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus. In some embodiments, the electrode may contact substantially the entire surface area at the superior gyrus, the middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus regions of the brain. In some embodiments, the electrode may additionally contact area(s) adjacent to the superior gyrus, the middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus regions. In embodiments where the neural recording device includes an electrode array, the electrode array may be centered on the central sulcus. In some embodiments, an electrode array arranged on a planar support substrate may be used for detecting electrical signals for neural activity from one or more of the brain regions specified herein. The surface area of the electrode array may be determined by the desired area of contact between the electrode array and the brain. An electrode for implanting on a brain surface, such as, a surface electrode or a surface electrode array may be obtained from a commercial supplier. A commercially obtained electrode / electrode array may be modified to achieve a desired contact area. In some cases, the non-brain penetrating electrode (also referred to as a surface electrode) that may be used in the methods disclosed herein may be an electrocorticography (ECoG) electrode or an electroencephalography (EEG) electrode. In some embodiments, the non-brain penetrating electrode may be an ECoG electrode. In certain cases, placing the electrode at a target area or site (e.g., a neural recording device electrode) may involve positioning a brain penetrating electrode (also referred to as depth electrode) in the specified region(s) of the brain. For example, a depth electrode may be placed in a selected region of the sensorimotor cortex involved in speech processing (e.g., the superior gyrus, the middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus region) or involved with movement such as, e.g., the movement of a specific muscle group. In some embodiments, the electrode may additionally contact area(s) adjacent to the selected region of the sensorimotor cortex involved in speech processing (e.g., adjacent to the superior gyrus, the middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus region) or involved Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 with movement such as, e.g., the movement of a specific muscle group. In some embodiments, an electrode array may be used for recording electrical signals at the selected region of the sensorimotor cortex involved in speech processing (e.g., the superior gyrus, the middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus region) or involved with movement such as, e.g., the movement of a specific muscle group as specified herein. The depth to which an electrode is inserted into the brain may be determined by the desired level of contact between the electrode array and the brain and the types of neural populations that the electrode would have access to for recording electrical signals. A brain- penetrating electrode array may be obtained from a commercial supplier. A commercially obtained electrode array may be modified to achieve a desired depth of insertion into the brain tissue. In some cases, the electrode may be a non-brain penetrating electrode and the depth to which an electrode is inserted into the brain is negligible and / or is the minimum depth the electrode can be inserted to ensure stable contact with a region of the brain. The precise number of electrodes contained in an electrode array (e.g., for recording of neural activity associated with an attempted action) may vary. In certain aspects, an electrode array may include 2 or more electrodes, such as 3 or more, 10 or more, 50 or more, 100 or more, 200 or more, 250 or more, 500 or more, including 253 or more, e.g., about 6 to 12 electrodes, about 12 to 18 electrodes, about 18 to 24 electrodes, about 24 to 30 electrodes, about 30 to 48 electrodes, about 48 to 72 electrodes, about 72 to 96 electrodes, about 96 to 128 electrodes, about 128 to 196 electrodes, about 196 to 294 electrodes, about 294 to 440 electrodes, or more electrodes. The electrodes may be arranged into a regular repeating pattern (e.g., a grid, such as a grid with about 3 mm center-to-center spacing between electrodes), or no pattern. An electrode that conforms to the target site for optimal recording of electrical signals from neural activity associated with attempted speech and / or an attempted movement by a subject may be used. One such example, is a non-brain penetrating high-density electrode array that consists of 253 disk- shaped electrodes arranged in a lattice formation with 3-mm center-to-center spacing, wherein each electrode has a 1-mm recording-contact diameter and a 2-mm overall diameter. Another example is an electrode with multiple contacts. Yet further, another example of an electrode that can be used in the present methods branched electrode such as a is a 2 or 3 branched electrode to cover the target site. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 In some embodiments, a high-density ECoG electrode array is used to record electrical signals from neural activity associated with attempted speech and / or an attempted movement by a subject. For example, a high-density ECoG electrode array may include at least 100 electrodes, at least 128 electrodes, at least 196 electrodes, at least 253 electrodes, at least 294 electrodes, at least 500 electrodes, or at least 1000 electrodes, or more. In some embodiments, the electrode center-to-center spacing in a high-density ECoG electrode array ranges from 250 µm to 4 mm, including any electrode center-to-center spacing within this range such as 250 µm, 300 µm, 350 µm, 400 µm, 500 µm, 550 µm, 600 µm, 650 µm, 700 µm, 800 µm, 900 µm, 1 mm, 1.5 mm, 2 mm, 2.5 mm, 3 mm, 3.5 mm, or 4 mm. In some embodiments, a high-density ECoG micro- electrode array is used. ECoG micro-electrode arrays may include electrodes having a diameter of 250 µm or less, 230 µm or less, or 200 µm or less, including electrodes having a diameter ranging from 150 µm to 250 µm, including any diameter within this range such as 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, or 250 µm. For a description of high-density ECoG electrode arrays and micro-electrode arrays, see, e.g., Muller et al. (2015) Annu Int Conf IEEE Eng Med Biol Soc 2016:1528-1531; Chiang et al (2020) J. Neural Eng. 17:046008; Escabi et al. (2014) J. Neurophysiol. 112(6): 1566-1583; herein incorporated by reference. The size of each electrode may also vary depending upon such factors as the number of electrodes in the array, the location of the electrodes, the material, the age of the patient, and other factors. In certain aspects, each electrode has a size (e.g., a diameter) of about 5 mm or less, such as about 4 mm or less, including 4 mm-0.25 mm, 3 mm-0.25 mm, 2 mm-0.25 mm, 1 mm-0.25 mm, or about 3 mm, about 2 mm, about 1 mm, about 0.5 mm, or about 0.25 mm. In certain embodiments, the method further includes mapping the brain of the subject to optimize positioning of an electrode. In some embodiments, the positioning of an electrode is optimized to detect brain activity features associated with attempted speech and / or an attempted movement by the subject and to achieve optimal decoding of attempted speech or the attempted movement. For example, patterns of electrical signals in specific frequency ranges (e.g., alpha, delta, beta, gamma, and / or high gamma) may be used for detecting attempted speech and / or an attempted movement and decoding speech sounds and movements intended by the subject. Thus, electrodes may be positioned to optimize detection and / or decoding of brain activity in specific frequency ranges to restore communication to a subject who has a communication disorder or mobility to a subject who has a mobility disorder. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 In certain aspects, the methods and systems of the present disclosure may include recording brain activity, for example, electrical activity in the ventral sensorimotor cortex, where patterns of gamma-frequency neural activity associated with words, phrases, and sentences of attempted speech, orofacial movements associated with attempted speech, or movements of a specific muscle or muscle group may be detected. In certain cases, electrical activity from a plurality of locations in the sensorimotor cortex may be measured. In some embodiments, electrical activity in the high gamma frequency range (such as 70 Hz to 150 Hz) and / or the low frequency range (such as 0.3 Hz to 100 Hz) may be measured. In some embodiments, electrical activity in the high gamma frequency range (such as 70 Hz to 150 Hz) and the low frequency range (such as 0.3 Hz to 100 Hz) may be measured the superior gyrus, the middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus region, or any combination thereof. Detection of brain activity may be performed by any method known in the art. For example, functional brain imaging of neural activity may be carried out by electrical methods such as electrocorticography (ECoG), electroencephalography (EEG), stereoelectroencephalography (sEEG), magnetoencephalography (MEG), single photon emission computed tomography (SPECT), as well as metabolic and blood flow studies such as functional magnetic resonance imaging (fMRI), positron emission tomography (PET), functional near- infrared spectroscopy (fNIRS), and time-domain functional near-infrared spectroscopy. In some embodiments, the sensorimotor cortex (such as, e.g., a specific region of the sensorimotor cortex including one or more of the superior gyrus, the middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus region) is mapped to determine optimal positioning for electrodes to detect neural activity associated with attempted speech and / or an attempted movement by the subject. One or more regions may be implanted with a neural recording device including electrodes to measure electrical signals from neural activity associated with attempted speech and / or an attempted movement. In some cases, electrical activity in one or more locations in the brain may be measured not only during attempted speech or an attempted movement but also during a period extending from just prior to attempted speech or attempted movement (i.e., period of preparation for speech or movement) to a period just after attempted speech or movement (i.e., rest period after attempted speech or movement). Assessment of the accuracy of the decoding of speech or movement from neural activity at a particular site may be determined by comparing decoded Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 speech sounds to the intended speech sounds of the subject or comparing decoded movements to the intended movements of the subject. For example, the patient may communicate the correct intended sounds or movements using an assistive typing device. Both detection of the onset and offset of speech or movement events and speech sound and movement classification accuracy from decoding neural activity may be evaluated. False positives include detected speech or movement events that are not associated with a true speech sound or movement production attempt and false negatives include speech sound or movement production attempts that are not associated with a detected speech or movement event. Lower error rates in detection of speech or movement events and decoding of speech sounds or movement from neural activity indicate better performance. In certain cases, the placement of electrodes or the number of electrodes may be altered to improve detection of electrical signals and decoding of attempted speech and / or movement by the subject. Application of the methods may include a prior step of selecting a patient for implantation with a neural recording device based on need as determined by clinical assessment of the severity of the communication or mobility disorder and the desire for assistance with communication or mobility, and may also include cognitive assessment, anatomical assessment, behavioral assessment and / or neurophysiological assessment. Patients who have difficulty with communication may be implanted with a neural recording device to assist communication, as described herein. Embodiments of the methods may include implanting an interface capable of communication with a computing device in the head of the subject or placing the interface on the head of the subject to provide an externally accessible platform through which brain electrical signals can be acquired from the neural recording device and transmitted to a data processor for decoding. In some embodiments, the interface includes a percutaneous pedestal connector anchored in the cranium of the subject. The interface can be connected, for example, to a computing device such as a computer or a handheld computing device (e.g., cell phone or tablet) with a detachable digital connector and cable. Alternatively, the interface may be connected to a computing device wirelessly. In some embodiments, the interface includes a first wireless communication unit in communication with a computing device including a second wireless communication unit. In some embodiments, the first wireless communication unit utilizes a wireless communication protocol using an electromagnetic carrier wave (e.g., a radio wave, Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 microwave, or an infrared carrier wave) or ultrasound to transfer data from the interface to the computing device including the second wireless communication unit. Brain-computer interfaces are commercially available, including the Neuroport™ system from Blackrock Microsystems (Salt Lake City, Utah), See also, e.g., Weiss et al. (2019) Brain-Computer Interfaces 6:106-117; herein incorporated by reference. The processor may be provided by a computer or a handheld computing device (e.g., cell phone or tablet) programmed to decode the attempted speech and / or attempted movement from the recorded brain electrical signal data. As discussed above, a neural recording device may be used to record brain electrical signal data associated with an attempted action (e.g., attempted speech or an attempted movement) by the subject. In certain embodiments, a series of go cues is provided to the subject indicating when the subject should initiate attempted an attempted action. In some embodiments, the series of go cues are provided visually on a display. Each go cue may be preceded by a countdown to the presentation of the go cue, wherein the countdown for the attempted action is provided visually on the display and automatically started after each go cue. In some embodiments, the series of go cues are provided with a set interval of time between each go cue, which may be adjustable by the user. In certain embodiments, the processor is programmed to use the recorded brain electrical signal data within a time window following a go cue. In some embodiments, the processor is programmed to use the recorded brain electrical signal data within a time window before a go cue and following a go cue. As discussed above, embodiments of the methods include positioning a neural recording device including an electrode at a location in a region of the brain of the subject to record brain electrical signal data associated with an attempted action by the subject. Embodiments of the methods further include positioning an interface in communication with a computing device at a location on the head of the subject, wherein the interface is connected to the neural recording device, and recording the brain electrical signal data associated with the attempted action by the subject using the neural recording device, wherein the interface receives the brain electrical signal data from the neural recording device and transmits the brain electrical signal data to a processor of the computing device. In certain embodiments, brain electrical signal data is recorded from a sensorimotor cortex region of the brain involved in movement, such as a region associated with movement of a specific muscle group or area of the body. In certain Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 embodiments, brain electrical signal data is recorded from a sensorimotor cortex region of the brain involved in speech processing such as the precentral gyrus, postcentral gyrus, posterior middle frontal gyrus, posterior superior frontal gyrus, or posterior inferior frontal gyrus region, or any combination thereof. In some embodiments, the recording device may include a non-brain penetrating surface array of electrodes. The number of electrodes contained in the electrode array may be 250 or more electrodes. In some cases, the non-brain penetrating electrode array may be an electrocorticography (ECoG) electrode array. In some embodiments, electrical activity in the high gamma frequency range (such as 70 Hz to 150 Hz) and / or the low frequency range (such as 0.3 Hz to 100 Hz) may be measured. In some cases, electrical activity in one or more locations in the brain may be measured during a period extending from just prior to attempted speech or movement to a period just after attempted speech or movement. Embodiments of the methods further include implanting an interface capable of communication with a computing device in the head of the subject or placing the interface on the head of the subject to provide an externally accessible platform through which brain electrical signals can be acquired from the neural recording device and transmitted to a data processor for decoding. The processor may be provided by a computer or a handheld computing device (e.g., cell phone or tablet) programmed to decode the attempted speech and / or attempted movement from the recorded brain electrical signal data. The processor is programmed to use a machine learning model to decode one or more speech sounds or one or more electronic output device or system actions (e.g., one or more actions of an animated humanoid avatar) from the recorded brain electrical signal data as is discussed in greater detail below. Decoding Actions from Brain Electrical Signals Embodiments of the methods include decoding one or more speech sounds, text (i.e., one or more words or sentences), or one or more electronic output device or system actions (e.g., one or more actions of an animated humanoid avatar or prosthetic arm) from brain electrical signal data recorded, e.g., as described above using a machine learning system. The decoding machine learning system, in accordance with embodiments of the methods, may vary and may include, but is not limited to, any of the models discussed below. Decoding of speech sounds, text, and / or actions from brain activity may be aided by the use of self-supervised machine learning techniques, which discretize each speech sound or action into one or more discrete or continuous Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 speech units or action representations that serve as an effective intermediary for decoding neural activity patterns and features into meaningful action outputs. In some embodiments, the methods further include training the decoding machine learning system to perform one or more tasks facilitating the decoding of one or more speech sounds or one or more electronic output device or system actions from the recorded brain electrical signal data. In some embodiments, the training may further include validating and testing. The term “decoding machine learning system” or “decoding system” as used herein refers to the machine learning model, or group of connected or interconnected models, used to generate one or more speech sounds, text, or one or more electronic output device or system actions from brain electrical signal data recorded, e.g., as described above (i.e., the model(s) used for decoding brain electrical signal data). The decoding machine learning system or decoding system may include other machine learning architectures (e.g., encoders) and may perform multiple different techniques (e.g., encoding) during the process of generating speech sounds / text / output actions from brain electrical signal data. Further, models not explicitly referred to as the “decoding machine learning system” or “decoding system” may be a component of the decoding machine learning system or decoding system based on context. For example, a model referred to as a “neural encoder” is a component of the decoding system if it actively takes part in the aforementioned decoding of speech / text / actions from brain electrical signal data (i.e., is part of the process of decoding after the model(s) of the decoding system have completed training). Further, the decoding machine learning system or decoding system may not be the only machine learning model(s) that performs the technique of decoding. As such, the terms “decoding machine learning system” or “decoding system” are used for the sake of clarity and are not meant to inherently be limiting beyond the definition supplied above. The term “encoding machine learning system” or “encoding system” as used herein refers to the machine learning model, or group of connected or interconnected models, that, e.g., discretizes reference speech sounds or reference electronic output device or system actions (e.g., reference avatar animations), or creates embedding spaces using said reference sounds or actions, in order to generate intermediate representations for decoding neural activity patterns and features into meaningful action outputs. In some embodiments, the encoding machine learning system or encoding system may include other machine learning architectures (e.g., decoders) and may perform multiple different techniques (e.g., decoding) during the process of generating Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 effective intermediate representations useful for decoding recorded neural activity. In some cases, the decoding machine learning system may include one or more components of the encoding machine learning system. For example, an encoding system may include both an encoder and a decoder and the decoding system may include the encoder and / or the decoder of the encoding system. Further, the encoding machine learning system or encoding system may not be the only machine learning model(s) that performs the technique of encoding. As such, the terms “encoding machine learning system” or “encoding system” are used for the sake of clarity and are not meant to inherently be limiting beyond the definition supplied above. The term “synthesizing machine learning system” or “synthesizing system” as used herein refers to the machine learning model, or group of connected or interconnected models, that synthesizes one or more speech sounds or speech audio from intermediary representations generated using the decoding system. In this way, the synthesizing system may be a component of the decoding system (i.e., the decoding system may include the synthesizing system) in embodiments wherein the decoding system is used to generate speech sounds from brain electrical signal data. The synthesizing machine learning system or synthesizing system may not be the only machine learning model(s) that performs the technique of synthesizing. As such, the terms “synthesizing machine learning system” or “synthesizing system” are used for the sake of clarity and are not meant to inherently be limiting beyond the definition supplied above. The decoding machine learning system, in accordance with embodiments of the methods, may vary and may include, but is not limited to, any of the models discussed below or any standard machine learning model known in the art, as well as combinations thereof, capable of performing the decoding tasks described below. In some embodiments, the decoding machine learning system may include a Logistic Regression, K-Nearest Neighbors, Decision Tree, Random Forest and / or XGBoost model. In some embodiments, the decoding machine learning system may include an artificial neural network (NN) (e.g., a convolutional NN (CNN)). In some embodiments, the machine learning system includes a deep learning model. In these cases, the model may be three or more layers deep, such as five or more layers deep, or ten or more, or twelve or more. In some embodiments, the decoding machine learning system is configured to process sequential input data. In these instances, the decoding machine learning model may include, or be based on, a recurrent neural network (RNN) model or a transformer model (e.g., encoders, decoders and / or attention). In Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 embodiments where the decoding system includes an RNN, the RNN may include, e.g., long short-term memory (LSTM) architecture, gated recurrent units (GRUs), and / or an attention mechanism. In some embodiments, the decoding machine learning system may include an RNN with one or more layers of GRUs, such as 2 layers or more of layers of GRUs, or 3 or more of layers of GRUs, or 5 or more of layers of GRUs. The decoding machine learning system may learn from the contextual information of the brain electrical signal data and, e.g., may learn from the past to present context of the data and / or the present to past context of the data. In some embodiments, the decoding system may learn from both the past to present context and the present to past context of the brain electrical signal data (i.e., the decoding machine learning model may be bidirectional). For example, the decoding system may include, or be based on, e.g., a bidirectional LSTM model, an RNN model with an attention, a convolutional recurrent neural network model with an attention (CRNN-A), a transformer model, a bidirectional RNN model, or a Transducer (e.g., an RNN Transducer) model. In some embodiments, the bidirectional decoding system includes, or is based on, an RNN model, and the bidirectional RNN model may include one or more layers of GRUs as described above. In some embodiments, the decoding machine learning system such as, e.g., a bidirectional decoding system including a bidirectional RNN model having one or more layers of GRUs, may include one or more convolutional layers. In some cases, the decoding machine learning system includes a linear readout layer. In some cases, a layer of the decoding system may comprise a plurality of hidden units such as, e.g., 10 or more units, or 50 or more units, or 100 or more, or 500 or more. In certain embodiments, brain electrical signal data is decoded into one or more speech sounds or electronic output device or system actions (e.g., one or more actions of an animated humanoid avatar) using intermediate representations. In some cases, multiple intermediate representations may be combined to generate or decode a single speech sound or electronic output device or system action. In other instances, a single intermediate representations may be used to generate or decode a single speech sound or electronic output device or system action. In some embodiments, the intermediate representations are a set of discrete action representations. In other words, the decoding system may decode one or discrete action representations from the brain electrical signal data, and the one or discrete action representations may then be used to decode the one or more speech sounds or electronic output device or system actions. The set of Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 discrete action representations may be obtained using self-supervised, supervised, or unsupervised machine learning techniques. In some cases, the set of discrete action representations may be obtained using self-supervised machine learning techniques. In some embodiments, the discrete action representations are generated during self-supervised training of an encoding machine learning system (i.e., one or more machine learning models) using reference speech sounds (e.g., the speech sounds of the word or phrase sets as discussed above) and / or speech sounds from a large database or reference electronic output device or system actions (e.g., reference avatar animations). In some instances, the encoding machine learning system may include an artificial neural network (ANN). In some embodiments, the encoding system includes a convolutional element such as, e.g., one or more convolutional layers. In embodiments where the attempted action is attempted speech, the intermediate representations may be discrete speech units. In these instances, the discrete speech units may be generated by applying a pretrained (i.e., using self-supervised machine learning techniques) encoding machine learning system to reference speech sounds (i.e., the speech sounds of the word or phrase sets as discussed above). In these cases, the encoding system may include a transformer encoder. For example, the encoding machine learning system may include a Hidden- Unit Bidirectional Encoder Representations from Transformers (HuBERT) model. In some embodiments, the HuBERT model generates 50 or more different discrete speech units, such as 80 or more different discrete speech units, or 100 or more different discrete speech units. In embodiments where the attempted action is performed in order to control the actions or animations of a computer-generated humanoid avatar, the intermediate representations may be discrete action representations. In these instances, the discrete action representations may be generated during self-supervised training of an encoding machine learning system using reference avatar animations. In these instances, the encoding machine learning system may include a neural network. In some embodiments, the neural network includes one or more components of a variational autoencoder and / or may include one or more rectified linear unit (ReLU) activations. In some embodiments, the encoding machine learning system may include an encoder and a decoder, e.g., for encoding each of the reference avatar animations into one or more discrete action representations (encoder) and for decoding one or more discrete action representations into an avatar animation (decoder). For example, the encoding machine learning system may include a vector-quantized variational autoencoder (VQ-VAE) model. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 As described above, brain electrical signal data may be decoded into one or more speech sounds or electronic output device or system actions (e.g., one or more actions of an animated humanoid avatar) using intermediate representations. In some embodiments, the intermediate representations are a set of continuous action representations. In other words, the decoding system may decode one or continuous action representations from the brain electrical signal data, and the one or continuous action representations may then be used to decode the one or more speech sounds or electronic output device or system actions. In some embodiments, the continuous action representations are generated using embedding. In these embodiments, the embeddings / embedding space may be generated using any context. For example, the embedding space may be generated using semantic context (i.e., wherein words of similar meanings tend to be near each other in the embedding space), movement context (i.e., wherein similar types of movements, or movements using similar muscles, tend to be near each other in the embedding space), sound context (i.e., wherein similar types of sounds tend to be near each other in the embedding space), and / or any other context. The embedding space for continuous action representations may be obtained using self-supervised, supervised, or unsupervised machine learning techniques. In some cases, the embedding space for continuous action representations may be obtained using supervised machine learning techniques. In some embodiments, the embedding space for continuous action representations is generated during supervised training of an encoding machine learning system using reference speech sounds (e.g., the speech sounds of the word or phrase sets as discussed above) and / or speech sounds from a large database or reference electronic output device or system actions (e.g., reference avatar animations). In some instances, the encoding machine learning system may include an artificial neural network (ANN). In some embodiments, the encoding system includes a convolutional element such as, e.g., one or more convolutional layers. In some embodiments, the decoding machine learning system may include a predictor model configured to predict a future intermediate representation or decoding system output (e.g., speech sound, text, and / or action) based, e.g., on previous representations / outputs output by the decoding system. In some cases, the predictor model may predict the most likely next intermediate representation or produce a probability distribution (i.e., discrete or continuous) over a number or range of possible future intermediate representations using, e.g., previously output representations. For example, in embodiments of the decoding system configured to Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 decode speech sounds or text via discrete representations, the predictor model may include a language model (e.g., BERT or ChatGTP) configured to predict the next word, or intermediate representation thereof, the subject will attempt to speak based, e.g., on what words would form a coherent sentence. In some embodiments, the decoding machine learning system includes a neural encoder (e.g., including one or more components of an encoding system as described above) configured to output one or more intermediate representations such as, e.g., a probability distribution of the most likely intermediate representations, based on recorded brain electrical signal data. In some embodiments, the decoding machine learning system includes a predictor model, a neural encoder, and a joiner module configured to combine the outputs of the predictor model and the neural encoder. In this way, the decoder system may be able to generate intermediate representations based both on contextual information (e.g., semantic context) and brain electrical signals. In some instances, the decoder system may include an RNN Transducer. In some embodiments, the decoding machine learning system may be configured to decode speech sounds and / or text from multiple different languages simultaneously. In other words, the user does not need to pre-specify which language intend to communicate in, and the system instead infers the desired language from their neural activity. In some embodiments, the decoding system may decode words from a multilingual vocabulary set as the user attempts to speak the words. In some embodiments, the decoding system may perform multilingual decoding of electrical brain signal data using multilingual linguistic representations that are shared across languages. These multilingual linguistic representations may include phonemes, syllables, word pieces, or other linguistic units. The decoded representations of linguistic units can vary across a variety of forms, including, but not limited to, a time series of predicted linguistic-unit probabilities. In some embodiments, one or more language models (e.g., natural language models) are used to score decoded candidate linguistic speech units, words, or sentences (e.g., using linguistic context specific to each language) in order to facilitate selection and output of the most likely decoded speech unit, word, or sentence. In some embodiments, the decoding system may include multiple different machine learning models each configured to generate context or information via electrical brain signals associated with a different type or category of attempted action. These different types or categories of attempted actions may include, e.g., the attempted actions of different attempted action sets as described above. In some cases, each of two or more models of the decoding Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 system are configured to generate information from the attempted action of a different part or region of the body. For example, the decoding system may include, e.g., a neural encoder trained to produce representations from neural activity features associated with attempted foot movements and a neural encoder trained to produce representations from neural activity features associated with attempted arm movements. In some embodiments, each of two or more models of the decoding system are configured to generate information from different categories of attempted action performed by the same body part. For example, the decoding system may include, e.g., a neural encoder trained to produce representations from neural activity features associated with attempted ASL hand movements and a neural encoder trained to produce representations from neural activity features associated with attempted cursor control movements. In some embodiments, the decoding system may include a joiner module configured to combine (and, e.g., weight) the outputs of the different machine learning models in order to decode speech sounds, text, and / or electronic output device actions from recorded brain electrical signal data. In some embodiments, the decoding system may be configured decode multiple outputs simultaneously, e.g., from context or information generated by one or multiple machine learning models (e.g., one or multiple neural encoders). For example, a user may attempt to speak while imagining moving their hand to simultaneously facilitate text decoding (e.g., entering a text message into a computer application) and cursor control (e.g., moving a computer cursor around on a screen), wherein the decoding system includes a first neural encoder for the attempted speech and a second neural encoder for the imagined hand movements. In another embodiment, a user may attempt to speak in order to simultaneously control an avatar and a keyboard, e.g., via a single neural encoder. In this way, the decoding system may be used to control or generate any number of different outputs or effectors using any number of different inputs or user intents (e.g., attempted action sets). In some cases, the decoding system may generate / control 1 output based on 2 or more inputs, or 2 or more outputs based on 1 input, or 3 or more outputs based on 1 input, or 1 output based on 3 or more inputs, or 5 or more outputs based on 1 input, or 1 output based on 5 or more inputs, or 2 or more outputs based on 2 or more inputs, or 3 or more outputs based on 3 or more inputs, or 5 or more outputs based on 2 or more inputs, or 2 or more outputs based on 5 or more inputs, etc. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 In some embodiments, the decoding system may include an intent detection system comprising one or more machine learning models trained to differentiate neural activity (e.g., brain electrical signal data) associated with volitional attempts or intents of a specific activity (e.g., the attempted movements of a specific movement set or attempted speech of a word or phrase set, as described above) from neural activity related to irrelevant activities. In some embodiments, the intent detection system may be configured to increase the accuracy of the decoding system (e.g., in real-world environments where environmental noise or other irrelevant activities or distractors may be present) and / or volitionally engage and / or disengage the decoding system via the neural activity (e.g., brain electrical signal data) of a user. For example, the decoding system may be inactive until it is activated by the intent detection system when neural features (e.g., HGA and / or LFS) associated with attempted speech are detected. In some embodiments, the intent detection system may be trained via neural activity collected as a user engages in activities unrelated to intended actions. For example, an intent detection system used for detecting neural activity associated with attempted speech may be trained using data collected as a user reads text on a screen or passively listens to speech played via loudspeaker, e.g., in order to better differentiate attempted speech from other language associated tasks. In some cases, a model of the intent detection system may be trained to distinguish time windows of neural activity associated with a specific activity from time windows of neural activity associated with a different activity. As discussed above, in some embodiments an intent detection system may be used to volitionally engage and disengage the decoding system via brain electrical signal data. In this way, the decoding system may be utilized by a subject as it is needed simply by attempting to speak (e.g., without any setup steps), allowing for the subject to more easily initiate spoken conversation or respond in real time to spontaneous interactions. In some embodiments, the decoding system may decode brain electrical signal data into one or more speech sounds or electronic output device or system actions in real-time, with low latency. In some embodiments, low latency is achieved by continuously streaming small increments of data through a transducer model, as described above. In these cases, the increments may be 500 ms or less, or 100 ms or less, or 80 ms or less, or 50 ms or less, or 10 ms or less. As discussed above, the methods further include training the decoding machine learning system to perform one or more tasks facilitating the decoding of one or more speech sounds or Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 one or more electronic output device or system actions from the recorded brain electrical signal data. In certain embodiments, the training includes: obtaining reference output device or system actions for a plurality of actions; encoding each reference action into a temporal sequence of discrete action representations using the encoding machine learning system (or, e.g., generating embeddings using the encoding machine learning system if continuous action representations are utilized); recording the brain electrical signal data associated with attempted action by the subject for each one of the plurality of actions; training the decoding machine learning system to perform one or more tasks facilitating the decoding of one or more electronic output device or system actions from the recorded brain electrical signal data using the action representations derived from the reference output device or system actions and the brain electrical signal data associated with the attempted action for each one of the plurality of actions. In some embodiments, the plurality of actions consist of the actions of an action set as described above. In some embodiments, the decoding machine learning system is trained to predict the most likely discrete action representation associated with a segment of electrical signal data using the action representations derived from the reference output device or system actions and the brain electrical signal data associated with the attempted action for each one of the plurality of actions. In some embodiments, the decoding machine learning system is trained to predict the probability of each discrete action representation for a segment of electrical signal data. In some instances, the decoding machine learning system is trained to learn mappings between neural activity patterns of electrical signals in the brain electrical signal data and the discrete action representations. In some embodiments, device or system actions include avatar animations. In some embodiments, the avatar animations are of the avatar performing the same action as the action attempted by the subject. In certain embodiments, the reference avatar animations are obtained from an avatar-animation system. In some embodiments, the decoding machine learning system is trained to predict the most likely continuous action representation associated with a segment of electrical signal data using the action representations (e.g., embeddings) derived from the reference output device or system actions and the brain electrical signal data associated with the attempted action for each one of the plurality of actions. In some embodiments, the decoding machine learning system is trained to predict the probability of different regions of an embedding space for a segment of electrical signal data. In some instances, the decoding machine learning system is trained to learn Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 mappings between neural activity patterns of electrical signals in the brain electrical signal data and the continuous action representations (e.g., embeddings). In embodiments where the attempted action is attempted speech, the training may include: obtaining reference electronic speech waveforms for a plurality of phrases; encoding each reference speech waveform into a temporal sequence of discrete speech units (or, e.g., generating embeddings using the encoding machine learning system if continuous action representations are utilized) using the encoding system; recording the brain electrical signal data associated with attempted speech by the subject for each one of the plurality of phrases; training the decoding machine learning system to perform one or more tasks facilitating the decoding of one or more speech sounds (or, e.g., text) from the recorded brain electrical signal data using the discrete speech units (or, e.g., embeddings) derived from the reference speech waveforms and the brain electrical signal data associated with the attempted speech for each one of the plurality of phrases. In some embodiments, the plurality of phrases consist of the words or phrases of a word or phrase set as described above. In certain embodiments, the reference electronic speech waveforms are obtained from a recruited speaker. In some embodiments, the reference electronic speech waveforms are obtained using a text-to-speech algorithm. In some embodiments, the decoding machine learning system is trained to predict the most likely discrete speech unit associated with a segment of electrical signal data using the discrete speech units derived from the reference speech waveforms and the brain electrical signal data associated with the attempted speech for each one of the plurality of phrases. In some embodiments, the decoding machine learning system is trained to predict the probability of each discrete speech unit for a segment of electrical signal data. In some instances, the decoding machine learning system is trained to learn mappings between neural activity patterns of electrical signals in the brain electrical signal data and the discrete speech units. In some embodiments, the decoding machine learning system is trained to predict the most likely continuous speech representation associated with a segment of electrical signal data using the speech representations (e.g., embeddings) derived from the reference speech waveforms and the brain electrical signal data associated with the attempted speech for each one of the plurality of phrases. In some embodiments, the decoding machine learning system is trained to predict the probability of different regions of a word embedding space for a segment of electrical signal data. In some instances, the decoding machine learning system is trained to learn Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 mappings between neural activity patterns of electrical signals in the brain electrical signal data and the continuous speech representations (e.g., word embeddings). In some embodiments, the training algorithms and hyperparameters used to control the training of the decoding system may depend on, e.g., the nature or architecture of the decoding machine learning system model(s), the tasks the machine learning model(s) is trained to perform, the desired accuracy or efficiency of the decoding machine learning system, and / or the nature or size of the training data set (e.g., the reference speech waveforms or reference avatar animations). In some embodiments, algorithms or techniques are used (e.g., during training) to account for the differences in timing between reference speech waveforms or reference electronic output device or system actions (e.g., reference avatar animations) and the brain electrical signal data associated with the attempted action (i.e., attempted speech or an attempted movement). In some embodiments, a CTC loss function is used during training. In some cases, algorithms or techniques are used during training that correct for overfitting issues such as, e.g., dropout layers. In embodiments wherein a decoding system is trained to decode speech sounds or text from electrical brain signal data, transfer learning may be used to train the decoding system for a language using a different decoding system trained using a different language. For example, a neural-decoding model may first be trained using brain electrical signal data and associated metadata from one language before the weights resulting from the aforementioned training are used to initialize the weights of a subsequent model trained using brain electrical signal data and activity from a second language. In some cases, transfer learning may be used to train a decoding system for a second language after the decoding system has first been trained using a first language. In some cases, multilingual transfer learning, as described above, may be used to reduce the amount of training data required to train a decoding system. As discussed above, algorithms or techniques may be used to account for the differences in timing between reference speech waveforms or reference electronic output device or system actions (e.g., reference avatar animations) and the brain electrical signal data associated with the attempted action (i.e., attempted speech or an attempted movement). In some cases, a silence / pause or blank token is used in order to account for the differences in timing between reference speech waveforms or reference electronic output device or system actions. For example, the intermediate representations (e.g., discrete speech units, continuous action Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 representations, etc.) may include a representation of the lack of speech or the lack of an action (e.g., representing periods of time where the subject is not attempting speech or an action) that can be referred to as a blank token. In embodiments wherein the intermediate representations are continuous, the blank token may be zero or may be the origin in the coordinate plane of the embedding space. In some cases, connectionist temporal classification (CTC) decoding is used with the blank token. In some cases, an RNN-transducer is used with the blank token. In some embodiments, training methods that do not rely on alignment are utilized (e.g., supervised machine learning techniques). In embodiments wherein supervised machine learning techniques are utilized, the labels may be generated in any fashion, e.g., from the task, using forced alignment, or using unsupervised learning. In some cases, breaks between actions (e.g., words or sentences for attempted speech) may be decoded using a blank token. In these embodiments, the decoding system may be continuously pinged or polled in order to determine when the blank or silence token is being decoded from the brain electrical data by the decoding system. In some cases, the decoding system may be polled every 500 ms or less, such as every 100 ms or less, or every 10 ms or less, or every 5 ms or less, or every 1 ms or less. In these cases, if the blank or silence token is determined to have been decoded by the decoding system for a certain number of polls in a row, the end of an action, or group of actions, may be determined (e.g., the end of a word, sentence, or attempted speech). In some cases, a certain number of blank or silence tokens may initiate the collapse of a CTC search beam. In some cases, a certain number of blank or silence tokens may initiate the next frame of an RNN-transducer search function. In some embodiments, the training may further include testing the trained decoding machine learning system or decoding machine learning systems. By testing in this context is meant evaluating the trained decoding machine learning system using, e.g., reference data (e.g., reference speech waveforms or reference avatar animations) different from the reference data used for training after the machine learning system has finished training. The testing may use one or more metrics to evaluate the performance of the trained decoding machine learning system . In some embodiments, the metric may be used to determine if the trained decoding machine learning system performs sufficiently using, e.g., a predetermined threshold (i.e., requirement). In these instances, if the trained decoding machine learning system does not meet the predetermined threshold, the model may be discarded and / or another model may be trained. In embodiments where another decoding machine learning system is trained, one or more of the Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 model architecture, training and / or the training data set may be modified prior to training. In some instances, decoding machine learning models are trained until a trained machine learning models meets the predetermined threshold. The division between the training set of reference data and the testing set of reference data may vary. In some cases, roughly 80% of the reference data may be used for training and roughly 20% for testing. In some instances, roughly 70% of the reference data may be used for training and roughly 30% for testing. In some embodiments, the training may further include validating the decoding machine learning model or decoding machine learning models. By validating in this context is meant evaluating the decoding machine learning system during training using reference data different from the reference data used for training and testing. The validating may use one or more metrics to evaluate the performance of the decoding machine learning model or decoding machine learning models. In some embodiments, the validating may be used to, e.g., select model parameters (e.g., select one or more machine learning algorithms to continue training), optimize or tune hyperparameters (e.g., model hyperparameters or algorithm hyperparameters), etc. The division between the training reference data, validating reference data, and testing reference data may vary. In some cases, roughly 80% of the reference data may be used for training, roughly 10% for testing, and roughly 10% for validating. As discussed above, embodiments of the methods include decoding one or more speech sounds or one or more electronic output device or system actions (e.g., one or more actions of a computer-generated humanoid avatar) from recorded brain electrical signal data using a machine learning model. In some embodiments, the decoding machine learning model may be bidirectional and may include an RNN with one or more layers of GRUs. In certain embodiments, brain electrical signal data is decoded into one or more speech sounds or electronic output device or system actions using a set of discrete action representations (e.g., discrete speech units). The set of discrete action representations may be obtained using self- supervised machine learning techniques. In some embodiments, the discrete action representations are generated during self-supervised training of an encoding machine learning model. In embodiments where the attempted action is attempted speech, the encoding machine learning model may be a Hidden-Unit Bidirectional Encoder Representations from Transformers (HuBERT) model. In embodiments where the attempted action is performed in order to control the actions or animations of a computer-generated humanoid avatar, the encoding machine Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 learning model may be a vector-quantized variational autoencoder (VQ-VAE) model. The methods may further include training the decoding machine learning model to perform one or more tasks facilitating the decoding of one or more speech sounds or one or more electronic output device or system actions from the recorded brain electrical signal data. In some embodiments, the decoding machine learning model is trained to predict the most likely discrete action representation associated with a segment of electrical signal data using the action representations derived from reference data (e.g., reference waveforms or reference output device or system actions) and the brain electrical signal data associated with the attempted action for each one of a plurality of actions. In some embodiments, a CTC loss function is used during training. In some cases, the training may further include validating and testing. An encoding machine learning model or a synthesizing machine learning model may be used to synthesize speech audio (e.g., one or more speech sounds or an electronic speech waveform) or generate electrical control signals for an electronic output device or system from the discrete action representations decoded from the recorded the brain electrical signal data as is discussed in greater detail below. Synthesizing Audio / Visual Stimuli and Controlling Electronic Output Devices Embodiments of the methods may include decoding one or more speech sounds or one or more electronic output device or system actions (e.g., one or more actions of a computer- generated humanoid avatar) from discrete action representations (e.g., discrete speech units) decoded from the recorded the brain electrical signal data using the decoding machine learning model as discussed above. In some embodiments, one or more speech sounds are decoded from discrete speech units using a speech synthesizer. In some embodiments, one or more avatar animations are decoded from discrete action representations using an encoding machine learning model. Embodiments of the methods may further include synthesizing an electronic speech waveform from the decoded speech sounds and / or generating an electrical control signal from the decoded electronic output device or system actions. In some embodiments, the electronic speech waveform may be transformed into the subject’s own voice. In some embodiments, electrical control signals may be generated from multiple machine learning models that cause a computer-generated avatar to speak, perform orofacial movements for speech, and perform non- speech associated actions. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 As discussed above, one or discrete action representations may be decoded from the brain electrical signal data using the decoding model, and the one or more discrete action representations may then be used to decode the one or more speech sounds or electronic output device or system actions. In embodiments where discrete speech units are decoded from brain electrical signal data associated with attempted speech, the discrete speech units may be decoded into one or more speech sounds and synthesized into an electronic speech waveform using a speech synthesizer. In some embodiments, the speech synthesizer includes a machine learning model. In some embodiments, the synthesizing machine learning model is configured to generate a spectrogram from one or more discrete speech units. In other words, the synthesizing machine learning model may be trained to generate a spectrogram from one or more discrete speech units. In some instances, the spectrogram is a mel spectrogram. In some cases, the synthesizing machine learning model may include a neural network. In these instances, the synthesizing machine learning model may be an RNN such as, e.g., a bidirectional RNN. In embodiments where the synthesizing machine learning model includes a bidirectional RNN, the bidirectional RNN may include GRUs or LSTM. In some cases, the bidirectional RNN includes one or more layers of GRUs. In some embodiments, the synthesizing machine learning model may include one or more convolutional layers and / or may include an attention mechanism. In some cases, the synthesizing machine learning model may include the Tacotron2 model. In some embodiments, the speech synthesizer further includes a vocoder configured to synthesize an electronic speech waveform from the spectrogram generated using the synthesizing machine learning model. In some instances, the vocoder may include a machine learning model. In some cases, the vocoder machine learning model may include a neural network. In these instances, the neural network may include one or more convolutional layers. In some cases, the vocoder machine learning model may include the WaveGlow model. In certain embodiments, the electronic speech waveform decoded from recorded brain electrical signal data and synthesized using the speech synthesizer is transformed or converted into a personalized electronic speech waveform. For example, the electronic speech waveform may be transformed such that it resembles the subject’s own voice. In some embodiments, the electronic speech waveform is transformed using a machine learning model, such as a machine learning model including a neural network. In some cases, the conversion machine learning Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 model may include an encoder and / or may be a zero-shot learning model. In these instances, the conversion machine learning model may include the YourTTS model. In some embodiments, the electronic speech waveform decoded from recorded brain electrical signal data and synthesized using the speech synthesizer is converted into an audible speech waveform using a loudspeaker. In some embodiments, an electronic output device or system is controlled using the electronic speech waveform. For example, an animated avatar (such as the humanoid avatar described herein) may be controlled to perform speech associated orofacial movements based on the electronic speech waveform. In some embodiments, the electronic speech waveform is played in sync with a humanoid avatar performing orifical facial movements decoded using a different decoding model than the decoding model used to decode the one or more speech sounds used to synthesize the electronic speech waveform. In embodiments where discrete action representations are decoded from brain electrical signal data associated with an attempted action (e.g., an attempted movement), the discrete action representations may be decoded into one or more electronic output device or system actions (e.g., one or more actions of a computer-generated humanoid avatar) using a decoder. In some cases, the decoder may be the decoder of the encoding machine learning model as discussed above. For example, a VQ-VAE may be used to discretize or encode each reference electronic output device or system action into one or more discrete action representations and the VQ-VAE’s decoder may be used to decode electronic output device or system actions from discrete action representations decoded from recorded brain electrical signal data. The decoded electronic output device or system actions may then be used to control the electronic output device or system (i.e., to perform the decoded action). For example, the decoded electronic output device or system actions may be used to generate electrical control signals that are transmitted to the electronic output device or system. In some embodiments, the attempted actions performed by the subject and the actions of the device or system include both speech associated actions and non-speech communicative gestures. In some embodiments, the decoding machine learning model is trained to discriminate between actions performed by different regions of the body and / or between speech associated actions and non-speech communicative gestures. In certain embodiments, the electronic output device or system is controlled using actions decoded from multiple different decoding machine learning models. In some embodiments, the Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 actions decoded from the different decoding machine learning models occur concurrently (i.e., the attempted actions associated with the decoded actions occur concurrently) and the electronic output device or system is controlled to perform the actions simultaneously. For example, multimodal speech decoding may be used to display animations of a humanoid avatar performing orofacial speech movements using a first decoding machine learning model (e.g., as described above), audibly play a speech waveform synthesized using a second decoding machine learning model (e.g., as described above), and display text of the speech using a third decoding machine learning model (see, e.g., the Examples section below) simultaneously as a subject attempts to speak. In some embodiments, the multiple decoding machine learning models each are trained using a different action set as described above and, e.g., are adapted for a specific region of the body (e.g., the subject’s body and the avatar’s body) or for a specific purpose. In some cases, the subject may be able to switch between action sets associated with specific regions of the body and with communication in a specific context as desired. For example, the subject may switch to a word or phrase set adapted for playing basketball for generating audible speech and orofacial movements, and to an action set adapted for playing basketball for lower limb movements when playing a basketball video game by controlling the humanoid avatar. In some embodiments, one decoding machine learning model may affect the manner in which the avatar performs an action decoded from a different machine learning model. In some embodiments, separate decoding machine learning models are trained for speech and orofacial movements for speech. In these embodiments, a decoded non-speech communicative gesture affects how the avatar performs a decoded speech associated action. For example, a decoded emotional expression may affect the inflection of decoded speech performed by an avatar. As discussed above, embodiments of the methods may include decoding one or more speech sounds or one or more electronic output device or system actions from discrete action representations decoded from the recorded the brain electrical signal data using the decoding machine learning model. In embodiments where discrete speech units are decoded from brain electrical signal data associated with attempted speech, the discrete speech units may be decoded into one or more speech sounds and synthesized into an electronic speech waveform using a speech synthesizer. In some embodiments, the speech synthesizer includes a machine learning model trained to generate a mel spectrogram from one or more discrete speech units. In some Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 cases, the synthesizing machine learning model may include the Tacotron2 model. In some embodiments, the speech synthesizer further includes a vocoder configured to synthesize an electronic speech waveform from the spectrogram generated using the synthesizing machine learning model. The electronic speech waveform decoded from recorded brain electrical signal data and synthesized using the speech synthesizer may be transformed or converted into a personalized electronic speech waveform using a zero-shot learning model such as, e.g., the YourTTS model. In embodiments where discrete action representations are decoded from brain electrical signal data associated with an attempted action (e.g., an attempted movement), the discrete action representations may be decoded into one or more actions of a computer-generated humanoid avatar using the decoder of the encoding machine learning model discussed above. The decoded avatar actions may then be used to control the avatar by generating electrical control signals that are transmitted to the device or system generating the avatar. In certain embodiments, the avatar is simultaneously controlled using actions decoded from multiple different decoding machine learning models. In some embodiments, one decoding machine learning model may affect the manner in which the avatar performs an action decoded from a different machine learning model. The speech sounds generated from recorded brain electrical signal data and the avatar controlled using recorded brain electrical signal data have a variety of applications, as discussed in greater detail below. Avatar Applications and Environments Embodiments of the methods may include methods of applying the speech sounds generated from recorded brain electrical signal data and the avatar controlled using recorded brain electrical signal data. In some embodiments, the speech sounds generated from recorded brain electrical signal data are used as the voice of the avatar controlled using recorded brain electrical signal data. In some embodiments, a virtual environment is created for the avatar, or the avatar is configured to interact in a virtual environment. In some cases, the avatar may be used for communication, entertainment, work, or therapy. The avatar may be highly personalized according to the needs and / or desires of the subject. In some embodiments, the computer-generated avatar is used by the subject to communicate with one or more individuals in person (e.g., using a TV screen, projector, computer monitor, or augmented reality goggles, glasses, or contacts). In some cases, the avatar Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 is used by the subject to communicate with one or more individuals in an interactive virtual environment (e.g., using one or more TV screens, projectors, computer monitors, or virtual reality goggles, glasses, or contacts). For example, the avatar may be used by the subject to communicate with other individuals in a metaverse, or in an online video game. In these instances, the other individuals may play the video game or interact in the metaverse using a personal computer or a gaming console (e.g., using an Xbox or PlayStation® controller). In some embodiments, the actions performed by the avatar are determined using recorded brain electrical signal data (e.g., as described above) and the virtual environment the avatar is rendered or generated in. For example, an attempted hand turn by the subject may result in the avatar opening a door when the avatar is within a certain proximity to the door and may result in the avatar waving when the avatar is not adjacent to a door. In some embodiments, the actions performed by the avatar are determined using recorded brain electrical signal data (e.g., as described above) and the previous actions performed by the avatar. For example, multiple consecutive attempted finger snaps may result in the avatar performing a dance. In some embodiments, the computer-generated avatar is used by the subject for therapy such as, e.g., physical therapy. In these instances, the physical therapy involves regaining mobility of a body part after an injury. For example, by watching the avatar perform leg movements when the subject attempts leg movements the subject may take advantage of neuroplasticity in order to regain leg mobility. In some embodiments, the computer-generated avatar is used by the subject to control one or more electronic devices in the subject’s environment such as, e.g., one or more electronic devices in the same room as the subject. For example, the subject may turn on a light by controlling the avatar to turn on the light in augmented reality. The electronic devices controlled by the avatar may include, but are not limited to, one or smart home appliances, or one or more electronic motors configured to move objects in the subject’s environment. In some embodiments, the computer-generated avatar is personalized. In these instances, one or more physical characteristics of the avatar’s body (e.g., facial feature shapes, height, weight, skin color, hair color, eye color, etc.) and / or the avatar’s outfits (e.g., shirts, hats, pants, skirts, shoes, socks, necklaces, earrings, etc.) may be customizable by or for the subject. FIG. 32 provides a flow diagram depicting a method of controlling a fully embodied virtual avatar using recorded brain electrical signal data in accordance with an embodiment of Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 the invention. At step 3201, the subject attempts a movement. The movement may be a hand or arm movement such as a wave performed in order to acknowledge another individual in the same physical or virtual environment as the subject. At step 3202, the neural activity occurring as a result of the subject attempting the movement is recorded, e.g., using a non-penetrating high- density ECoG electrode array positioned on the sensorimotor cortex of the subject’s brain, and high-gamma activity and low-frequency signals are extracted to produce brain electrical signal data associated with the attempted action. At step 3203, avatar gestures are decoded from the brain electrical signal data using the decoding machine learning model and the decoder of the encoding machine learning model as described above. At step 3204, the decoded avatar gestures are used to animate an avatar in a virtual environment and at step 3205 the animated avatar and, e.g., the virtual environment are displayed on a visual display device. The subject may then attempt another movement using feedback obtained by watching the displayed avatar animation. FIG. 33 provides a flow diagram depicting a method of controlling and displaying a fully embodied virtual avatar using recorded brain electrical signal data in accordance with an embodiment of the invention. At step 3301, the neural activity occurring as a result of the subject attempting a movement is recorded, e.g., using an intracortical non-penetrating high-density ECoG electrode, and high-gamma activity and low-frequency signals are extracted to produce brain electrical signal data associated with the attempted action. At step 3302, the decoding machine learning model decodes one or more discrete action representations from the brain electrical signal data. At step 3303, a decoder such as, e.g., the decoder of the autoencoder used to generate the discrete action representations from reference avatar animations is used to decode the one or more discrete action representations (i.e., the representations decoded from the brain electrical signal data in step 3302) into one or more avatar animations. At step 3304, the decoded avatar movements are used to generate an avatar movement animation and at step 3205 the is played. At step 3304, the virtual environment, including the avatar animation, is reproduced on a display device. FIG. 34 provides a flow diagram depicting a method for training a machine learning model to predict avatar movements using recorded neural activity in accordance with an embodiment of the invention. At step 3401, a go cue is provided to the subject indicating when the subject should initiate an attempted action. The go cue may be provided visually on a display and may be preceded by a countdown to the presentation of the go cue. At step 3402, the neural Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 activity occurring as a result of the subject attempting a movement is recorded, e.g., using an intracortical non-penetrating high-density ECoG electrode, and high-gamma activity and low- frequency signals are extracted to produce brain electrical signal data associated with the attempted action. In some embodiments, the processor is programmed to use the recorded brain electrical signal data within a time window before the go cue and following the go cue. At step 3403, a decoding machine learning model and, e.g., an autoencoder are trained to predict avatar movements from the recorded brain electrical signal data (e.g., within a time window before the go cue and following the go cue). Training for the decoding model may include: obtaining reference avatar animations for a plurality of movements; encoding each reference animation into a temporal sequence of discrete action representations using the autoencoder; training the decoding machine learning model to predict the most likely discrete action representation associated with a segment of electrical signal data using the action representations derived from the reference avatar animations and the brain electrical signal data recorded at step 3402. At step 3404, an avatar movement decoder model is produced that can decode avatar movements from brain electrical signal data. FIG. 35 illustrates an overview of a non-speech communicative gesture neural-decoding pipeline in accordance with embodiments of the invention. FIG. 36 provides a block diagram of a pipeline for concurrent multi-effector neural- decoding in accordance with embodiments of the invention. FIG. 37 provides a block diagram of a pipeline for non-verbal linguistic neural-decoding in accordance with embodiments of the invention. SYSTEMS ANDCOMPUTERIMPLEMENTEDMETHODSAspects of the present disclosure further include systems, such as computer-controlled systems, for practicing embodiments of the above methods. Aspects of the systems may include: a neural recording device including an electrode adapted for positioning at a location in a sensorimotor cortex region of the brain of the subject to record brain electrical signal data associated with attempted speech by the subject; a processor programmed to use a machine learning model to decode one or more speech sounds from the recorded brain electrical signal data; an interface in communication with a computing device, said interface adapted for positioning at a location on the head of the subject, wherein the interface receives the brain Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 electrical signal data from the neural recording device and transmits the brain electrical signal data to the processor; and an audio speaker for playing the one or more speech sounds from the recorded brain electrical signal data. Aspects of the systems may also include: a neural recording device including an electrode adapted for positioning at a location in a sensorimotor cortex region of the brain of the subject to record brain electrical signal data associated with an attempted action by the subject; a processor programmed to use a machine learning model to decode an avatar animation from the recorded brain electrical signal data; an interface in communication with a computing device, said interface adapted for positioning at a location on the head of the subject, wherein the interface receives the brain electrical signal data from the neural recording device and transmits the brain electrical signal data to the processor; and a display component for displaying the avatar action from the recorded brain electrical signal data. For example, electrical activity in the high gamma frequency range (such as 70 Hz to 150 Hz) and / or low frequency range (e.g., 0.3 Hz to 100 Hz) from the precentral gyrus, postcentral gyrus, superior temporal gyrus, middle temporal gyrus, or any combination thereof may be recorded with the neural recording device using this system, wherein the interface receives the brain electrical signal data from the neural recording device and transmits the brain electrical signal data to a processor. The processor may run programming for speech sounds or avatar actions from the recorded brain electrical signal data using one or more machine learning models, as described herein. In some embodiments, a computer implemented method is used for decoding speech sounds from recorded brain electrical signal data associated with attempted speech by a subject. The processor may be programmed to perform steps of the computer implemented method including: receiving the recorded brain electrical signal data associated with the attempted speech by the subject; and decoding one or more speech sounds from the recorded brain electrical signal data using a machine learning model. In some embodiments, a computer implemented method is used for decoding avatar actions from recorded brain electrical signal data associated with an attempted action by a subject. The processor may be programmed to perform steps of the computer implemented method including: receiving the recorded brain electrical signal data associated with the Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 attempted action by the subject; and decoding one or more electronic output device or system actions from the recorded brain electrical signal data using a machine learning model. In certain embodiments, the computer implemented method further includes storing a user profile for the subject including information regarding the patterns of electrical signals in the recorded brain electrical signal data associated with attempted speech or an attempted action by the subject. The recorded brain electrical signal data may be processed in various ways before decoding. For example, data processing may include, without limitation, real-time sample-by- sample processing of neural feature streams, the use of common-average referencing across individual electrode channels, the use of finite impulse response (FIR) filters to perform digital signal filtering, a running sliding-window normalization procedure, e.g., using Welford’s method, automatic artifact rejection, and parallelization and linear pipelining to improve computational efficiency. Processing of neural features may be performed in real-time to extract one or more feature streams for use during speech / action decoding. For a description of data processing methods, see, e.g., Moses et al. (2018) J. Neural. Eng. 15(3):036005, Moses et al. (2019) Nat. Commun. 201910(1):3096, Moses et al. (2021) N. Engl. J. Med. 385(3):217-227, Sun et al. (2020) J. Neural. Eng. 17(6), and Makin et al. (2020) Nature Neuroscience 23:575- 582; herein incorporated by reference in their entireties. In some instances the systems further include one or more computers for complete automation or partial automation of the methods described herein. In some embodiments, systems include a computer having a computer readable storage medium with a computer program stored thereon. The methods described herein can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, a data processing apparatus. The computer readable medium can be a machine- readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or any combination thereof. A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network. In a further aspect, the system for performing the computer implemented method, as described, may include a computer containing a processor, a storage component (i.e., memory), a display component, and other components typically present in general purpose computers. The storage component stores information accessible by the processor, including instructions that may be executed by the processor and data that may be retrieved, manipulated or stored by the processor. The storage component includes instructions. For example, the storage component includes instructions for decoding a speech sound or action from recorded brain electrical signal data associated with attempted speech and / or attempted action by a subject. The computer processor is coupled to the storage component and configured to execute the instructions stored in the storage component in order to receive brain electrical signal data associated with attempted speech by the subject and analyze the data according to one or more algorithms, as described herein. The display component displays the sentence decoded from the recorded brain electrical signal data. The storage component may be of any type capable of storing information accessible by the processor, such as a hard-drive, memory card, ROM, RAM, DVD, CD-ROM, USB Flash drive, write-capable, and read-only memories. The processor may be any well-known processor, such as processors from Intel Corporation. Alternatively, the processor may be a dedicated controller such as an ASIC or an FPGA. The instructions may be any set of instructions to be executed directly (such as machine code) or indirectly (such as scripts) by the processor. In that regard, the terms "instructions," "steps" and "programs" may be used interchangeably herein. The instructions may be stored in Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 object code form for direct processing by the processor, or in any other computer language including scripts or collections of independent source code modules that are interpreted on demand or compiled in advance. Data may be retrieved, stored or modified by the processor in accordance with the instructions. For instance, although the system is not limited by any particular data structure, the data may be stored in computer registers, in a relational database as a table having a plurality of different fields and records, XML documents, or flat files. The data may also be formatted in any computer-readable format such as, but not limited to, binary values, ASCII or Unicode. Moreover, the data may include any information sufficient to identify the relevant information, such as numbers, descriptive text, proprietary codes, pointers, references to data stored in other memories (including other network locations) or information which is used by a function to calculate the relevant data. In certain embodiments, the processor and storage component may include multiple processors and storage components that may or may not be stored within the same physical housing. For example, some of the instructions and data may be stored on removable CD-ROM and others within a read-only computer chip. Some or all of the instructions and data may be stored in a location physically remote from, yet still accessible by, the processor. Similarly, the processor may include a collection of processors which may or may not operate in parallel. The system also includes an interface capable of communication with a computing device. The interface may be implanted in the cranium or placed on the head of the subject to provide an externally accessible platform through which brain electrical signals can be acquired from the neural recording device and transmitted to a computing device for decoding. In some embodiments, the interface includes a percutaneous pedestal connector anchored in the cranium of the subject. The interface can be connected, for example, to a computing device such as a computer or a handheld computing device (e.g., cell phone or tablet) with a detachable digital connector and cable. Alternatively, the interface may be connected to a computing device wirelessly. In some embodiments, the interface includes a first wireless communication unit in communication with a computing device including a second wireless communication unit. In some embodiments, the first wireless communication unit utilizes a wireless communication protocol using an electromagnetic carrier wave (e.g., a radio wave, microwave, or an infrared carrier wave) or ultrasound to transfer data from the interface to the computing device including Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 the second wireless communication unit. Brain-computer interfaces are commercially available, including the Neuroport™ system from Blackrock Microsystems (Salt Lake City, Utah), See also, e.g., Weiss et al. (2019) Brain-Computer Interfaces 6:106-117; herein incorporated by reference. Aspects of the present disclosure further include non-transitory computer readable storage mediums having instructions for practicing the subject methods. Computer readable storage mediums may be employed on one or more computers for complete automation or partial automation of a system for practicing methods described herein. In certain embodiments, instructions in accordance with the method described herein can be coded onto a computer- readable medium in the form of “programming”, where the term "computer readable medium" as used herein refers to any non-transitory storage medium that participates in providing instructions and data to a computer for execution and processing. Examples of suitable non- transitory storage media include a floppy disk, hard disk, optical disk, magneto-optical disk, CD- ROM, CD-R, magnetic tape, non-volatile memory card, ROM, DVD-ROM, Blue-ray disk, solid state disk, and network attached storage (NAS), whether or not such devices are internal or external to the computer. A file containing information can be “stored” on computer readable medium, where “storing” means recording information such that it is accessible and retrievable at a later date by a computer. The computer-implemented method described herein can be executed using programming that can be written in one or more of any number of computer programming languages. Such languages include, for example, Python, Java, Java Script, C, C#, C++, Go, R, Swift, PHP, as well as many others. The non-transitory computer readable storage medium may be employed on one or more computer systems having a display and operator input device. Operator input devices may, for example, be a keyboard, mouse, or the like. The processing module includes a processor which has access to a memory having instructions stored thereon for performing the steps of the subject methods. The processing module may include an operating system, a graphical user interface (GUI) controller, a system memory, memory storage devices, input-output controllers, cache memory, a data backup unit, and many other devices. The processor may be a commercially available processor, or it may be one of other processors that are or will become available. The processor executes the operating system and the operating system interfaces with firmware and hardware in a well-known manner, and facilitates the processor in coordinating and executing the Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 functions of various computer programs that may be written in a variety of programming languages, such as those mentioned above, other high level or low-level languages, as well as combinations thereof, as is known in the art. The operating system, typically in cooperation with the processor, coordinates and executes functions of the other components of the computer. The operating system also provides scheduling, input-output control, file and data management, memory management, and communication control and related services, all in accordance with known techniques. Components of systems for carrying out the presently disclosed methods are further described in the examples below. KITS Kits are also provided for carrying out the methods described herein. In some embodiments, the kit includes software for carrying out the computer implemented methods for decoding a speech sound or avatar action from recorded brain electrical signal data associated with attempted speech and / or attempted action by a subject, as described herein. In some embodiments, the kit includes a system for assisting a subject with communication as described herein. Such a system may include: a neural recording device including an electrode adapted for positioning at a location in a sensorimotor cortex region of the subject to record brain electrical signal data associated with attempted speech and / or attempted action by the subject; a processor programmed to decode a speech sound or avatar action from the recorded brain electrical signal data according to a computer implemented method described herein; an interface capable of communication with a computing device, said interface adapted for positioning at a location on the head of the subject, wherein the interface receives the brain electrical signal data from the neural recording device and transmits the brain electrical signal data to the processor; and a display component for displaying the speech sound or avatar action decoded from the recorded brain electrical signal data. In addition, the kits may further include (in certain embodiments) instructions for practicing the subject methods. These instructions may be present in the subject kits in a variety of forms, one or more of which may be present in the kit. For example, instructions may be present as printed information on a suitable medium or substrate, e.g., a piece or pieces of paper on which the information is printed, in the packaging of the kit, in a package insert, and the like. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 Another form of these instructions is a computer readable medium, e.g., diskette, compact disk (CD), flash drive, and the like, on which the information has been recorded. Yet another form of these instructions that may be present is a website address which may be used via the internet to access the information at a removed site. UTILITY The methods, devices, and systems of the invention find use in assisting individuals with communication. In particular, methods, devices, and systems are provided that facilitate full, embodied communication to people living with severe paralysis by restoring the ability to produce speech sounds and facial movements related to speaking, as well as by restoring the ability to perform non-speech communicative gestures. In some embodiments, the methods, devices, and systems of the present disclosure find use in restoring aspects of an individual’s personhood and identity by providing them the ability to communicate with naturalistic speed and expressivity, and by providing them with highly personalizable audio-visual representations. Embodiments of the present disclosure enable an individual with a communication or mobility disorder to interface with evolving technology to communicate with family and friends, facilitate community involvement and occupational participation, and engage in virtual, internet-based social contexts (such as social media, videogames, and metaverses). In some embodiments, the subject methods, devices, and systems find use in applications where it is desirable to increase the independence of an individual with a communication or mobility disorder. The methods, devices, and systems disclosed herein may be used to assist individuals who have difficulty with communication (e.g., using word-based and / or action-based communication) caused by conditions and diseases including, without limitation, anarthria, strokes, traumatic brain injuries, brain tumors, amyotrophic lateral sclerosis, multiple sclerosis, Huntington's disease, Niemann-Pick disease, Friedreich's ataxia, Wilson's disease, cerebral palsy, Guillain-Barré syndrome, Tay-Sachs disease, encephalopathy, central pontine myelinolysis, and other conditions causing dysfunction or paralysis of the muscles of the head, neck, arms, or chest resulting in anarthria. The methods disclosed herein may be used to restore communication to such individuals and improve autonomy and quality of life. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 EXAMPLES OF NON-LIMITING ASPECTS OF THE DISCLOSURE Aspects, including embodiments, of the present subject matter described above may be beneficial alone or in combination, with one or more other aspects or embodiments. Without limiting the foregoing description, certain non-limiting aspects of the disclosure numbered 1-406 are provided below. As will be apparent to those of skill in the art upon reading this disclosure, each of the individually numbered aspects may be used or combined with any of the preceding or following individually numbered aspects. This is intended to provide support for all such combinations of aspects and is not limited to combinations of aspects explicitly provided below: 1. A method of assisting a subject with communication, the method comprising: positioning a neural recording device comprising an electrode at a location in a sensorimotor cortex region of the brain of the subject to record brain electrical signal data associated with attempted speech by the subject; positioning an interface in communication with a computing device at a location on the head of the subject, wherein the interface is connected to the neural recording device; recording the brain electrical signal data associated with attempted speech by the subject using the neural recording device, wherein the interface receives the brain electrical signal data from the neural recording device and transmits the brain electrical signal data to a processor of the computing device; and decoding one or more speech sounds from the recorded brain electrical signal data using the processor, wherein the processor is programmed to use a machine learning model for the decoding. 2. The method of aspect 1, wherein the subject has difficulty with said communication because of anarthria, a stroke, a traumatic brain injury, a brain tumor, or amyotrophic lateral sclerosis. 3. The method of aspect 1 or 2, wherein the subject is paralyzed. 4. The method of any of aspects 1-3, wherein the subject has a speech intelligibility of 10% or less for prompted words. 5. The method of any of aspects 1-4, wherein the location of the neural recording device is on a surface of the sensorimotor cortex region or within the sensorimotor cortex region. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 6. The method of any of aspects 1-5, wherein the electrode of the neural recording device is positioned on the pial surface of the sensorimotor cortex. 7. The method of any of aspects 1-6, wherein the neural recording device is positioned such that the recording device covers regions associated with speech production and language perception. 8. The method of aspect 7, wherein the covered regions include a middle portion of the superior and / or middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus. 9. The method of any of aspects 1-8, wherein the neural recording device is centered on the central sulcus. 10. The method of any of aspects 1-9, wherein the neural recording device comprises an electrocorticography (ECoG) electrode array. 11. The method of aspect 10, wherein the ECoG electrode array is a high-density array. 12. The method of aspect 11, wherein the high-density array comprises 200 electrodes or more. 13. The method of any of aspects 10-12, wherein the electrode array comprises non- penetrating surface electrodes. 14. The method of any of aspects 1-13, wherein the interface comprises a percutaneous pedestal connector attached to the subject's cranium. 15. The method of any of aspects 1-14, wherein the interface further comprises a headstage connected to the percutaneous pedestal connector. 16. The method of aspect 15, wherein the headstage processes and digitizes the brain electrical signal data before transmitting the data to the processor. 17. The method of any of aspects 1-16, wherein the electrical signal data comprises high- gamma frequency content features. 18. The method of any of aspects 1-17, wherein the electrical signal data comprises neural oscillations in a range from 70 Hz to 150 Hz. 19. The method of any of aspects 1-18, wherein the electrical signal data comprises low- frequency signals. 20. The method of aspect 19, wherein the electrical signal data comprises neural oscillations in a range from 0.3 Hz to 17 Hz. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 21. The method of any of aspects 1-20, wherein the brain electrical signal data is recorded from a sensorimotor cortex region selected from the precentral gyrus, postcentral gyrus, superior temporal gyrus, middle temporal gyrus, or any combination thereof. 22. The method of any of aspects 1-21, wherein the method further comprises mapping the brain of the subject to identify an optimal location for positioning the electrode for recording the brain electrical signals associated with the attempted speech by the subject. 23. The method of any of aspects 1-22, wherein the one or more speech sounds form a word. 24. The method of aspect 23, wherein the one or more speech sounds form a sentence. 25. The method of aspects 23 or 24, wherein the subject is limited to a specified word set for the attempted speech. 26. The method of aspect 25, wherein the word set comprises words for expressing basic concepts and / or caregiving needs. 27. The method of aspects 25 or 26, wherein the word set comprises 100 words or more. 28. The method of aspect 27, wherein the word set comprises 350 words or more. 29. The method of aspect 28, wherein the word set comprises 1000 words or more. 30. The method of any of aspects 25-29, wherein the subject is limited to two or more word sets for the attempted speech. 31. The method of aspect 30, wherein the subject may switch between word sets. 32. The method of any of aspects 1-31, wherein the decoding machine learning model comprises a neural network. 33. The method of aspect 32, wherein the neural network comprises one or more convolutional layers. 34. The method of aspects 32 or 33, wherein the neural network is bidirectional. 35. The method of aspect 34, wherein the neural network is a recurrent neural network (RNN). 36. The method of aspect 35, wherein the RNN comprises gated recurrent units (GRUs) or long-short term memory (LSTM). 37. The method of aspect 36, wherein the RNN comprises GRUs. 38. The method of any of aspects 1-37, wherein one or more discrete speech units of a set of discrete speech units are decoded from the recorded brain electrical signal data. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 39. The method of aspect 38, wherein each speech sound is decoded from one or more discrete speech units. 40. The method of aspects 38 or 39, wherein the set of discrete speech units comprises 50 or more discrete speech units. 41. The method of aspect 40, wherein the set of discrete speech units comprises 100 or more discrete speech units. 42. The method of any of aspects 36-39, wherein discrete speech units are continuously decoded at a uniform frequency. 43. The method of aspect 42, wherein the frequency is 50 Hz or more. 44. The method of aspect 43, wherein the frequency is 200 Hz or more. 45. The method of any of aspects 38-44, wherein the set of discrete speech units are generated using an encoding machine learning model. 46. The method of aspect 45, wherein the encoding machine learning model comprises a neural network. 47. The method of aspect 46, wherein the neural network comprises one or more convolutional layers. 48. The method of aspects 46 or 47, wherein the neural network is bidirectional. 49. The method of aspect 48, wherein the neural network comprises a transformer encoder. 50. The method of aspect 49, wherein the machine learning model is a Hidden-Unit Bidirectional Encoder Representations from Transformers (HuBERT) model. 51. The method of any of aspects 45-50, wherein the set of discrete speech units are generated by training the encoding machine learning model. 52. The method of aspect 51, wherein the training is self-supervised training. 53. The method of any of aspects 45-52, wherein the method further comprises: obtaining reference electronic speech waveforms for a plurality of phrases; encoding each reference speech waveform into a temporal sequence of discrete speech units using the encoding model; recording the brain electrical signal data associated with attempted speech by the subject for each one of the plurality of phrases; training the decoding machine learning model to predict the most likely discrete speech unit associated with a segment of electrical signal data using the speech units derived from the Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 reference speech waveforms and the brain electrical signal data associated with attempted speech for each one of the plurality of phrases. 54. The method of aspect 53, wherein the reference speech waveforms are obtained from a recruited speaker. 55. The method of aspect 53, wherein the reference speech waveforms are obtained using a text-to-speech algorithm. 56. The method of any of aspects 53-55, wherein the decoding machine learning model is trained to learn mappings between neural activity patterns of electrical signals in the brain electrical signal data and the discrete speech units. 57. The method of any of aspects 53-56, wherein the training uses a CTC loss function. 58. The method of any of aspects 39-57, wherein each speech sound is decoded from one or more discrete speech units using a speech synthesizer. 59. The method of aspect 58, wherein the speech synthesizer comprises a machine learning model. 60. The method of aspect 59, wherein the synthesizing machine learning model comprises a neural network. 61. The method of aspect 60, wherein the neural network comprises one or more convolutional layers. 62. The method of aspects 60 or 61, wherein the neural network is bidirectional. 63. The method of aspect 62, wherein the neural network is an RNN. 64. The method of aspect 63, wherein the RNN comprises GRUs or LSTM. 65. The method of aspect 64, wherein the RNN comprises one or more LSTM layers. 66. The method of any of aspects 60-65, wherein the neural network comprises an attention mechanism. 67. The method of any of aspects 59-66, wherein the synthesizing machine learning model is configured to generate a spectrogram from the one or more discrete speech units. 68. The method of aspect 67, wherein the spectrogram is a mel spectrogram. 69. The method of aspects 67 or 68, wherein the speech synthesizer further comprises a vocoder configured to synthesize an electronic speech waveform from the spectrogram. 70. The method of aspect 69, wherein the vocoder comprises a machine learning model. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 71. The method of aspect 70, wherein the vocoder machine learning model comprises a neural network. 72. The method of aspect 71, wherein the neural network is an RNN. 73. The method of any of aspects 69-72, wherein the method further comprises transforming the electronic speech waveform into a personalized electronic speech waveform. 74. The method of aspect 73, wherein the personalized electronic speech waveform resembles speech in the subject’s own voice. 75. The method of aspects 73 or 74, wherein the transforming is performed using a machine learning model. 76. The method of aspect 75, wherein the transforming machine learning model comprises a neural network. 77. The method of aspect 76, wherein the neural network is based on transformer architecture. 78. The method of any of aspects 69-77, wherein the method further comprises converting the electronic speech waveform into an audible speech waveform. 79. The method of aspect 78, wherein the electronic speech waveform is converted into an audible speech waveform using a loudspeaker. 80. The method of any of the preceding aspects, wherein the processor is provided by a computer or handheld device. 81. The method of aspect 80, wherein the handheld device is a cell phone or a tablet. 82. A non-transitory computer-readable medium comprising program instructions that, when executed by a processor, cause the processor to perform the method of any one of aspects 1-81. 83. A kit comprising the non-transitory computer-readable medium of aspect 82 and instructions for decoding brain electrical signal data associated with attempted speech by a subject. 84. A system for decoding speech sounds from recorded brain electrical signal data configured to perform the method according to any of Aspects 1-81. 85. A computer implemented method for decoding speech audio from recorded brain electrical signal data associated with attempted speech by a subject, the computer performing steps comprising: Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 receiving the recorded brain electrical signal data associated with the attempted speech by the subject; and decoding one or more speech sounds from the recorded brain electrical signal data using a machine learning model. 86. The computer implemented method of aspect 85, wherein the decoding machine learning model comprises a neural network. 87. The computer implemented method of aspect 86, wherein the neural network comprises one or more convolutional layers. 88. The computer implemented method of aspects 86 or 87, wherein the neural network is bidirectional. 89. The computer implemented method of aspect 88, wherein the neural network is a recurrent neural network (RNN). 90. The computer implemented method of aspect 89, wherein the RNN comprises gated recurrent units (GRUs) or long-short term memory (LSTM). 91. The computer implemented method of aspect 90, wherein the RNN comprises GRUs. 92. The computer implemented method of any of aspects 85-91, wherein one or more discrete speech units of a set of discrete speech units are decoded from the recorded brain electrical signal data. 93. The computer implemented method of aspect 92, wherein each speech sound is decoded from one or more discrete speech units. 94. The computer implemented method of aspects 92 or 93, wherein the set of discrete speech units comprises 50 or more discrete speech units. 95. The computer implemented method of aspect 94, wherein the set of discrete speech units comprises 100 or more discrete speech units. 96. The computer implemented method of any of aspects 92-95, wherein discrete speech units are continuously decoded at a uniform frequency. 97. The computer implemented method of aspect 96, wherein the frequency is 50 Hz or more. 98. The computer implemented method of aspect 97, wherein the frequency is 200 Hz or more. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 99. The computer implemented method of any of aspects 92-98, wherein the set of discrete speech units are generated using an encoding machine learning model. 100. The computer implemented method of aspect 99, wherein the encoding machine learning model comprises a neural network. 101. The computer implemented method of aspect 100, wherein the neural network comprises one or more convolutional layers. 102. The computer implemented method of aspects 100 or 101, wherein the neural network is bidirectional. 103. The computer implemented method of aspect 102, wherein the neural network comprises a transformer encoder. 104. The computer implemented method of aspect 103, wherein the machine learning model is a Hidden-Unit Bidirectional Encoder Representations from Transformers (HuBERT) model. 105. The computer implemented method of any of aspects 99-104, wherein the set of discrete speech units are generated by training the encoding machine learning model. 106. The computer implemented method of aspect 105, wherein the training is self-supervised training. 107. The computer implemented method of any of aspects 99-106, wherein the method further comprises: receiving reference electronic speech waveforms for a plurality of phrases; encoding each reference speech waveform into a temporal sequence of discrete speech units using the encoding model; receiving brain electrical signal data associated with attempted speech by the subject for each one of the plurality of phrases; training the decoding machine learning model to predict the most likely discrete speech unit associated with a segment of electrical signal data using the speech units derived from the reference speech waveforms and the brain electrical signal data associated with attempted speech for each one of the plurality of phrases. 108. The computer implemented method of aspect 107, wherein the reference speech waveforms are generated using a text-to-speech algorithm. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 109. The computer implemented method of aspects 107 or 108, wherein the decoding machine learning model is trained to learn mappings between neural activity patterns of electrical signals in the brain electrical signal data and the discrete speech units. 110. The computer implemented method of any of aspects 93-109, wherein each speech sound is decoded from one or more discrete speech units using a speech synthesizer. 111. The computer implemented method of aspect 110, wherein the speech synthesizer comprises a machine learning model. 112. The computer implemented method of aspect 111, wherein the synthesizing machine learning model comprises a neural network. 113. The computer implemented method of aspect 112, wherein the neural network comprises one or more convolutional layers. 114. The computer implemented method of aspects 112 or 113, wherein the neural network is bidirectional. 115. The computer implemented method of aspect 114, wherein the neural network is an RNN. 116. The computer implemented method of aspect 115, wherein the RNN comprises GRUs or LSTM. 117. The computer implemented method of aspect 116, wherein the RNN comprises one or more LSTM layers. 118. The computer implemented method of any of aspects 112-117, wherein the neural network comprises an attention mechanism. 119. The computer implemented method of any of aspects 111-118, wherein the synthesizing machine learning model is configured to generate a spectrogram from the one or more discrete speech units. 120. The computer implemented method of aspect 119, wherein the spectrogram is a mel spectrogram. 121. The computer implemented method of aspects 119 or 120, wherein the speech synthesizer further comprises a vocoder configured to synthesize an electronic speech waveform from the spectrogram. 122. The computer implemented method of aspect 121, wherein the vocoder comprises a machine learning model. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 123. The computer implemented method of aspect 122, wherein the vocoder machine learning model comprises a neural network. 124. The computer implemented method of aspect 123, wherein the neural network is an RNN. 125. The computer implemented method of any of aspects 121-124, wherein the method further comprises transforming the electronic speech waveform into a personalized electronic speech waveform. 126. The computer implemented method of aspect 125, wherein the personalized electronic speech waveform resembles speech in the subject’s own voice. 127. The computer implemented method of aspects 125 or 126, wherein the transforming is performed using a machine learning model. 128. The computer implemented method of aspect 127, wherein the transforming machine learning model comprises a neural network. 129. The computer implemented method of aspect 128, wherein the neural network is based on transformer architecture. 131. A non-transitory computer-readable medium comprising program instructions that, when executed by a processor, cause the processor to perform the computer implemented method of any one of aspects 85-129. 132. A kit comprising the non-transitory computer-readable medium of aspect 131 and instructions for decoding brain electrical signal data associated with attempted speech by a subject. 133. A system for producing speech audio directly from neural activity, the system comprising: a neural recording device comprising an electrode adapted for positioning at a location in a sensorimotor cortex region of the brain of the subject to record brain electrical signal data associated with attempted speech by the subject; a processor programmed to use a machine learning model to decode one or more speech sounds from the recorded brain electrical signal data; an interface in communication with a computing device, said interface adapted for positioning at a location on the head of the subject, wherein the interface receives the brain Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 electrical signal data from the neural recording device and transmits the brain electrical signal data to the processor; and an audio speaker for playing the one or more speech sounds from the recorded brain electrical signal data. 134. The system of aspect 133, wherein the subject has difficulty with said speech because of anarthria, a stroke, a traumatic brain injury, a brain tumor, or amyotrophic lateral sclerosis. 135. The system of any of aspects 133-134, wherein the neural recording device is adapted for positioning on a surface of the sensorimotor cortex region or within the sensorimotor cortex region. 136. The system of any of aspects 133-135, wherein the electrode of the neural recording device is adapted for positioning on the pial surface of the sensorimotor cortex. 137. The system of any of aspects 133-136, wherein the neural recording device is adapted to be positioned such that the recording device covers regions associated with speech production and language perception. 138. The system of aspect 137, wherein the covered regions include a middle portion of the superior and / or middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus. 139. The system of any of aspects 133-138, wherein the neural recording device is adapted to be centered on the central sulcus. 140. The system of any of aspects 133-139, wherein the neural recording device comprises an electrocorticography (ECoG) electrode array. 141. The system of aspect 140, wherein the ECoG electrode array is a high-density array. 142. The system of aspect 141, wherein the high-density array comprises 200 electrodes or more. 143. The system of any of aspects 140-142, wherein the electrode array comprises non- penetrating surface electrodes. 144. The system of any of aspects 133-143, wherein the interface comprises a percutaneous pedestal connector attached to the subject's cranium. 145. The system of any of aspects 133-144, wherein the interface further comprises a headstage connected to the percutaneous pedestal connector. 146. The system of aspect 145, wherein the headstage is configured to process and digitize the brain electrical signal data before transmitting the data to the processor. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 147. The system of any of aspects 133-146, wherein the electrical signal data comprises high- gamma frequency content features. 148. The system of any of aspects 133-147, wherein the electrical signal data comprises neural oscillations in a range from 70 Hz to 150 Hz. 149. The system of any of aspects 133-148, wherein the electrical signal data comprises low- frequency signals. 150. The system of aspect 149, wherein the electrical signal data comprises neural oscillations in a range from 0.3 Hz to 17 Hz. 151. The system of any of aspects 133-150, wherein the brain electrical signal data is recorded from a sensorimotor cortex region selected from the precentral gyrus, postcentral gyrus, superior temporal gyrus, middle temporal gyrus, or any combination thereof. 153. The system of any of aspects 133-152, wherein the one or more speech sounds form a word. 154. The system of aspect 153, wherein the one or more speech sounds form a sentence. 155. The system of aspects 153 or 154, wherein the subject is limited to a specified word set for the attempted speech. 156. The system of aspect 155, wherein the word set comprises words for expressing basic concepts and / or caregiving needs. 157. The system of aspects 155 or 156, wherein the word set comprises 100 words or more. 158. The system of aspect 157, wherein the word set comprises 350 words or more. 159. The system of aspect 158, wherein the word set comprises 1000 words or more. 160. The system of any of aspects 155-159, wherein the subject is limited to two or more word sets for the attempted speech. 161. The system of aspect 160, wherein the processor is configured to switch between word sets. 162. The system of any of aspects 133-161, wherein the decoding machine learning model comprises a neural network. 163. The system of aspect 162, wherein the neural network comprises one or more convolutional layers. 164. The system of aspects 162 or 163, wherein the neural network is bidirectional. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 165. The system of aspect 164, wherein the neural network is a recurrent neural network (RNN). 166. The system of aspect 165, wherein the RNN comprises gated recurrent units (GRUs) or long-short term memory (LSTM). 167. The system of aspect 166, wherein the RNN comprises GRUs. 168. The system of any of aspects 133-167, wherein one or more discrete speech units of a set of discrete speech units are decoded from the recorded brain electrical signal data. 169. The system of aspect 168, wherein each speech sound is decoded from one or more discrete speech units. 170. The system of any of aspects 168-169, wherein discrete speech units are continuously decoded at a uniform frequency. 171. The system of aspect 170, wherein the frequency is 200 Hz or more. 172. The system of any of aspects 168-171, wherein the set of discrete speech units are generated using an encoding machine learning model. 173. The system of aspect 172, wherein the encoding machine learning model comprises a neural network. 174. The system of aspect 173, wherein the neural network comprises one or more convolutional layers. 175. The system of aspects 173 or 174, wherein the neural network is bidirectional. 176. The system of aspect 175, wherein the neural network comprises a transformer encoder. 177. The system of aspect 176, wherein the machine learning model is a Hidden-Unit Bidirectional Encoder Representations from Transformers (HuBERT) model. 178. The system of any of aspects 172-177, wherein the set of discrete speech units are generated by training the encoding machine learning model. 179. The system of aspect 178, wherein the training is self-supervised training. 180. The system of any of aspects 172-179, wherein the processor is further programmed to: obtain reference electronic speech waveforms for a plurality of phrases; encode each reference speech waveform into a temporal sequence of discrete speech units using the encoding model; record the brain electrical signal data associated with attempted speech by the subject for each one of the plurality of phrases; Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 train the decoding machine learning model to predict the most likely discrete speech unit associated with a segment of electrical signal data using the speech units derived from the reference speech waveforms and the brain electrical signal data associated with attempted speech for each one of the plurality of phrases. 181. The system of aspect 180, wherein the system further comprises a microphone configured to generate reference speech waveforms from a recruited speaker. 182. The system of aspect 180, wherein the processor further comprises a text-to-speech algorithm configured to generate the reference speech waveforms. 183. The system of any of aspects 180-182, wherein the decoding machine learning model is trained to learn mappings between neural activity patterns of electrical signals in the brain electrical signal data and the discrete speech units. 184. The system of any of aspects 172 to 183, wherein each speech sound is decoded from one or more discrete speech units using a speech synthesizer. 185. The system of aspect 184, wherein the speech synthesizer comprises a machine learning model. 186. The system of aspect 185, wherein the synthesizing machine learning model comprises a neural network. 187. The system of aspect 186, wherein the neural network comprises one or more convolutional layers. 188. The system of aspects 186 or 187, wherein the neural network is bidirectional. 189. The system of aspect 188, wherein the neural network is an RNN. 190. The system of aspect 189, wherein the RNN comprises GRUs or LSTM. 191. The system of aspect 190, wherein the RNN comprises one or more LSTM layers. 192. The system of any of aspects 186 to 191, wherein the neural network comprises an attention mechanism. 193. The system of any of aspects 185 to 192, wherein the synthesizing machine learning model is configured to generate a spectrogram from the one or more discrete speech units. 194. The system of aspect 193, wherein the spectrogram is a mel spectrogram. 195. The system of aspects 193 or 194, wherein the speech synthesizer further comprises a vocoder configured to synthesize an electronic speech waveform from the spectrogram. 196. The system of aspect 195, wherein the vocoder comprises a machine learning model. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 197. The system of aspect 196, wherein the vocoder machine learning model comprises a neural network. 198. The system of aspect 197, wherein the neural network is an RNN. 199. The system of any of aspects 195 to 198, wherein the processor is further programmed to transform the electronic speech waveform into a personalized electronic speech waveform. 200. The system of aspect 199, wherein the personalized electronic speech waveform resembles speech in the subject’s own voice. 201. The system of aspects 199 or 200, wherein the transforming is performed using a machine learning model. 202. The system of aspect 201, wherein the transforming machine learning model comprises a neural network. 203. The system of aspect 202, wherein the neural network is based on transformer architecture. 204. The system of any of aspects 195-203, wherein the processor is further programmed to convert the electronic speech waveform into an audible speech waveform. 205. The system of aspect 204, wherein the electronic speech waveform is converted into an audible speech waveform using a loudspeaker. 206. The system of any of aspects 133-205, wherein the processor is provided by a computer or handheld device. 207. The system of aspect 206, wherein the handheld device is a cell phone or a tablet. 208. A kit comprising the system of any of aspects 133-207 and instructions for decoding brain electrical signal data associated with attempted speech by a subject. 209. A method of controlling an electronic output device or system to perform one or more actions using brain electrical signals, the method comprising: positioning a neural recording device comprising an electrode at a location in a sensorimotor cortex region of the brain of the subject to record brain electrical signal data associated with an attempted action by the subject; positioning an interface in communication with a computing device at a location on the head of the subject, wherein the interface is connected to the neural recording device; recording the brain electrical signal data associated with the attempted action by the subject using the neural recording device, wherein the interface receives the brain electrical Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 signal data from the neural recording device and transmits the brain electrical signal data to a processor of the computing device; decoding one or more electronic output device or system actions from the recorded brain electrical signal data, wherein the processor is programmed to use a machine learning model for the decoding; and controlling the electronic output device or system to perform the one or more decoded electronic output device or system actions. 210. The method of aspect 209, wherein the subject has difficulty communicating because of anarthria, a stroke, a traumatic brain injury, a brain tumor, or amyotrophic lateral sclerosis. 211. The method of aspect 209 or 210, wherein the subject is paralyzed. 212. The method of any of aspects 209-211, wherein the subject is quadriplegic and / or experiences partial or total facial paralysis. 213. The method of any of aspects 209-212, wherein the location of the neural recording device is on a surface of the sensorimotor cortex region or within the sensorimotor cortex region. 214. The method of any of aspects 209-213, wherein the electrode of the neural recording device is positioned on the pial surface of the sensorimotor cortex. 215. The method of any of aspects 209-214, wherein the neural recording device is positioned such that the recording device covers regions associated with speech production and language perception. 216. The method of aspect 215, wherein the covered regions include a middle portion of the superior and / or middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus. 217. The method of any of aspects 209-216, wherein the neural recording device is centered on the central sulcus. 218. The method of any of aspects 209-217, wherein the neural recording device comprises an electrocorticography (ECoG) electrode array. 219. The method of aspect 218, wherein the ECoG electrode array is a high-density array. 220. The method of aspect 219, wherein the high-density array comprises 200 electrodes or more. 221. The method of any of aspects 218-220, wherein the electrode array comprises non- penetrating surface electrodes. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 222. The method of any of aspects 209-221, wherein the interface comprises a percutaneous pedestal connector attached to the subject's cranium. 223. The method of any of aspects 209-222, wherein the interface further comprises a headstage connected to the percutaneous pedestal connector. 224. The method of aspect 223, wherein the headstage processes and digitizes the brain electrical signal data before transmitting the data to the processor. 225. The method of any of aspects 209-224, wherein the electrical signal data comprises high- gamma frequency content features. 226. The method of any of aspects 209-225, wherein the electrical signal data comprises neural oscillations in a range from 70 Hz to 150 Hz. 227. The method of any of aspects 209-226, wherein the electrical signal data comprises low- frequency signals. 228. The method of aspect 227, wherein the electrical signal data comprises neural oscillations in a range from 0.3 Hz to 17 Hz. 229. The method of any of aspects 209-228, wherein the brain electrical signal data is recorded from a sensorimotor cortex region selected from the precentral gyrus, postcentral gyrus, superior temporal gyrus, middle temporal gyrus, or any combination thereof. 230. The method of any of aspects 209-229, wherein the method further comprises mapping the brain of the subject to identify an optimal location for positioning the electrode for recording the brain electrical signals associated with the attempted speech by the subject. 231. The method of any of aspects 209-230, wherein the subject is limited to a specified action set for the attempted action. 232. The method of aspect 231, wherein the action set comprises actions for expressing basic emotions and / or communicating caregiving needs. 233. The method of aspects 231 or 232, wherein the action set comprises 6 actions or more. 234. The method of aspect 233, wherein the action set comprises 100 actions or more. 235. The method of aspect 234, wherein the action set comprises 500 actions or more. 236. The method of any of aspects 231-235, wherein the subject is limited to two or more action sets for the attempted action. 237. The method of aspect 28236 wherein the subject may switch between action sets. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 238. The method of any of aspects 209-237, wherein the decoding machine learning model comprises a neural network. 239. The method of aspect 238, wherein the neural network comprises one or more convolutional layers. 240. The method of aspects 238 or 239, wherein the neural network is bidirectional. 241. The method of aspect 240, wherein the neural network is a recurrent neural network (RNN). 242. The method of aspect 241, wherein the RNN comprises gated recurrent units (GRUs) or long-short term memory (LSTM). 243. The method of aspect 242, wherein the RNN comprises GRUs. 244. The method of any of aspects 209-243, wherein each action of the device or system is discretized into one or more action representations. 245. The method of aspect 244, wherein the action representations are decoded from the recorded brain electrical signal data. 246. The method of aspect 245, wherein each action of the device or system is decoded from one or more action representations. 247. The method of aspect 246, wherein discrete action representations are continuously decoded from the recorded brain electrical signal data at a uniform frequency. 248. The method of aspect 247, wherein the frequency is 50 Hz or more. 249. The method of aspect 248, wherein the frequency is 200 Hz or more. 250. The method of any of aspects 244-249, wherein the discretization is performed using an autoencoder. 251. The method of aspect 250, wherein the autoencoder comprises one or more convolutional layers. 252. The method of aspects 250 or 251, wherein the autoencoder uses one or more rectified linear unit (ReLU) activations. 253. The method of aspect 252, wherein the autoencoder comprises a vector-quantized variational autoencoder (VQ-VAE). 254. The method of any of aspects 250-253, wherein the discretization occurs during training of the autoencoder. 255. The method of aspect 254, wherein the training is self-supervised training. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 256. The method of any of aspects 250-255, wherein the attempted action performed by the subject is different than the one or more electronic output device or system actions. 257. The method of aspect 256, wherein the attempted action performed by the subject is a hand gesture and the action of the electronic output device or system is powering on or off. 258. The method of any of aspects 250-255, wherein the attempted action performed by the subject corresponds to the one or more electronic output device or system actions. 259. The method of aspect 258, wherein the electronic output device or system is a prosthetic limb. 260. The method of aspect 258, wherein the electronic output device or system comprises a visual display and / or a loudspeaker. 261. The method of aspect 260, wherein the visual display and / or loudspeaker is configured to present a humanoid avatar. 262. The method of aspect 261, wherein the one or more electronic output device or system actions comprise actions performed by the avatar. 263. The method of aspect 262, wherein actions performed by the avatar include speech, orofacial movements for speech, and / or non-speech communicative gestures. 264. The method of aspect 263, wherein the speech orofacial movements comprise one or more of: a tongue tip raise, tongue retraction, tongue body raise, tongue advance, lip rounding, pinching nostril flare, upper lip pull, lower lip tuck, lower lip push, lower lip pull, lip flare, jaw opening, lip compression, and / or lip adduction. 265. The method of aspect 263, wherein the non-speech communicative gestures comprise the abduction, adduction, flexion, extension, and / or circumduction of one or more body parts. 266. The method of aspects 263 or 265, wherein the non-speech communicative gestures include emotional expressions using facial muscles. 267. The method of aspect 266, wherein the emotional expressions comprise happy, sad, and surprised expressions. 268. The method of any of aspects 262-267, wherein the method further comprises: obtaining reference avatar animations for a plurality of actions; encoding each reference animation into a temporal sequence of discrete action representations using the autoencoder; Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 recording the brain electrical signal data associated with attempted action by the subject for each one of the plurality of actions; training the decoding machine learning model to predict the most likely discrete action representation associated with a segment of electrical signal data using the action representations derived from the reference avatar animations and the brain electrical signal data associated with attempted action for each one of the plurality of actions. 269. The method of aspect 268, wherein the reference avatar animations are obtained from an avatar-animation system. 270. The method of any of aspects 268 or 268, wherein the decoding machine learning model is trained to learn mappings between neural activity patterns of electrical signals in the brain electrical signal data and the discrete action representations. 271. The method of any of aspects 268-270, wherein the training uses a CTC loss function. 272. The method of any of aspects 250-271, wherein each action of the avatar is decoded from one or more action representations using the autoencoder. 273. The method of any of aspects 268-272, wherein the plurality of actions includes at least two of speech, orofacial movements for speech, and / or non-speech communicative gestures. 274. The method of aspect 273, wherein the decoding machine learning model is trained to discriminate between actions performed by different regions of the body and / or between speech associated actions and non-speech communicative gestures. 275. The method of aspect 274, wherein the decoding machine learning model is trained to discriminate between finger-flexions and actions associated with attempted speech. 276. The method of any of aspects 268-275, wherein the machine learning model is trained to decode orofacial movements for speech using action representations derived from reference avatar animations of the avatar performing the orofacial movements and brain electrical signal data associated with orofacial movements attempted by the subject. 277. The method of any of aspects 268-275, wherein the avatar is controlled to perform orofacial movements based on speech decoded from the recorded brain electrical signal data. 278. The method of aspect 277, wherein the avatar is controlled to perform orofacial movements using a speech-to-gesture algorithm. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 279. The method of any of aspects 268-277, wherein a single decoding machine learning model is trained for speech, orofacial speech movement, and non-speech communicative gesture actions. 280. The method of any of aspects 268-277, wherein separate decoding machine learning models are trained for speech associated actions and non-speech communicative gestures. 281. The method of any of aspects 268-277 and 280, wherein separate decoding machine learning models are trained for speech and orofacial movements for speech. 282. The method of aspects 280 or 281, wherein the avatar is controlled using actions decoded from multiple machine learning models. 283. The method of aspect 282, wherein the actions decoded from the multiple machine learning models occur concurrently and the avatar is controlled to perform the actions simultaneously. 284. The method of any of aspects 273-283, wherein a decoded non-speech communicative gesture action affects how the avatar performs a decoded speech associated action. 285. The method of aspect 284, wherein a decoded emotional expression affects the inflection of decoded speech performed by the avatar. 286. The method of any of aspects 262-285, wherein the avatar is used by the subject to communicate with one or more individuals in person. 287. The method of any of aspects 262-286, wherein the avatar is used by the subject to communicate with one or more individuals in an interactive virtual environment. 288. The method of any of aspects 262-287, wherein the avatar is used by the subject to play a video game. 289. The method of any of aspects 262-288, wherein the avatar is used by the subject for therapy. 290. The method of aspect 289, wherein the avatar is used by the subject for physical therapy. 291. The method of aspect 290, wherein the physical therapy involves regaining mobility of a body part after an injury. 292. The method of any of aspects 262-291, wherein actions performed by the avatar control one or more electronic devices in the subject’s environment. 293. The method of aspect 292, wherein actions performed by the avatar control one or more electronic devices in the same room as the subject. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 294. The method of aspect 293, wherein actions performed by the avatar control one or smart home appliances. 295. The method of any of aspects 262-294, wherein the visual display comprises a computer monitor, a television, and / or a visual projection device. 296. The method of any of aspects 262-295, wherein the visual display comprises a virtual reality headset, goggles, or contacts. 297. The method of any of aspects 262-295, wherein the visual display comprises an augmented reality headset, goggles, or contacts. 298. The method of any of aspects 262-297, wherein the method further comprises generating a virtual environment for the avatar. 299. The method of aspect 298, wherein the action performed by the avatar is determined using the virtual environment. 300. The method of any of aspects 262-299, wherein previous actions performed by the avatar are used to determine the action performed by the avatar. 301. The method of any of aspects 209-300, wherein the processor is provided by a computer, a handheld device, or a headset. 302. The method of aspect 301, wherein the handheld device is a cell phone or a tablet. 303. A non-transitory computer-readable medium comprising program instructions that, when executed by a processor, cause the processor to perform the method of any one of aspects 209- 302. 304. A kit comprising the non-transitory computer-readable medium of aspect 303 and instructions for decoding brain electrical signal data associated with attempted speech by a subject. 305. A system for controlling an electronic device or system using brain electrical signals configured to perform the method according to any of Aspects 209-302. 306. A computer implemented method for controlling an electronic output device or system to perform one or more actions using brain electrical signals, the computer performing steps comprising: receiving the recorded brain electrical signal data associated with the attempted action by the subject; and Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 decoding one or more electronic output device or system actions from the recorded brain electrical signal data using a machine learning model. 307. The computer implemented method of aspect 306, wherein the decoding machine learning model comprises a neural network. 308. The computer implemented method of aspect 307, wherein the neural network comprises one or more convolutional layers. 309. The computer implemented method of aspects 307 or 308, wherein the neural network is bidirectional. 310. The computer implemented method of aspect 309, wherein the neural network is a recurrent neural network (RNN). 311. The computer implemented method of aspect 310, wherein the RNN comprises gated recurrent units (GRUs) or long-short term memory (LSTM). 312. The computer implemented method of aspect 311, wherein the RNN comprises GRUs. 313. The computer implemented method of any of aspects 306-312, wherein the subject is limited to a specified action set for the attempted action. 314. The computer implemented method of aspect 313, wherein the action set comprises actions for expressing basic emotions and / or communicating caregiving needs. 315. The computer implemented method of aspect 314, wherein the action set comprises 6 actions or more. 316. The computer implemented method of aspect 315, wherein the action set comprises 100 actions or more. 317. The computer implemented method of aspect 316, wherein the action set comprises 500 actions or more. 318. The computer implemented method of any of aspects 313-317, wherein the subject is limited to two or more action sets for the attempted action. 319. The computer implemented method of aspect 318, wherein the subject may switch between action sets. 320. The computer implemented method of any of aspects 306-319, wherein each action of the device or system is discretized into one or more action representations. 321. The computer implemented method of aspect 320, wherein the action representations are decoded from the recorded brain electrical signal data. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 322. The computer implemented method of aspect 321, wherein each action of the output device or system is decoded from one or more action representations. 323. The computer implemented method of any of aspects 320-322, wherein the discretization is performed using an autoencoder. 324. The computer implemented method of aspect 323, wherein the autoencoder comprises one or more convolutional layers. 325. The computer implemented method of aspects 323 or 324, wherein the autoencoder uses one or more rectified linear unit (ReLU) activations. 326. The computer implemented method of aspect 325, wherein the autoencoder comprises a vector-quantized variational autoencoder (VQ-VAE). 327. The computer implemented method of any of aspects 323-326, wherein the discretization occurs during training of the autoencoder. 328. The computer implemented method of aspect 327, wherein the training is self-supervised training. 329. The computer implemented method of any of aspects 306-328, wherein the electronic output device or system comprises a visual display and / or a loudspeaker. 330. The computer implemented method of aspect 329, wherein the visual display and / or loudspeaker is configured to present a humanoid avatar. 331. The computer implemented method of aspect 330, wherein the one or more electronic output device or system actions comprise actions performed by the avatar. 332. The computer implemented method of aspect 331, wherein actions performed by the avatar include speech, orofacial movements for speech, and / or non-speech communicative gestures. 333. The method of aspect 332, wherein the non-speech communicative gestures include emotional expressions using facial muscles. 334. The computer implemented method of any of aspects 331-333, wherein the method further comprises: receiving reference avatar animations for a plurality of actions; encoding each reference avatar animation into a temporal sequence of discrete action representations using the autoencoder; Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 receiving brain electrical signal data associated with the attempted action by the subject for each one of the plurality of actions; training the decoding machine learning model to predict the most likely discrete action representation associated with a segment of electrical signal data using the action representations derived from the reference avatar animations and the brain electrical signal data associated with the attempted action for each one of the plurality of actions. 335. The computer implemented method of aspect 334, wherein separate decoding machine learning models are trained for speech associated actions and non-speech communicative gestures. 336. The computer implemented method of aspects 334 or 335, wherein the avatar is used by the subject to communicate with one or more individuals in person. 337. The computer implemented method of any of aspects 334-336, wherein the avatar is used by the subject to communicate with one or more individuals in an interactive virtual environment. 338. The computer implemented method of aspect 337, wherein the method further comprises generating a virtual environment for the avatar. 339. The computer implemented method of aspect 338, wherein the action performed by the avatar is determined using the virtual environment. 340. The computer implemented method of any of aspects 331-339, wherein previous actions performed by the avatar are used to determine the action performed by the avatar. 341. A non-transitory computer-readable medium comprising program instructions that, when executed by a processor, cause the processor to perform the computer implemented method of any one of aspects 306-340. 342. A kit comprising the non-transitory computer-readable medium of aspect 341 and instructions for decoding brain electrical signal data associated with attempted speech by a subject. 343. A system for controlling an avatar to perform one or more actions using brain electrical signals, the system comprising: a neural recording device comprising an electrode adapted for positioning at a location in a sensorimotor cortex region of the brain of the subject to record brain electrical signal data associated with an attempted action by the subject; Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 a processor programmed to use a machine learning model to decode an avatar animation from the recorded brain electrical signal data; an interface in communication with a computing device, said interface adapted for positioning at a location on the head of the subject, wherein the interface receives the brain electrical signal data from the neural recording device and transmits the brain electrical signal data to the processor; and a display component for displaying the avatar action from the recorded brain electrical signal data. 344. The system of aspect 343, wherein the subject has difficulty with said speech because of anarthria, a stroke, a traumatic brain injury, a brain tumor, or amyotrophic lateral sclerosis. 345. The system of aspects 343 or 344, wherein the neural recording device is adapted for positioning on a surface of the sensorimotor cortex region or within the sensorimotor cortex region. 346. The system of any of aspects 343-345, wherein the electrode of the neural recording device is adapted for positioning on the pial surface of the sensorimotor cortex. 347. The system of any of aspects 343-345, wherein the neural recording device is adapted to be positioned such that the recording device covers regions associated with speech production and language perception. 348. The system of aspect 347, wherein the covered regions include a middle portion of the superior and / or middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus. 349. The system of any of aspects 343-348, wherein the neural recording device is adapted to be centered on the central sulcus. 350. The system of any of aspects 343-349, wherein the neural recording device comprises an electrocorticography (ECoG) electrode array. 351. The system of aspect 350, wherein the ECoG electrode array is a high-density array. 352. The system of aspect 351, wherein the high-density array comprises 200 electrodes or more. 353. The system of any of aspects 350-352, wherein the electrode array comprises non- penetrating surface electrodes. 354. The system of any of aspects 343-353, wherein the interface comprises a percutaneous pedestal connector attached to the subject's cranium. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 355. The system of any of aspects 343-354, wherein the interface further comprises a headstage connected to the percutaneous pedestal connector. 356. The system of aspect 355, wherein the headstage is configured to process and digitize the brain electrical signal data before transmitting the data to the processor. 357. The system of any of aspects 343-356, wherein the electrical signal data comprises high- gamma frequency content features. 358. The system of any of aspects 343-357, wherein the electrical signal data comprises neural oscillations in a range from 70 Hz to 150 Hz. 359. The system of any of aspects 343-358, wherein the electrical signal data comprises low- frequency signals. 360. The system of aspect 359, wherein the electrical signal data comprises neural oscillations in a range from 0.3 Hz to 17 Hz. 361. The system of any of aspects 343-360, wherein the brain electrical signal data is recorded from a sensorimotor cortex region selected from the precentral gyrus, postcentral gyrus, superior temporal gyrus, middle temporal gyrus, or any combination thereof. 362. The system of any of aspects 343-361, wherein the subject is limited to a specified action set for the attempted action. 363. The system of aspect 362, wherein the action set comprises actions for expressing basic emotions and / or communicating caregiving needs. 364. The system of aspects 362 or 363, wherein the action set comprises 6 actions or more. 365. The system of aspect 364, wherein the action set comprises 100 actions or more. 366. The system of any of aspects 362-365, wherein the subject is limited to two or more action sets for the attempted action. 367. The system of aspect 366, wherein the processor is configured to switch between action sets. 368. The system of any of aspects 343-367, wherein the decoding machine learning model comprises a neural network. 369. The system of aspect 368 wherein the neural network comprises one or more convolutional layers. 370. The system of aspects 368 or 369, wherein the neural network is bidirectional. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 371. The system of aspect 370, wherein the neural network is a recurrent neural network (RNN). 372. The system of aspect 371, wherein the RNN comprises gated recurrent units (GRUs) or long-short term memory (LSTM). 373. The system of aspect 372, wherein the RNN comprises GRUs. 374. The system of any of aspects 343-373, wherein the processor is further programmed to discretize each avatar action into one or more action representations. 375. The system of aspect 374, wherein the action representations are decoded from the recorded brain electrical signal data. 376. The system of aspect 375, wherein each avatar action is decoded from one or more action representations. 377. The system of any of aspects 374-376, wherein the discretization is performed using an autoencoder. 378. The system of aspect 377, wherein the autoencoder comprises one or more convolutional layers. 379. The system of aspects 377 or 378, wherein the autoencoder uses one or more rectified linear unit (ReLU) activations. 380. The system of aspect 379, wherein the autoencoder comprises a vector-quantized variational autoencoder (VQ-VAE). 381. The system of any of aspects 377-380, wherein the discretization occurs during training of the autoencoder. 382. The system of aspect 381, wherein the training is self-supervised training. 383. The system of any of aspects 343-382, wherein actions performed by the avatar include speech, orofacial movements for speech, and / or non-speech communicative gestures. 384. The system of aspect 383, wherein the speech orofacial movements comprise one or more of: a tongue tip raise, tongue retraction, tongue body raise, tongue advance, lip rounding, pinching nostril flare, upper lip pull, lower lip tuck, lower lip push, lower lip pull, lip flare, jaw opening, lip compression, and / or lip adduction. 385. The system of aspect 383, wherein the non-speech communicative gestures comprise the abduction, adduction, flexion, extension, and / or circumduction of one or more body parts. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 386. The system of aspects 383 or 385, wherein the non-speech communicative gestures include emotional expressions using facial muscles. 387. The system of aspect 58, wherein the emotional expressions comprise happy, sad, and / or surprised expressions. 388. The system of any of aspects 377-387, wherein the processor is further programmed to: obtain reference avatar animations for a plurality of actions; encode each reference animation into a temporal sequence of discrete action representations using the autoencoder; record brain electrical signal data associated with the attempted action by the subject for each one of the plurality of actions; train the decoding machine learning model to predict the most likely discrete action representation associated with a segment of electrical signal data using the action representations derived from the reference avatar animations and the brain electrical signal data associated with attempted action for each one of the plurality of actions. 389. The system of aspect 388, wherein the decoding machine learning model is trained to learn mappings between neural activity patterns of electrical signals in the brain electrical signal data and the discrete action representations. 390. The system of aspects 388 or 389, wherein the training uses a CTC loss function. 391. The system of any of aspects 377-390, wherein each action of the avatar is decoded from one or more action representations using the autoencoder. 392. The system of any of aspects 388-391, wherein the plurality of actions includes at least two of speech, orofacial movements for speech, and / or non-speech communicative gestures. 393. The system of aspect 392, wherein the decoding machine learning model is trained to discriminate between actions performed by different regions of the body and / or between speech associated actions and non-speech communicative gestures. 394. The system of any of aspects 377-393, wherein separate decoding machine learning models are trained for speech associated actions and non-speech communicative gestures. 395. The system of aspect 394, wherein separate decoding machine learning models are trained for speech and orofacial movements for speech. 396. The system of aspects 394 or 395, wherein the avatar is controlled using actions decoded from multiple machine learning models. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 397. The system of aspect 396, wherein the actions decoded from the multiple machine learning models occur concurrently and the avatar is controlled to perform the actions simultaneously. 398. The system of any of aspects 343-397, wherein the visual display comprises a computer monitor, a television, and / or a visual projection device. 399. The system of any of aspects 343-398, wherein the visual display comprises a virtual reality headset, goggles, or contacts. 400. The system of any of aspects 343-399, wherein the visual display comprises an augmented reality headset, goggles, or contacts. 401. The system of any of the aspects 343-400, wherein the processor is further programmed to generate a virtual environment for the avatar. 402. The system of aspect 401, wherein the action performed by the avatar is determined using the virtual environment. 403. The system of any of the aspects 343-402, wherein previous actions performed by the avatar are used to determine the action performed by the avatar. 404. The system of any of aspects 343-403, wherein the processor is provided by a computer, a handheld device, or a headset. 405. The system of aspect 404, wherein the handheld device is a cell phone or a tablet. 406. A kit comprising the system of any of aspects 343-405 and instructions for decoding brain electrical signal data associated with attempted speech by a subject.
[0002] Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 EXPERIMENTAL As demonstrated in the above disclosure, the present invention has a wide variety of applications. The following examples are put forth so as to provide those of ordinary skill in the art with a complete disclosure and description of how to make and use the present invention and are not intended to limit the scope of what the inventors regard as their invention nor are they intended to represent that the experiments below are all or the only experiments performed. Those of skill in the art will readily recognize a variety of noncritical parameters that could be changed or modified to yield essentially similar results. Efforts have been made to ensure accuracy with respect to numbers used (e.g. amounts, percentages, etc.) but some experimental errors and deviations should be accounted for. Example 1: A High-Performance Neuroprosthesis for Speech Decoding and Avatar Control 1.1. Overview Speech neuroprostheses have the potential to restore communication to people living with paralysis, but naturalistic speed and expressivity remain elusive. Here, high-density cortical- surface recordings were used in a clinical-trial participant with severe limb and vocal paralysis to achieve high-performance real-time decoding across three complementary speech-related output modalities: text, speech audio, and facial-avatar animation. Deep-learning models were trained and evaluated using neural data collected as the participant attempted to silently speak sentences. For text, accurate and rapid large-vocabulary decoding was demonstrated with a median rate of 77.6 words per minute and median word error rate of 25.5%. For speech sounds, intelligible speech synthesis of high-utility phrases was demonstrated, with untrained listeners achieving a median perceptual word error rate of 29.2% during transcription. For facial avatar, the control of virtual orofacial movements for speech and non-speech communicative gestures was demonstrated. The decoders reached high performance with fewer than two weeks of training. The findings introduce a new multimodal speech-neuroprosthetic approach that has significant promise to restore full, embodied communication to people living with severe paralysis. 1.2. Introduction Speech is the ability to express thoughts and ideas through spoken words. Speech loss after neurological injury is devastating because it significantly impairs communication and Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 causes social isolation [1]. Previous demonstrations have shown that it is possible to decode speech from the brain activity of a person with paralysis, but only in the form of text and with limited speed and vocabulary [2] [3]. A compelling goal is to both enable faster large-vocabulary text-based communication and restore the produced speech sounds and facial movements related to speaking. While text outputs are good for basic communication and messaging, speaking has rich prosody, expressiveness, and identity that can enhance embodied communication beyond what can be conveyed in text alone. To address this, a multimodal speech neuroprosthesis that uses broad-coverage, high-density electrocorticography was designed to decode text and audio- visual speech outputs from articulatory vocal-tract representations distributed throughout the sensorimotor cortex. Due to severe paralysis caused by a brainstem stroke that occurred over 18 years ago, the participant cannot speak or vocalize speech sounds given the severe weakness of her orofacial muscles (anarthria) and cannot type given the weakness in her arms and hands (quadriplegia). Instead, she uses commercially available assistive technology to communicate, primarily relying on a head-tracking interface to generate intended messages at about 14 words per minute. A speech-decoding system was designed that enabled a clinical-trial participant (ClinicalTrials.gov; NCT03698149) with severe paralysis and anarthria to communicate by decoding intended sentences from signals acquired by a 253-channel high-density electrocorticography (ECoG) array implanted over her sensorimotor cortex (FIG. 1 and FIGS. 2A to 2B). The array was positioned over cortical areas relevant for orofacial movements, and simple movement tasks demonstrated differentiable activations associated with attempted movements of the lips, tongue, and jaw (FIG. 2C). For speech decoding, the participant was presented with a sentence as a text prompt on a screen and was instructed to silently attempt to say the sentence to the best of her ability after a visual go cue. Meanwhile, neural signals recorded from all 253 ECoG electrodes were processed to extract high-gamma activity (HGA; between 70–150 Hz) and low-frequency signals (LFS; between 0.3–17 Hz) [3]. Deep-learning models were trained to learn mappings between these ECoG features and phones, speech-sound features, and articulatory gestures, which were then used to output text, synthesize speech audio, and animate a virtual avatar, respectively (FIG. 1). The system was evaluated using three custom sentence sets containing varying amounts of unique words and sentences named “50-phrase-AAC,” “529-phrase-AAC,” and “1024-word- Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 General.” The first two sets closely mirror corpora preloaded on commercially available augmentative and alternative communication (AAC) devices, designed to let patients express basic concepts and caregiving needs [4]. These two sets were chosen to assess the ability to decode high-utility sentences at a limited and expanded vocabulary level. The 529-phrase-AAC set contained 529 sentences composed of 372 unique words, and 50 high-utility sentences composed of 119 unique words were selected to create the 50-phrase-AAC set. To evaluate how well the system performed with a larger vocabulary containing common English words, the 1024-word-General set was created, containing 9,512 sentences composed of 1,024 unique words sampled from Twitter and movie transcriptions. This set was primarily used to assess how well the decoders could generalize to sentences that the participant did not attempt to say during training with a vocabulary size large enough to facilitate general-purpose communication. To train the neural-decoding models prior to real-time testing, ECoG data was recorded as the participant silently attempted to speak individual sentences. Learning statistical mappings between the ECoG features and the sequences of phones and speech sound features in the sentences was challenged by the absence of clear timing information in the attempted speech. To overcome the inability to definitively know when the phones and speech units began and ended, a connectionist temporal classification (CTC) loss function was used during training of the neural decoders, which is commonly used in automatic speech recognition to infer sequences of sub-word units (such as phones or letters) from speech waveforms when precise time alignment between the units and the waveforms is unknown [5]. CTC loss was used during training of the text-decoding, speech-synthesis, and articulatory-decoding models to enable prediction of phone probabilities, discrete speech-sound units, and discrete articulator movements, respectively, from the ECoG signals. FIG. 1 and FIGS. 2A to 2C depict multimodal speech decoding in a participant with vocal-tract paralysis. FIG. 1 illustrates an overview of the speech-decoding pipeline. A brainstem-stroke survivor with anarthria was implanted with a 253-channel high-density electrocorticography (ECoG) array 18 years after injury as part of a clinical trial to decode speech. Neural activity was processed and used to train deep learning models to predict phone probabilities, speech-sound features, and articulatory gestures. These outputs were used to decode text, synthesize audible speech, and animate a virtual avatar, respectively. FIG. 2A provides a sagittal MRI showing brainstem atrophy (in the bilateral pons; red arrow) resulting Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 from stroke. FIG. 2B shows MRI reconstruction of the participant’s brain overlaid with the locations of implanted electrodes. The ECoG array was implanted over the participant’s lateral cortex, centered on the central sulcus. FIG. 2C shows an example of ECoG features resulting from attempted orofacial movements. The top illustrations depict simple articulatory movements attempted by the participant. The middle depictions illustrate electrode-activation maps demonstrating robust electrode tunings across articulators during attempted movements. Only the electrodes with the strongest responses (top 20%) are shown for each movement type. Color indicates the magnitude of the average evoked high-gamma activity (HGA) response with each type of movement. The bottom plots depict Z-scored trial-averaged evoked HGA responses with each movement type for each of the boxed electrodes in the electrode-activation maps. In each plot, each response trace shows mean + / - standard error across trials and is aligned to the peak activation time. 1.3. Results 1.3.1. Text decoding To evaluate real-time performance during the text task condition, text was decoded as the participant attempted to silently say 250 randomly selected sentences from the 1024-word- General set that were not used during model training (FIG. 3A, Video 1). To decode text, features extracted from ECoG signals were streamed starting 500 ms prior to the go cue into a bidirectional recurrent neural network (RNN). Prior to testing, the RNN was trained to predict the probabilities of 39 phones and silence at each time step. A CTC beam search then determined the most likely sentence given these probabilities. First, it created a set of candidate phone sequences that were constrained to form valid words within the 1,024-word vocabulary. Then, it evaluated candidate sentences by combining each candidate’s underlying phone probabilities with its linguistic probability using a natural-language model. To quantify text-decoding performance, standard metrics were used in automatic speech recognition: word error rate (WER), phone error rate (PER), character error rate (CER), and words per minute (WPM). WER, PER, and CER measure the percentage of decoded words, phones, and characters, respectively, that were incorrect. During real-time evaluation 250 randomly selected test sentences were decoded from the corpus that the participant did not attempt to produce for model training, and then error rates Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 were computed across sequential pseudo-blocks of 10-sentence segments. Sentences were decoded with a median PER of 18.5% (99% CI [14.9, 28.5]; FIG. 3B), a median WER of 25.5% (99% CI [19.3, 34.5]; FIG. 3C), and a median CER of 19.9% (99% CI [15.0, 30.1]; FIG. 3C; see Table 1 for example decodes). For all metrics, performance was better than chance, which was computed by re-evaluating performance after using temporally shuffled neural data as the input to the decoding pipeline (P < 0.0001 for all three comparisons, two-sided Wilcoxon rank-sum tests with 5-way Holm-Bonferroni correction). The average WER passes the 30% threshold below which speech-recognition applications generally become useful [6] while providing access to a large vocabulary of over 1,000 words, indicating that the approach may be viable in clinical applications. Furthermore, in offline simulations using the same neural data and decoder but with a modified language model and vocabulary containing 42,391 words, an offline WER of 29.8% was achieved (99% CI [21.1, 36.00]; FIG. 16) showing that the approach retains high performance in a large-vocabulary setting. WER Percentile Target sentence Decoded sentence(%) (%) You should have let me do the talking You should have let me do the talking 0 46.0 I think I need a little air I think I need a little air 0 46.0 Do you want to get some coffee Do you want to get some coffee 0 46.0 What do you get if you finish Why do you get if you finish 14 46.8 What do you want from us What do you want for us 17 51.6 You got your wish You get your wish 25 62.4 No tell me why So tell me why 25 62.4 You have no right to keep us here You have no right to be out here 25 62.4 Why would they come to me Why would they have to be 33 64.8 Why are you looking at me like that Why are you looking at that 38 66.4 All I told them was the truth Can I do that was the truth 43 71.6 You got it all in your head You got here all your right 43 71.6 I would like to watch television I was that no one television 67 86.0 Do you mind me talking about your stuff Do you make it out to yourself 75 89.2 Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 Table 1: Illustrative text decoding examples for the 1024-word-General set. Examples are shown for various levels of word error rate (WER) during real-time decoding with the 1024-word-General set. Each percentile value indicates the percent of decoded sentences that had a WER less than or equal to the WER of the provided example sentence. A median real-time decoding rate of 77.6 WPM was observed (99% CI [76.7, 78.4]; FIG. 4A) and a maximum rate of 83.7 WPM. This decoding rate exceeds the participant’s typical communication rate using her previous assistive device (14.2 WPM) and is closer to naturalistic speaking rates than has been previously reported with communication neuroprostheses [2] [3] [7- 9]. To assess how well the system could decode phones in the absence of a language model and constrained vocabulary, performance was evaluated using just the RNN neural-decoding model (using the most likely phone prediction at each time step instead of the CTC beam search) in an offline analysis. This yielded a median PER of 29.4% (99% CI [26.3, 33.0]; FIG. 3B). This median PER is only 10.9 percentage points higher than the full model, demonstrating that the primary contributor to phone-decoding performance was the neural-decoding RNN model and not the CTC beam search or language model (P < 0.0001 for all comparisons to chance and to the full model, two-sided Wilcoxon signed-rank tests with 5-way Holm-Bonferroni correction; Table 2). Table 2: Real-time text-decoding comparisons with the 1024- word-General sentence set. The relationship between quantity of training data and text-decoding performance was also characterized in offline analyses. For each day of data collection, models were trained on all the data collected on or before that date, then performance was simulated on the real-time blocks. Steadily declining error rates were observed over the course of 13 days of training-data collection (FIG. 4A), during which 9,512 sentence trials were collected at an average rate of Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 about 1.6 hours of training data per day. Overall, these results show that functional speech- decoding performance can be achieved after a relatively short period of data collection compared to prior work [2] [3] and is likely to continue to improve with more data
[0010] . To assess signal stability, real-time classification performance was measured during a separate NATO-motor task that was collected during each research session with the participant. In each trial of this task, the participant was prompted to attempt to either silently say one of the 26 code words from the NATO phonetic alphabet (alpha, bravo, charlie, and so forth) or perform one of four hand movements (described and analyzed below). A neural-network classifier was trained to predict the most likely NATO code word from a 4-second window of ECoG features (aligned to the task go cue), and real-time performance was evaluated with the classifier during the NATO-motor task (FIG. 4B, Video 2). The model continued to be retrained using available data prior to real-time testing in subsequent sessions until day 40, at which point the classifier was frozen after training it on data from the 1,196 available trials (46 per code word). Across 19 sessions after freezing the classifier, a mean classification accuracy of 96.8% (99% CI [94.5, 98.6]) was observed, with accuracies of 100% obtained on 8 of these sessions. Accuracy remained high after a 61 day pause in recording for the participant to travel. These results illustrate the stability of the cortical-surface neural interface and demonstrate that high performance can be achieved with relatively few training trials and without requiring recalibration. To evaluate model performance on closed sets with repeats across sentences and without pauses in the participant’s speech models were trained, then simulated text decoding with the AAC-50 (FIG. 17) and AAC-529 (FIG. 18) sentence sets (Method 1). With the AAC-529 set, a median WER of 17.1% was observed across sentences (99% CI [8.89%, 28.9%]), with a median decoding rate of 89.9 WPM (99% CI [83.6, 93.3]). With the AAC-50 set, a median WER of 4.92% (99% CI [3.18, 14.04]) was observed with median decoding speeds of 101 WPM (99% CI: [95.6, 103]). PERs and CERs for each set are given in FIGS. 17 and 18. The results show the models make more accurate predictions as the set’s sizes are reduced, and that the approach of the invention can directly decode faster, continuous speech in the context of closed sentence sets, which are highly valuable for common assistive-communication needs. FIGS. 3A to 3E and FIGS. 4A to 4B depict high-performance text decoding from neural activity. FIG. 3A is a schematic diagram of the text-decoding algorithm. During attempts to Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 silently speak, a bidirectional recurrent neural network (RNN) decodes neural features into a time series of phone and silence probabilities. Only the first 5 target phones and the silence token (denoted as Ø) are depicted. A CTC beam search processes this time series, searching for the most likely sequence of phones that can be translated into valid words in the vocabulary. An n- gram language model rescores sentences created from these phone sequences to yield the most likely sentence. FIG. 3B depicts median phone error rates with the 1024-word-General sentence set, calculated using shuffled neural data (Chance), neural decoding with the RNN without applying vocabulary constraints or language modeling (Neural decoding only), and the full real- time system (Real-time results) across 25 pseudo-blocks each containing 10 sentence trials. FIG. 3C depicts word error rates for chance and real-time results. FIG. 3D depicts character error rates for chance and real-time results. FIGS. 3B to 3D used a ****P < 0.0001, Two-sided Wilcoxon Signed-Rank test with 5-way Holm-Bonferroni correction for multiple comparisons; P-values and statistics are found in Table 2. FIG. 3E depicts communication rates, measured by decoded words per minute. Dashed line denotes previous state-of-the art rate for a speech brain-computer interface (BCI) in a person with paralysis [2]. FIG. 4A depicts offline evaluation of error rates as a function of number of recording days and data quantity. FIG. 4B depicts real-time classification accuracy during attempts to silently say 26 NATO code words across many recording days. The classifier was retrained after every session until it was frozen and no longer updated (vertical line). All box plots in all figures depict mean (horizontal line inside box), 25th and 75th percentiles (box), 25th and 75th percentiles + / - 1.5 times the interquartile range (whiskers), and outliers (diamonds). FIG. 16 provides results of simulated text decoding with a larger vocabulary. Text decoding results were simulated using a 42,391 word vocabulary on the blocks used for real-time evaluation with the 1024-word-General set. Across pseudo-blocks, a median WER of 29.82% (99% CI [21.1, 36.0], median CER of 22.17% (99% CI [17.5, 28.4]), and median PER of 21.4% (99% CI [16.8%, 26.6%]) was achieved. These results demonstrate consistent performance, even with a vocabulary 40x larger than the one used in real-time, demonstrating the natural ability of the phone decoding approach to scale up. The PER, WER, and CER were also significantly better than chance (P < .0001 for all metrics, Wilcoxon signed-rank test with 3-way Holm- Bonferonni Correction for multiple comparisons, n = 25 pseudo-blocks). For PER: stat = 0, P = 1.79e-8. For CER: stat = 0, P=1.79e-8. For WER: stat = 0, P=1.79e-8. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 FIG. 17 provides results of simulated text decoding on the 50-phrase-AAC sentence set. Text decoding results were simulated on the real-time blocks used for evaluation with the synthesis models. On the 50-phrase-AAC sentence set, extremely high accuracy was achieved. A median PER of 5.63% (99% CI [2.10, 12.0]) was observed. Median WER was 4.92% (99% CI [3.18, 14.0]), and median CER was 5.91% (99% CI [2.21, 11.4]). Speech was decoded at high rates with a median WPM of 101 (99% CI [95.6, 103]). The PER, WER, and CER were also significantly better than chance (P < .001 for all metrics, Wilcoxon signed-rank test with 3-way Holm-Bonferonni Correction for multiple comparisons). Statistics compare n = 15 total pseudo- blocks. For PER: stat=0, P = 1.83e-4. For CER: stat = 0, P=1.83e-4. For WER: stat = 0, P=1.83e- 4. FIG. 18 provides results of simulated text decoding on the 529-phrase-AAC sentence set. Text decoding results were simulated on the real-time blocks used for evaluation of the 529- phrase-AAC sentence set with the synthesis models. A median PER of 17.3 (99% CI [12.6, 20.1]) was observed. Median WER was 17.1% (99% CI [8.89, 28.9]), and median CER was 15.2% (99% CI [10.1, 22.7]). Speech was decoded at high rates with a median WPM of 89.9 (99% CI [83.6, 93.3]). The PER, WER, and CER were also significantly better than chance (P < .001 for all metrics, two-sided Wilcoxon signed-rank test with 3-way Holm-Bonferonni Correction for multiple comparisons). Statistics compare n = 15 total pseudo-blocks. For PER: stat=0, P = 1.83e-4. For CER: stat = 0, P=1.83e-4. For WER: stat = 0, P=1.83e-4. 1.3.2. Speech synthesis An alternative approach to text decoding is to synthesize speech sounds directly from recorded neural activity, which could offer a pathway towards more naturalistic and expressive communication for someone who is unable to speak. It has not previously been shown that intelligible speech can be synthesized from neural activity in someone who is paralyzed. To assess this, real-time speech synthesis was performed by transforming the participant’s neural activity directly into audible speech as she attempted to silently speak during the audio-visual task condition (FIG. 5, Videos 3 and 4). To synthesize speech, time windows of neural activity were passed around the go cue into a bidirectional RNN. Prior to testing, the RNN was trained to predict the probabilities of 100 discrete speech units at each time step. To create the reference speech-unit sequences for training, HuBERT, a self-supervised speech- Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 representation learning model, was used
[0013] . HuBERT was chosen for its ability to encode a continuous speech waveform into a temporal sequence of discrete speech units that captures latent phonetic and articulatory representations
[0014] . Because the participant cannot speak, reference speech waveforms were acquired from a recruited speaker for the AAC sentence sets or via a text-to-speech algorithm for the 1024-word-General set. A CTC loss function was used during training to enable the RNN to learn mappings between the ECoG features and speech units derived from these reference waveforms without having precise alignment between the participant’s silent speech attempts and the reference waveforms. Note the participant never heard the basis or reference waveforms and was not instructed to alter the style or pacing of her attempts to match the basis or reference waveforms in any way. After predicting the unit probabilities, the most likely unit at each time step was passed into a pre-trained unit-to-speech model that first generated a mel spectrogram and vocoded this mel spectrogram into an audible speech waveform in real time
[0015]
[0016] . Offline, a voice-conversion model trained on a brief segment of the participant’s speech (recorded before her injury) was used to process the decoded speech into the participant’s own personalized synthetic voice (Video 5). From speech-synthesis outputs during real-time testing, it was qualitatively observed that decoded spectrograms shared both fine-grained and broad time-scale information with corresponding reference spectrograms (FIG. 6). To quantitatively assess the quality of the decoded speech, the mel-cepstral distortion (MCD) metric was used, which measures the similarity between two sets of mel-cepstral coefficients (which are speech-relevant acoustic features) and is commonly used to evaluate speech-synthesis performance
[0017]
[0018] . Lower MCD indicates stronger similarity. Mean MCDs of 3.45 (99% CI [3.25, 3.82]), 4.49 (99% CI [4.07, 4.67]), and 5.21 (99% CI [4.74, 5.51]) dB were achieved for the 50-phrase-AAC, 529- phrase-AAC, and 1024-word-General sets, respectively (FIG. 7A). Similar MCD performance was observed on the participant’s personalized voice (FIG. 19; Table 3). Performance increased as the number of unique words and sentences in the sentence set decreased but was always better than chance (all P < 0.0001, two-sided Wilcoxon rank-sum tests with 19-way Holm-Bonferroni correction; chance MCDs were measured using waveforms generated by passing temporally shuffled ECoG features through the synthesis pipeline). Further, these MCDs are comparable to those observed with text-to-speech synthesizers
[0018] and better than what has been reported in prior neural-decoding work with participants that were able to speak naturally
[0012] . Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 Human-transcription assessments are a standard method to quantify the perceptual accuracy of synthesized speech
[0019] . To directly assess the intelligibility of the synthesized speech waveforms, perceptual free-transcription accuracy was measured from evaluators on a crowd-sourcing platform (Amazon Mechanical Turk). In this assessment, evaluators listened to the synthesized speech waveforms and then transcribed what they heard into text. Perceptual word and character error rates (WERs and CERs, respectively) were then computed by comparing these transcriptions to the ground-truth sentence texts. For each test trial, 11 evaluators transcribed the same synthesized speech waveform, and then the median error rate was used across evaluators to obtain a single WER and CER value for that trial. Median perceptual WERs of 8.33% (99% CI [5.17, 14.0]), 29.2% (99% CI [22.5, 33.3]), 61.9% (99% CI [46.4, 65.4]) and median perceptual CERs of 6.30% (99% CI [3.69, 10.1]), 24.6% (99% CI [16.5, 29.3]), and 46.7% (99% CI [36.6, 51.7]) were achieved across test trials for the 50-phrase- AAC, 529-phrase-AAC, and 1024-word-General sets, respectively (FIGS. 7B and 7C). Similar to the MCD results, perceptual WERs and CERs improved as the number of unique words and sentences in the sentence set decreased (all P < 0.0001, 2-sided Wilcoxon rank-sum tests with 19-way Holm-Bonferroni correction; chance perceptual WERs and CERs were measured by shuffling the mapping between the transcriptions and the ground-truth sentence texts). For the AAC sentence sets, evaluators often transcribed the decoded speech with perfect or near-perfect accuracy (119 / 150 trials and 79 / 150 trials had a perceptual WER of less than 15% for the 50- phrase-AAC and 500-phrase-AAC sets respectively). Together, these results demonstrate it is possible to synthesize intelligible speech from brain activity in a person with paralysis. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 Table 4: FIG. 5, FIG. 6, and FIGS. 7A to 7C depict intelligible speech synthesis from neural activity. FIG. 5 is a schematic diagram of the speech-synthesis decoding algorithm. During attempts to silently speak, a bidirectional recurrent neural network (RNN) decodes neural features into a time series of discrete speech units. The RNN was trained using reference speech units computed by applying a large pretrained acoustic model (HuBERT) on basis waveforms. Predicted speech units are then transformed into the mel spectrogram and vocoded into audible speech. The decoded waveform is played back to the participant in real time after a brief delay. Offline, the decoded speech was transformed to be in the participant’s personalized synthetic voice using a voice-conversion model. FIG. 5 top: three example decoded spectrograms and waveforms (top) from the 529-phrase-AAC sentence set. Bottom: the corresponding reference spectrograms and waveforms representing the decoding targets. FIG. 7A depicts mel-cepstral distortions (MCDs) for the decoded waveforms. Lower MCD indicates better performance. Chance waveforms were computed by shuffling electrode indices in the test data for the 50- phrase-AAC set with the same synthesis pipeline. FIG. 7B shows perceptual word error rates from untrained human evaluators via a transcription task. FIG. 7C shows perceptual character error rates from the same human-evaluation results as FIG. 7B. In FIGS 7A to 7C, ****P < 0.0001, Mann-Whitney U-test with 19-way Holm-Bonferroni correction for multiple comparisons; all non-adjacent comparisons were also significant (P < 0.0001; not depicted); Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 n=15 pseudo-blocks for the AAC sets, n=20 pseudo-blocks for 1024-word-General set. P-values and statistics in Table 4. In FIG. 6 and FIGS. 7A to 7C, all decoded waveforms, spectrograms, and quantitative results use the non-personalized voice (see FIG. 19 and Table 3 for results with the personalized voice). FIG. 19 provides Mel-cepstral distortions (MCDs) using a personalized voice tailored to the participant. The Mel-cepstral distortion (MCDs) was calculated between decoded speech with the participant’s personalized voice and reference waveforms for the 529-phrase-AAC, 50- phrase-AAC, and 1024-word-General set. Lower MCD indicates better performance. Mean MCDs of 3.87 (99% CI [3.83, 4.45]), 5.12 (99% CI [4.41, 5.35]), and 5.57 (99% CI [5.17, 5.90]) dB were achieved for the 50-phrase-AAC (N = 15 pseudo-blocks), 529-phrase-AAC (N = 15 pseudo-blocks), and 1024- word-General sets (N = 20 pseudo-blocks). Chance MCDs were computed by shuffling electrode indices in the test data with the same synthesis pipeline. The MCDs of all sets are significantly lower than the chance. 529-phrase-AAC vs. 1024-word- General ∗ ∗ ∗ = P ≤ 0.001, otherwise all ∗ ∗ ∗∗ = P ≤ 0.0001. Two-sided Wilcoxon rank-sum tests for comparisons within-dataset and Mann-Whitney U-test outside of dataset with 9-way Holm-Bonferroni correct. 1.3.3. Animating articulation for a personalized BCI avatar Face-to-face audio-visual communication offers multiple advantages over solely audio- based communication. Previous studies show that non-verbal facial gestures often account for a significant portion of the perceived feeling and attitude of a speaker
[0020]
[0021] and that face-to- face communication enhances social connectivity
[0022] and intelligibility
[0023] . Therefore, animation of a facial avatar to accompany synthesized speech and further embody the user is a promising means toward naturalistic communication. To this end, a facial-avatar BCI was developed to decode neural activity into articulatory speech gestures and render a dynamically moving virtual face during the audio-visual task condition (FIG. 8). To synthesize the avatar's motion an avatar-animation system was used that was designed to transform speech signals into accompanying facial-movement animations for applications in games, TV, film and communication (Speech Graphics Ltd, Edinburgh, Scotland). This technology uses speech-to-gesture methods that predict articulatory gestures (Table 5) from sound waveforms then synthesizes the avatar animation from these gestures
[0028] . A 3-D virtual Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 environment was designed to display the avatar to the participant during testing. Before testing, the participant selected an avatar from multiple potential candidates. Two approaches were implemented for animating the avatar: a direct approach and an acoustic approach. The direct approach was used for offline analyses to evaluate if articulatory movements could be directly inferred from neural activity without the use of a speech-based intermediate, which has implications for potential future uses of an avatar that are not based on speech representations, including non-verbal facial expressions. The acoustic approach was used for real-time audio-visual synthesis because it provided low-latency synchronization between decoded speech audio and avatar movements. For the direct approach, a bidirectional RNN was trained with CTC loss to learn a mapping between ECoG features and reference discretized articulatory gestures. These articulatory gestures were obtained by passing the reference waveforms through the animation system’s speech-to-gesture model. The articulatory gestures were then discretized using a vector quantized variational autoencoder (VQ-VAE)
[0029] . During testing, the RNN was used to decode the discretized articulatory gestures from neural activity and then de-quantized them into continuous articulatory gestures using the VQ-VAE’s decoder. Finally, the gesture-to-animation subsystem was used to animate the avatar face from the continuous gestures. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 It was found that the direct approach produced articulatory gestures that were strongly correlated with reference articulatory gestures across all datasets (FIG. 20 and FIG. 21; Table 6), highlighting the system’s ability to decode articulatory information from brain activity. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 Table 6: Comparisons for articulatory gesture decoding using the direct-decoding approach. Direct-decoding results were then evaluated by measuring the perceptual accuracy of the avatar. Evaluators on a crowd-sourcing platform (Amazon Mechanical Turk) watched silent videos of the decoded avatar animations and, for each video, were asked to identify to which of two sentences the video corresponded. One sentence was the ground-truth sentence while the other was randomly selected from the set of test sentences. The median bootstrapped accuracy was used across six evaluators to represent the final accuracy for each sentence. Median accuracies of 85.7% (99% CI [79.0, 92.0]), 87.7% (99% CI [79.7, 93.7]), and 74.3% (99% CI [66.7, 80.8]) were obtained across the 50-phrase-AAC, 529-phrase-AAC, and 1024-word- General sets, demonstrating that the avatar renderings conveyed perceptually meaningful speech- related facial movements (FIG. 9B). Next, the facial-avatar movements generated during direct decoding were compared with real movements made by healthy speakers. Videos of eight healthy volunteers were recorded as they read aloud sentences from the 1024-word-General set. A facial-keypoint recognition model was then applied (dlib)
[0030] to avatar and healthy-speaker videos to extract trajectories important for speech; jaw opening, lip aperture, and mouth width. For each pseudo-block of 10 test sentences and each trajectory, the mean correlations were computed across sentences between the trajectory values for each possible pair of corresponding videos (36 total combinations with one avatar and eight healthy-speaker videos). Prior to calculating correlations between two trajectories for the same sentence, dynamic time warping (DTW) was applied to account for variability in timing. It was found that the jaw opening, lip aperture, and mouth width of the avatar and healthy speakers were well correlated with median values of 0.733 (99% CI [0.711, 0.748]), 0.690 (99% CI [0.663, 0.714]), and 0.446 (99% CI [0.417, 0.470]) respectively (FIG. 9B). Although correlations amongst pairs of healthy speakers were significantly higher than between the avatar and healthy speakers (all P < 0.0001, two-sided Mann-Whitney U-test with 9- way Holm-Bonferroni correction; Table 7) there was a large degree of overlap between the two distributions, illustrating that the avatar reasonably approximated the expected articulatory Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 trajectories relative to natural variances between healthy speakers. Correlations for both distributions were significantly above chance, which was calculated by temporally shuffling the human trajectories and then recomputing correlations with DTW (all P < 0.0001, two-sided Mann-Whitney U-test with 9-way Holm-Bonferroni correction; Table 7). Avatar animations rendered in real time using the acoustic approach also exhibited strong correlations with reference articulatory gestures (FIG. 22; Table 8), high perceptual accuracy (FIG. 23), and visual facial-landmark trajectories that were closely correlated with healthy- speaker trajectories (FIG 24; Table 9). These findings emphasize the strong performance of the speech-synthesis neural decoder when used with the commercial speech-to-gesture rendering system, although this approach cannot be used to generate meaningful facial gestures in the absence of a speech waveform.
[0003] Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 Table 9: Comparisons for dlib traces with the acoustic approach.In addition to articulatory gestures to visually accompany synthesized speech, a fully embodying avatar BCI would also enable the user to portray non-speech orofacial gestures, including movements of particular orofacial muscles and expressions that convey emotion. To this end, neural data was collected from the participant as she performed two additional tasks: an articulatory-movement task and an emotional-expression task. In the articulatory-movement task, the participant attempted to produce 6 orofacial movements: jaw opening, lip puckering, lip retraction (smiling), tongue raising, tongue lowering, and rest (idle with mouth closed). In the emotional-expression task, the participant attempted to produce 3 types of expressions — happy, sad, and surprised — with either low, medium, or high intensity, resulting in 9 unique expressions in total. Offline, for the articulatory-movement task a small feed-forward neural- network model was trained to learn the mapping between the ECoG features and each of the targets. For the articulatory-movement task, a median classification accuracy of 87.8% (99% CI [85.1, 90.5]; across n=10 cross-validation folds; FIG. 10A) was observed when classifying between the 6 movements. For the emotional-expression task, a small RNN to was trained learn the mapping between ECoG features and each of the expression targets. A median classification accuracy of 74.0% (99% CI [70.8, 77.1]; across n=15 cross-validation folds; FIG. 10B) was observed when classifying between the 9 possible expressions and a median classification accuracy of 96.9% (99% CI [93.8,100]) when only considering the classifier’s outputs for the Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 strong-intensity versions of the 3 expression types (FIG. 25). In separate, qualitative task blocks, it was shown that the participant could control the avatar BCI to portray the articulatory movements (Video 6) and strong-intensity emotional expressions (Video 7), illustrating the potential of multimodal communication BCIs to restore the ability to express meaningful orofacial gestures. FIG. 8, FIGS. 9A to 9B, and FIGS. 10A to 10B depict direct decoding of orofacial articulatory gestures from neural activity to drive an avatar. FIG. 8 is a schematic diagram of the avatar decoding algorithm. Offline, a bidirectional recurrent neural network (RNN) decodes neural activity recorded during attempts to silently speak into discretized articulatory gestures (quantized via a vector quantized variational autoencoder, abbreviated VQ-VAE). A convolutional neural network de-quantizer (VQ-VAE decoder) is then applied to generate the final predicted gestures, which are then passed through a pre-trained gesture-animation model to animate the avatar in a virtual environment. FIG. 9A binary perceptual accuracies from human evaluators on avatar animations generated from neural activity. Evaluators see the decoded avatar animations (with no accompanying audio) and, for each one, choose between the correct reference text target and a randomly selected incorrect text string. FIG. 9B correlations for jaw, lip, and mouth-width movements between decoded avatar renderings and videos of real human speakers on the 1024-word-General sentence set across all pseudo-blocks for each comparison (n=152 for avatar-person comparison, n=532 for person-person comparisons; ****P < 0.0001, Mann-Whitney U-test with 9-way Holm-Bonferroni correction; p-values and U-statistics in Table 7). A facial-landmark detector (dlib) was used to measure orofacial movements from the videos. FIG. 10A top: snapshots of avatar animations of 6 non-speech articulatory movements in the articulatory-movement task. Bottom: confusion matrix depicting classification accuracy across the movements. The classifier was trained to predict which movement the participant was attempting from her neural activity, and the prediction was used to animate the avatar. FIG. 10B top: snapshots of avatar animations of 3 non-speech emotional expressions in the emotional- expression task. Bottom: confusion matrix depicting classification accuracy across 3 intensity levels (high, medium, and low) of the 3 expressions, ordered via hierarchical agglomerative clustering on the confusion values. The classifier was trained to predict which expression the participant was attempting from her neural activity, and the prediction was used to animate the avatar. Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 FIG. 20 provides examples of directly decoded avatar articulatory gestures. Examples of directly decoded articulatory gestures (colored) compared with reference articulatory gestures (black). Examples were taken from the 50-phrase-AAC sentence set. Dynamic time warping
[0086] was applied to align traces prior to plotting and computation of Pearson r correlation, which is displayed to the right of each gesture. Reference articulatory gestures were computed using the speech-to-gesture technology from SG Com. FIG. 21 provides correlations of directly decoded avatar articulatory gestures with reference articulatory gestures. Pearson correlation (R) of decoded articulatory gestures with reference articulatory gestures using the direct decoding approach after applying dynamic time warping using fast-dtw
[0086] to align the reference and decoded gestures, since the participant never heard the reference waveform used to derive reference gestures. Chance values are derived by shuffling the neural data temporally then feeding it through our decoding pipeline. The resulting traces are then warped using fast-dtw and compared with reference traces. Correlations were significantly above chance for all comparisons except comparisons of nostril flare for all sentence sets and pinching for the 1024-word-General sentence set, two-sided Wilcoxon Signed Rank test with 16-way HolmBonferroni correction across n=20 pseudo-blocks for the 1024- word-General sentence set, n = 15 pseudo-blocks for AAC sets. See Supplementary Table 6 for all p-values and statistics. **** P< .0001, *** P < .001, ** P<.01. FIG. 22 provides correlations between avatar articulatory gestures with reference articulatory gestures using the acoustic approach. Pearson correlation (R) of decoded articulatory gestures with reference articulatory gestures during acoustic approach after applying dynamic time warping using fast-dtw
[0086] to align the reference and decoded gestures, since the participant never heard the reference waveform used to derive reference gestures. Chance values are derived by shuffling the neural data temporally then feeding it through our decoding pipeline. The resulting traces are then warped using fast-dtw and compared with reference traces. Co...
Claims
Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 WHAT IS CLAIMED IS:
1. A method of assisting a subject with communication, the method comprising: positioning a neural recording device comprising an electrode at a location in a sensorimotor cortex region of the brain of the subject to record brain electrical signal data associated with attempted speech by the subject; positioning an interface in communication with a computing device at a location on the head of the subject, wherein the interface is connected to the neural recording device; recording the brain electrical signal data associated with attempted speech by the subject using the neural recording device, wherein the interface receives the brain electrical signal data from the neural recording device and transmits the brain electrical signal data to a processor of the computing device; and decoding one or more speech sounds from the recorded brain electrical signal data using the processor, wherein the processor is programmed to use a machine learning model for the decoding.
2. The method of claim 1, wherein the subject has difficulty with said communication because of anarthria, a stroke, a traumatic brain injury, a brain tumor, or amyotrophic lateral sclerosis.
3. The method of claim 1 or 2, wherein the subject is paralyzed.
4. The method of any of claims 1-3, wherein the subject has a speech intelligibility of 10% or less for prompted words.
5. The method of any of claims 1-4, wherein the location of the neural recording device is on a surface of the sensorimotor cortex region or within the sensorimotor cortex region.
6. The method of any of claims 1-5, wherein the electrode of the neural recording device is positioned on the pial surface of the sensorimotor cortex.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 7. The method of any of claims 1-6, wherein the neural recording device is positioned such that the recording device covers regions associated with speech production and language perception.
8. The method of claim 7, wherein the covered regions include a middle portion of the superior and / or middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus.
9. The method of any of claims 1-8, wherein the neural recording device is centered on the central sulcus.
10. The method of any of claims 1-9, wherein the neural recording device comprises an electrocorticography (ECoG) electrode array.
11. The method of claim 10, wherein the ECoG electrode array is a high-density array.
12. The method of claim 11, wherein the high-density array comprises 200 electrodes or more.
13. The method of any of claims 10-12, wherein the electrode array comprises non- penetrating surface electrodes.
14. The method of any of claims 1-13, wherein the interface comprises a percutaneous pedestal connector attached to the subject's cranium.
15. The method of any of claims 1-14, wherein the interface further comprises a headstage connected to the percutaneous pedestal connector.
16. The method of claim 15, wherein the headstage processes and digitizes the brain electrical signal data before transmitting the data to the processor.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 17. The method of any of claims 1-16, wherein the electrical signal data comprises high- gamma frequency content features.
18. The method of any of claims 1-17, wherein the electrical signal data comprises neural oscillations in a range from 70 Hz to 150 Hz.
19. The method of any of claims 1-18, wherein the electrical signal data comprises low- frequency signals.
20. The method of claim 19, wherein the electrical signal data comprises neural oscillations in a range from 0.3 Hz to 17 Hz.
21. The method of any of claims 1-20, wherein the brain electrical signal data is recorded from a sensorimotor cortex region selected from the precentral gyrus, postcentral gyrus, superior temporal gyrus, middle temporal gyrus, or any combination thereof.
22. The method of any of claims 1-21, wherein the method further comprises mapping the brain of the subject to identify an optimal location for positioning the electrode for recording the brain electrical signals associated with the attempted speech by the subject.
23. The method of any of claims 1-22, wherein the one or more speech sounds form a word.
24. The method of claim 23, wherein the one or more speech sounds form a sentence.
25. The method of claims 23 or 24, wherein the subject is limited to a specified word set for the attempted speech.
26. The method of claim 25, wherein the word set comprises words for expressing basic concepts and / or caregiving needs.
27. The method of claims 25 or 26, wherein the word set comprises 100 words or more.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 28. The method of claim 27, wherein the word set comprises 350 words or more.
29. The method of claim 28, wherein the word set comprises 1000 words or more.
30. The method of any of claims 25-29, wherein the subject is limited to two or more word sets for the attempted speech.
31. The method of claim 30, wherein the subject may switch between word sets.
32. The method of any of claims 1-31, wherein the decoding machine learning model comprises a neural network.
33. The method of claim 32, wherein the neural network comprises one or more convolutional layers.
34. The method of claims 32 or 33, wherein the neural network is bidirectional.
35. The method of claim 34, wherein the neural network is a recurrent neural network (RNN).
36. The method of claim 35, wherein the RNN comprises gated recurrent units (GRUs) or long-short term memory (LSTM).
37. The method of claim 36, wherein the RNN comprises GRUs.
38. The method of any of claims 1-37, wherein one or more discrete speech units of a set of discrete speech units are decoded from the recorded brain electrical signal data.
39. The method of claim 38, wherein each speech sound is decoded from one or more discrete speech units.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 40. The method of claims 38 or 39, wherein the set of discrete speech units comprises 50 or more discrete speech units.
41. The method of claim 40, wherein the set of discrete speech units comprises 100 or more discrete speech units.
42. The method of any of claims 36-39, wherein discrete speech units are continuously decoded at a uniform frequency.
43. The method of claim 42, wherein the frequency is 50 Hz or more.
44. The method of claim 43, wherein the frequency is 200 Hz or more.
45. The method of any of claims 38-44, wherein the set of discrete speech units are generated using an encoding machine learning model.
46. The method of claim 45, wherein the encoding machine learning model comprises a neural network.
47. The method of claim 46, wherein the neural network comprises one or more convolutional layers.
48. The method of claims 46 or 47, wherein the neural network is bidirectional.
49. The method of claim 48, wherein the neural network comprises a transformer encoder.
50. The method of claim 49, wherein the machine learning model is a Hidden-Unit Bidirectional Encoder Representations from Transformers (HuBERT) model.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 51. The method of any of claims 45-50, wherein the set of discrete speech units are generated by training the encoding machine learning model.
52. The method of claim 51, wherein the training is self-supervised training.
53. The method of any of claims 45-52, wherein the method further comprises: obtaining reference electronic speech waveforms for a plurality of phrases; encoding each reference speech waveform into a temporal sequence of discrete speech units using the encoding model; recording the brain electrical signal data associated with attempted speech by the subject for each one of the plurality of phrases; training the decoding machine learning model to predict the most likely discrete speech unit associated with a segment of electrical signal data using the speech units derived from the reference speech waveforms and the brain electrical signal data associated with attempted speech for each one of the plurality of phrases.
54. The method of claim 53, wherein the reference speech waveforms are obtained from a recruited speaker.
55. The method of claim 53, wherein the reference speech waveforms are obtained using a text-to-speech algorithm.
56. The method of any of claims 53-55, wherein the decoding machine learning model is trained to learn mappings between neural activity patterns of electrical signals in the brain electrical signal data and the discrete speech units.
57. The method of any of claims 53-56, wherein the training uses a CTC loss function.
58. The method of any of claims 39-57, wherein each speech sound is decoded from one or more discrete speech units using a speech synthesizer.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 59. The method of claim 58, wherein the speech synthesizer comprises a machine learning model.
60. The method of claim 59, wherein the synthesizing machine learning model comprises a neural network.
61. The method of claim 60, wherein the neural network comprises one or more convolutional layers.
62. The method of claims 60 or 61, wherein the neural network is bidirectional.
63. The method of claim 62, wherein the neural network is an RNN.
64. The method of claim 63, wherein the RNN comprises GRUs or LSTM.
65. The method of claim 64, wherein the RNN comprises one or more LSTM layers.
66. The method of any of claims 60-65, wherein the neural network comprises an attention mechanism.
67. The method of any of claims 59-66, wherein the synthesizing machine learning model is configured to generate a spectrogram from the one or more discrete speech units.
68. The method of claim 67, wherein the spectrogram is a mel spectrogram.
69. The method of claims 67 or 68, wherein the speech synthesizer further comprises a vocoder configured to synthesize an electronic speech waveform from the spectrogram.
70. The method of claim 69, wherein the vocoder comprises a machine learning model.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 71. The method of claim 70, wherein the vocoder machine learning model comprises a neural network.
72. The method of claim 71, wherein the neural network is an RNN.
73. The method of any of claims 69-72, wherein the method further comprises transforming the electronic speech waveform into a personalized electronic speech waveform.
74. The method of claim 73, wherein the personalized electronic speech waveform resembles speech in the subject’s own voice.
75. The method of claims 73 or 74, wherein the transforming is performed using a machine learning model.
76. The method of claim 75, wherein the transforming machine learning model comprises a neural network.
77. The method of claim 76, wherein the neural network is based on transformer architecture.
78. The method of any of claims 69-77, wherein the method further comprises converting the electronic speech waveform into an audible speech waveform.
79. The method of claim 78, wherein the electronic speech waveform is converted into an audible speech waveform using a loudspeaker.
80. The method of any of the preceding claims, wherein the processor is provided by a computer or handheld device.
81. The method of claim 80, wherein the handheld device is a cell phone or a tablet.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 82. A non-transitory computer-readable medium comprising program instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1-81.
83. A kit comprising the non-transitory computer-readable medium of claim 82 and instructions for decoding brain electrical signal data associated with attempted speech by a subject.
84. A system for decoding speech sounds from recorded brain electrical signal data configured to perform the method according to any of Claims 1-81.
85. A computer implemented method for decoding speech audio from recorded brain electrical signal data associated with attempted speech by a subject, the computer performing steps comprising: receiving the recorded brain electrical signal data associated with the attempted speech by the subject; and decoding one or more speech sounds from the recorded brain electrical signal data using a machine learning model.
86. The computer implemented method of claim 85, wherein the decoding machine learning model comprises a neural network.
87. The computer implemented method of claim 86, wherein the neural network comprises one or more convolutional layers.
88. The computer implemented method of claims 86 or 87, wherein the neural network is bidirectional.
89. The computer implemented method of claim 88, wherein the neural network is a recurrent neural network (RNN).Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 90. The computer implemented method of claim 89, wherein the RNN comprises gated recurrent units (GRUs) or long-short term memory (LSTM).
91. The computer implemented method of claim 90, wherein the RNN comprises GRUs.
92. The computer implemented method of any of claims 85-91, wherein one or more discrete speech units of a set of discrete speech units are decoded from the recorded brain electrical signal data.
93. The computer implemented method of claim 92, wherein each speech sound is decoded from one or more discrete speech units.
94. The computer implemented method of claims 92 or 93, wherein the set of discrete speech units comprises 50 or more discrete speech units.
95. The computer implemented method of claim 94, wherein the set of discrete speech units comprises 100 or more discrete speech units.
96. The computer implemented method of any of claims 92-95, wherein discrete speech units are continuously decoded at a uniform frequency.
97. The computer implemented method of claim 96, wherein the frequency is 50 Hz or more.
98. The computer implemented method of claim 97, wherein the frequency is 200 Hz or more.
99. The computer implemented method of any of claims 92-98, wherein the set of discrete speech units are generated using an encoding machine learning model.
100. The computer implemented method of claim 99, wherein the encoding machine learning model comprises a neural network.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 101. The computer implemented method of claim 100, wherein the neural network comprises one or more convolutional layers.
102. The computer implemented method of claims 100 or 101, wherein the neural network is bidirectional.
103. The computer implemented method of claim 102, wherein the neural network comprises a transformer encoder.
104. The computer implemented method of claim 103, wherein the machine learning model is a Hidden-Unit Bidirectional Encoder Representations from Transformers (HuBERT) model.
105. The computer implemented method of any of claims 99-104, wherein the set of discrete speech units are generated by training the encoding machine learning model.
106. The computer implemented method of claim 105, wherein the training is self-supervised training.
107. The computer implemented method of any of claims 99-106, wherein the method further comprises: receiving reference electronic speech waveforms for a plurality of phrases; encoding each reference speech waveform into a temporal sequence of discrete speech units using the encoding model; receiving brain electrical signal data associated with attempted speech by the subject for each one of the plurality of phrases; training the decoding machine learning model to predict the most likely discrete speech unit associated with a segment of electrical signal data using the speech units derived from the reference speech waveforms and the brain electrical signal data associated with attempted speech for each one of the plurality of phrases.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 108. The computer implemented method of claim 107, wherein the reference speech waveforms are generated using a text-to-speech algorithm.
109. The computer implemented method of claims 107 or 108, wherein the decoding machine learning model is trained to learn mappings between neural activity patterns of electrical signals in the brain electrical signal data and the discrete speech units.
110. The computer implemented method of any of claims 93-109, wherein each speech sound is decoded from one or more discrete speech units using a speech synthesizer.
111. The computer implemented method of claim 110, wherein the speech synthesizer comprises a machine learning model.
112. The computer implemented method of claim 111, wherein the synthesizing machine learning model comprises a neural network.
113. The computer implemented method of claim 112, wherein the neural network comprises one or more convolutional layers.
114. The computer implemented method of claims 112 or 113, wherein the neural network is bidirectional.
115. The computer implemented method of claim 114, wherein the neural network is an RNN.
116. The computer implemented method of claim 115, wherein the RNN comprises GRUs or LSTM.
117. The computer implemented method of claim 116, wherein the RNN comprises one or more LSTM layers.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 118. The computer implemented method of any of claims 112-117, wherein the neural network comprises an attention mechanism.
119. The computer implemented method of any of claims 111-118, wherein the synthesizing machine learning model is configured to generate a spectrogram from the one or more discrete speech units.
120. The computer implemented method of claim 119, wherein the spectrogram is a mel spectrogram.
121. The computer implemented method of claims 119 or 120, wherein the speech synthesizer further comprises a vocoder configured to synthesize an electronic speech waveform from the spectrogram.
122. The computer implemented method of claim 121, wherein the vocoder comprises a machine learning model.
123. The computer implemented method of claim 122, wherein the vocoder machine learning model comprises a neural network.
124. The computer implemented method of claim 123, wherein the neural network is an RNN.
125. The computer implemented method of any of claims 121-124, wherein the method further comprises transforming the electronic speech waveform into a personalized electronic speech waveform.
126. The computer implemented method of claim 125, wherein the personalized electronic speech waveform resembles speech in the subject’s own voice.
127. The computer implemented method of claims 125 or 126, wherein the transforming is performed using a machine learning model.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 128. The computer implemented method of claim 127, wherein the transforming machine learning model comprises a neural network.
129. The computer implemented method of claim 128, wherein the neural network is based on transformer architecture.
131. A non-transitory computer-readable medium comprising program instructions that, when executed by a processor, cause the processor to perform the computer implemented method of any one of claims 85-129.
132. A kit comprising the non-transitory computer-readable medium of claim 131 and instructions for decoding brain electrical signal data associated with attempted speech by a subject.
133. A system for producing speech audio directly from neural activity, the system comprising: a neural recording device comprising an electrode adapted for positioning at a location in a sensorimotor cortex region of the brain of the subject to record brain electrical signal data associated with attempted speech by the subject; a processor programmed to use a machine learning model to decode one or more speech sounds from the recorded brain electrical signal data; an interface in communication with a computing device, said interface adapted for positioning at a location on the head of the subject, wherein the interface receives the brain electrical signal data from the neural recording device and transmits the brain electrical signal data to the processor; and an audio speaker for playing the one or more speech sounds from the recorded brain electrical signal data.
134. The system of claim 133, wherein the subject has difficulty with said speech because of anarthria, a stroke, a traumatic brain injury, a brain tumor, or amyotrophic lateral sclerosis.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 135. The system of any of claims 133-134, wherein the neural recording device is adapted for positioning on a surface of the sensorimotor cortex region or within the sensorimotor cortex region.
136. The system of any of claims 133-135, wherein the electrode of the neural recording device is adapted for positioning on the pial surface of the sensorimotor cortex.
137. The system of any of claims 133-136, wherein the neural recording device is adapted to be positioned such that the recording device covers regions associated with speech production and language perception.
138. The system of claim 137, wherein the covered regions include a middle portion of the superior and / or middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus.
139. The system of any of claims 133-138, wherein the neural recording device is adapted to be centered on the central sulcus.
140. The system of any of claims 133-139, wherein the neural recording device comprises an electrocorticography (ECoG) electrode array.
141. The system of claim 140, wherein the ECoG electrode array is a high-density array.
142. The system of claim 141, wherein the high-density array comprises 200 electrodes or more.
143. The system of any of claims 140-142, wherein the electrode array comprises non- penetrating surface electrodes.
144. The system of any of claims 133-143, wherein the interface comprises a percutaneous pedestal connector attached to the subject's cranium.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 145. The system of any of claims 133-144, wherein the interface further comprises a headstage connected to the percutaneous pedestal connector.
146. The system of claim 145, wherein the headstage is configured to process and digitize the brain electrical signal data before transmitting the data to the processor.
147. The system of any of claims 133-146, wherein the electrical signal data comprises high- gamma frequency content features.
148. The system of any of claims 133-147, wherein the electrical signal data comprises neural oscillations in a range from 70 Hz to 150 Hz.
149. The system of any of claims 133-148, wherein the electrical signal data comprises low- frequency signals.
150. The system of claim 149, wherein the electrical signal data comprises neural oscillations in a range from 0.3 Hz to 17 Hz.
151. The system of any of claims 133-150, wherein the brain electrical signal data is recorded from a sensorimotor cortex region selected from the precentral gyrus, postcentral gyrus, superior temporal gyrus, middle temporal gyrus, or any combination thereof.
153. The system of any of claims 133-152, wherein the one or more speech sounds form a word.
154. The system of claim 153, wherein the one or more speech sounds form a sentence.
155. The system of claims 153 or 154, wherein the subject is limited to a specified word set for the attempted speech.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 156. The system of claim 155, wherein the word set comprises words for expressing basic concepts and / or caregiving needs.
157. The system of claims 155 or 156, wherein the word set comprises 100 words or more.
158. The system of claim 157, wherein the word set comprises 350 words or more.
159. The system of claim 158, wherein the word set comprises 1000 words or more.
160. The system of any of claims 155-159, wherein the subject is limited to two or more word sets for the attempted speech.
161. The system of claim 160, wherein the processor is configured to switch between word sets.
162. The system of any of claims 133-161, wherein the decoding machine learning model comprises a neural network.
163. The system of claim 162, wherein the neural network comprises one or more convolutional layers.
164. The system of claims 162 or 163, wherein the neural network is bidirectional.
165. The system of claim 164, wherein the neural network is a recurrent neural network (RNN).
166. The system of claim 165, wherein the RNN comprises gated recurrent units (GRUs) or long-short term memory (LSTM).
167. The system of claim 166, wherein the RNN comprises GRUs.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 168. The system of any of claims 133-167, wherein one or more discrete speech units of a set of discrete speech units are decoded from the recorded brain electrical signal data.
169. The system of claim 168, wherein each speech sound is decoded from one or more discrete speech units.
170. The system of any of claims 168-169, wherein discrete speech units are continuously decoded at a uniform frequency.
171. The system of claim 170, wherein the frequency is 200 Hz or more.
172. The system of any of claims 168-171, wherein the set of discrete speech units are generated using an encoding machine learning model.
173. The system of claim 172, wherein the encoding machine learning model comprises a neural network.
174. The system of claim 173, wherein the neural network comprises one or more convolutional layers.
175. The system of claims 173 or 174, wherein the neural network is bidirectional.
176. The system of claim 175, wherein the neural network comprises a transformer encoder.
177. The system of claim 176, wherein the machine learning model is a Hidden-Unit Bidirectional Encoder Representations from Transformers (HuBERT) model.
178. The system of any of claims 172-177, wherein the set of discrete speech units are generated by training the encoding machine learning model.
179. The system of claim 178, wherein the training is self-supervised training.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 180. The system of any of claims 172-179, wherein the processor is further programmed to: obtain reference electronic speech waveforms for a plurality of phrases; encode each reference speech waveform into a temporal sequence of discrete speech units using the encoding model; record the brain electrical signal data associated with attempted speech by the subject for each one of the plurality of phrases; train the decoding machine learning model to predict the most likely discrete speech unit associated with a segment of electrical signal data using the speech units derived from the reference speech waveforms and the brain electrical signal data associated with attempted speech for each one of the plurality of phrases.
181. The system of claim 180, wherein the system further comprises a microphone configured to generate reference speech waveforms from a recruited speaker.
182. The system of claim 180, wherein the processor further comprises a text-to-speech algorithm configured to generate the reference speech waveforms.
183. The system of any of claims 180-182, wherein the decoding machine learning model is trained to learn mappings between neural activity patterns of electrical signals in the brain electrical signal data and the discrete speech units.
184. The system of any of claims 172 to 183, wherein each speech sound is decoded from one or more discrete speech units using a speech synthesizer.
185. The system of claim 184, wherein the speech synthesizer comprises a machine learning model.
186. The system of claim 185, wherein the synthesizing machine learning model comprises a neural network.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 187. The system of claim 186, wherein the neural network comprises one or more convolutional layers.
188. The system of claims 186 or 187, wherein the neural network is bidirectional.
189. The system of claim 188, wherein the neural network is an RNN.
190. The system of claim 189, wherein the RNN comprises GRUs or LSTM.
191. The system of claim 190, wherein the RNN comprises one or more LSTM layers.
192. The system of any of claims 186 to 191, wherein the neural network comprises an attention mechanism.
193. The system of any of claims 185 to 192, wherein the synthesizing machine learning model is configured to generate a spectrogram from the one or more discrete speech units.
194. The system of claim 193, wherein the spectrogram is a mel spectrogram.
195. The system of claims 193 or 194, wherein the speech synthesizer further comprises a vocoder configured to synthesize an electronic speech waveform from the spectrogram.
196. The system of claim 195, wherein the vocoder comprises a machine learning model.
197. The system of claim 196, wherein the vocoder machine learning model comprises a neural network.
198. The system of claim 197, wherein the neural network is an RNN.
199. The system of any of claims 195 to 198, wherein the processor is further programmed to transform the electronic speech waveform into a personalized electronic speech waveform.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 200. The system of claim 199, wherein the personalized electronic speech waveform resembles speech in the subject’s own voice.
201. The system of claims 199 or 200, wherein the transforming is performed using a machine learning model.
202. The system of claim 201, wherein the transforming machine learning model comprises a neural network.
203. The system of claim 202, wherein the neural network is based on transformer architecture.
204. The system of any of claims 195-203, wherein the processor is further programmed to convert the electronic speech waveform into an audible speech waveform.
205. The system of claim 204, wherein the electronic speech waveform is converted into an audible speech waveform using a loudspeaker.
206. The system of any of claims 133-205, wherein the processor is provided by a computer or handheld device.
207. The system of claim 206, wherein the handheld device is a cell phone or a tablet.
208. A kit comprising the system of any of claims 133-207 and instructions for decoding brain electrical signal data associated with attempted speech by a subject.
209. A method of controlling an electronic output device or system to perform one or more actions using brain electrical signals, the method comprising:Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 positioning a neural recording device comprising an electrode at a location in a sensorimotor cortex region of the brain of the subject to record brain electrical signal data associated with an attempted action by the subject; positioning an interface in communication with a computing device at a location on the head of the subject, wherein the interface is connected to the neural recording device; recording the brain electrical signal data associated with the attempted action by the subject using the neural recording device, wherein the interface receives the brain electrical signal data from the neural recording device and transmits the brain electrical signal data to a processor of the computing device; decoding one or more electronic output device or system actions from the recorded brain electrical signal data, wherein the processor is programmed to use a machine learning model for the decoding; and controlling the electronic output device or system to perform the one or more decoded electronic output device or system actions.
210. The method of claim 209, wherein the subject has difficulty communicating because of anarthria, a stroke, a traumatic brain injury, a brain tumor, or amyotrophic lateral sclerosis.
211. The method of claim 209 or 210, wherein the subject is paralyzed.
212. The method of any of claims 209-211, wherein the subject is quadriplegic and / or experiences partial or total facial paralysis.
213. The method of any of claims 209-212, wherein the location of the neural recording device is on a surface of the sensorimotor cortex region or within the sensorimotor cortex region.
214. The method of any of claims 209-213, wherein the electrode of the neural recording device is positioned on the pial surface of the sensorimotor cortex.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 215. The method of any of claims 209-214, wherein the neural recording device is positioned such that the recording device covers regions associated with speech production and language perception.
216. The method of claim 215, wherein the covered regions include a middle portion of the superior and / or middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus.
217. The method of any of claims 209-216, wherein the neural recording device is centered on the central sulcus.
218. The method of any of claims 209-217, wherein the neural recording device comprises an electrocorticography (ECoG) electrode array.
219. The method of claim 218, wherein the ECoG electrode array is a high-density array.
220. The method of claim 219, wherein the high-density array comprises 200 electrodes or more.
221. The method of any of claims 218-220, wherein the electrode array comprises non- penetrating surface electrodes.
222. The method of any of claims 209-221, wherein the interface comprises a percutaneous pedestal connector attached to the subject's cranium.
223. The method of any of claims 209-222, wherein the interface further comprises a headstage connected to the percutaneous pedestal connector.
224. The method of claim 223, wherein the headstage processes and digitizes the brain electrical signal data before transmitting the data to the processor.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 225. The method of any of claims 209-224, wherein the electrical signal data comprises high- gamma frequency content features.
226. The method of any of claims 209-225, wherein the electrical signal data comprises neural oscillations in a range from 70 Hz to 150 Hz.
227. The method of any of claims 209-226, wherein the electrical signal data comprises low- frequency signals.
228. The method of claim 227, wherein the electrical signal data comprises neural oscillations in a range from 0.3 Hz to 17 Hz.
229. The method of any of claims 209-228, wherein the brain electrical signal data is recorded from a sensorimotor cortex region selected from the precentral gyrus, postcentral gyrus, superior temporal gyrus, middle temporal gyrus, or any combination thereof.
230. The method of any of claims 209-229, wherein the method further comprises mapping the brain of the subject to identify an optimal location for positioning the electrode for recording the brain electrical signals associated with the attempted speech by the subject.
231. The method of any of claims 209-230, wherein the subject is limited to a specified action set for the attempted action.
232. The method of claim 231, wherein the action set comprises actions for expressing basic emotions and / or communicating caregiving needs.
233. The method of claims 231 or 232, wherein the action set comprises 6 actions or more.
234. The method of claim 233, wherein the action set comprises 100 actions or more.
235. The method of claim 234, wherein the action set comprises 500 actions or more.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 236. The method of any of claims 231-235, wherein the subject is limited to two or more action sets for the attempted action.
237. The method of claim 28236 wherein the subject may switch between action sets.
238. The method of any of claims 209-237, wherein the decoding machine learning model comprises a neural network.
239. The method of claim 238, wherein the neural network comprises one or more convolutional layers.
240. The method of claims 238 or 239, wherein the neural network is bidirectional.
241. The method of claim 240, wherein the neural network is a recurrent neural network (RNN).
242. The method of claim 241, wherein the RNN comprises gated recurrent units (GRUs) or long-short term memory (LSTM).
243. The method of claim 242, wherein the RNN comprises GRUs.
244. The method of any of claims 209-243, wherein each action of the device or system is discretized into one or more action representations.
245. The method of claim 244, wherein the action representations are decoded from the recorded brain electrical signal data.
246. The method of claim 245, wherein each action of the device or system is decoded from one or more action representations.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 247. The method of claim 246, wherein discrete action representations are continuously decoded from the recorded brain electrical signal data at a uniform frequency.
248. The method of claim 247, wherein the frequency is 50 Hz or more.
249. The method of claim 248, wherein the frequency is 200 Hz or more.
250. The method of any of claims 244-249, wherein the discretization is performed using an autoencoder.
251. The method of claim 250, wherein the autoencoder comprises one or more convolutional layers.
252. The method of claims 250 or 251, wherein the autoencoder uses one or more rectified linear unit (ReLU) activations.
253. The method of claim 252, wherein the autoencoder comprises a vector-quantized variational autoencoder (VQ-VAE).
254. The method of any of claims 250-253, wherein the discretization occurs during training of the autoencoder.
255. The method of claim 254, wherein the training is self-supervised training.
256. The method of any of claims 250-255, wherein the attempted action performed by the subject is different than the one or more electronic output device or system actions.
257. The method of claim 256, wherein the attempted action performed by the subject is a hand gesture and the action of the electronic output device or system is powering on or off.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 258. The method of any of claims 250-255, wherein the attempted action performed by the subject corresponds to the one or more electronic output device or system actions.
259. The method of claim 258, wherein the electronic output device or system is a prosthetic limb.
260. The method of claim 258, wherein the electronic output device or system comprises a visual display and / or a loudspeaker.
261. The method of claim 260, wherein the visual display and / or loudspeaker is configured to present a humanoid avatar.
262. The method of claim 261, wherein the one or more electronic output device or system actions comprise actions performed by the avatar.
263. The method of claim 262, wherein actions performed by the avatar include speech, orofacial movements for speech, and / or non-speech communicative gestures.
264. The method of claim 263, wherein the speech orofacial movements comprise one or more of: a tongue tip raise, tongue retraction, tongue body raise, tongue advance, lip rounding, pinching nostril flare, upper lip pull, lower lip tuck, lower lip push, lower lip pull, lip flare, jaw opening, lip compression, and / or lip adduction.
265. The method of claim 263, wherein the non-speech communicative gestures comprise the abduction, adduction, flexion, extension, and / or circumduction of one or more body parts.
266. The method of claims 263 or 265, wherein the non-speech communicative gestures include emotional expressions using facial muscles.
267. The method of claim 266, wherein the emotional expressions comprise happy, sad, and surprised expressions.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 268. The method of any of claims 262-267, wherein the method further comprises: obtaining reference avatar animations for a plurality of actions; encoding each reference animation into a temporal sequence of discrete action representations using the autoencoder; recording the brain electrical signal data associated with attempted action by the subject for each one of the plurality of actions; training the decoding machine learning model to predict the most likely discrete action representation associated with a segment of electrical signal data using the action representations derived from the reference avatar animations and the brain electrical signal data associated with attempted action for each one of the plurality of actions.
269. The method of claim 268, wherein the reference avatar animations are obtained from an avatar-animation system.
270. The method of any of claims 268 or 268, wherein the decoding machine learning model is trained to learn mappings between neural activity patterns of electrical signals in the brain electrical signal data and the discrete action representations.
271. The method of any of claims 268-270, wherein the training uses a CTC loss function.
272. The method of any of claims 250-271, wherein each action of the avatar is decoded from one or more action representations using the autoencoder.
273. The method of any of claims 268-272, wherein the plurality of actions includes at least two of speech, orofacial movements for speech, and / or non-speech communicative gestures.
274. The method of claim 273, wherein the decoding machine learning model is trained to discriminate between actions performed by different regions of the body and / or between speech associated actions and non-speech communicative gestures.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 275. The method of claim 274, wherein the decoding machine learning model is trained to discriminate between finger-flexions and actions associated with attempted speech.
276. The method of any of claims 268-275, wherein the machine learning model is trained to decode orofacial movements for speech using action representations derived from reference avatar animations of the avatar performing the orofacial movements and brain electrical signal data associated with orofacial movements attempted by the subject.
277. The method of any of claims 268-275, wherein the avatar is controlled to perform orofacial movements based on speech decoded from the recorded brain electrical signal data.
278. The method of claim 277, wherein the avatar is controlled to perform orofacial movements using a speech-to-gesture algorithm.
279. The method of any of claims 268-277, wherein a single decoding machine learning model is trained for speech, orofacial speech movement, and non-speech communicative gesture actions.
280. The method of any of claims 268-277, wherein separate decoding machine learning models are trained for speech associated actions and non-speech communicative gestures.
281. The method of any of claims 268-277 and 280, wherein separate decoding machine learning models are trained for speech and orofacial movements for speech.
282. The method of claims 280 or 281, wherein the avatar is controlled using actions decoded from multiple machine learning models.
283. The method of claim 282, wherein the actions decoded from the multiple machine learning models occur concurrently and the avatar is controlled to perform the actions simultaneously.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 284. The method of any of claims 273-283, wherein a decoded non-speech communicative gesture action affects how the avatar performs a decoded speech associated action.
285. The method of claim 284, wherein a decoded emotional expression affects the inflection of decoded speech performed by the avatar.
286. The method of any of claims 262-285, wherein the avatar is used by the subject to communicate with one or more individuals in person.
287. The method of any of claims 262-286, wherein the avatar is used by the subject to communicate with one or more individuals in an interactive virtual environment.
288. The method of any of claims 262-287, wherein the avatar is used by the subject to play a video game.
289. The method of any of claims 262-288, wherein the avatar is used by the subject for therapy.
290. The method of claim 289, wherein the avatar is used by the subject for physical therapy.
291. The method of claim 290, wherein the physical therapy involves regaining mobility of a body part after an injury.
292. The method of any of claims 262-291, wherein actions performed by the avatar control one or more electronic devices in the subject’s environment.
293. The method of claim 292, wherein actions performed by the avatar control one or more electronic devices in the same room as the subject.
294. The method of claim 293, wherein actions performed by the avatar control one or smart home appliances.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 295. The method of any of claims 262-294, wherein the visual display comprises a computer monitor, a television, and / or a visual projection device.
296. The method of any of claims 262-295, wherein the visual display comprises a virtual reality headset, goggles, or contacts.
297. The method of any of claims 262-295, wherein the visual display comprises an augmented reality headset, goggles, or contacts.
298. The method of any of claims 262-297, wherein the method further comprises generating a virtual environment for the avatar.
299. The method of claim 298, wherein the action performed by the avatar is determined using the virtual environment.
300. The method of any of claims 262-299, wherein previous actions performed by the avatar are used to determine the action performed by the avatar.
301. The method of any of claims 209-300, wherein the processor is provided by a computer, a handheld device, or a headset.
302. The method of claim 301, wherein the handheld device is a cell phone or a tablet.
303. A non-transitory computer-readable medium comprising program instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 209- 302.
304. A kit comprising the non-transitory computer-readable medium of claim 303 and instructions for decoding brain electrical signal data associated with attempted speech by a subject.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 305. A system for controlling an electronic device or system using brain electrical signals configured to perform the method according to any of Claims 209-302.
306. A computer implemented method for controlling an electronic output device or system to perform one or more actions using brain electrical signals, the computer performing steps comprising: receiving the recorded brain electrical signal data associated with the attempted action by the subject; and decoding one or more electronic output device or system actions from the recorded brain electrical signal data using a machine learning model.
307. The computer implemented method of claim 306, wherein the decoding machine learning model comprises a neural network.
308. The computer implemented method of claim 307, wherein the neural network comprises one or more convolutional layers.
309. The computer implemented method of claims 307 or 308, wherein the neural network is bidirectional.
310. The computer implemented method of claim 309, wherein the neural network is a recurrent neural network (RNN).
311. The computer implemented method of claim 310, wherein the RNN comprises gated recurrent units (GRUs) or long-short term memory (LSTM).
312. The computer implemented method of claim 311, wherein the RNN comprises GRUs.
313. The computer implemented method of any of claims 306-312, wherein the subject is limited to a specified action set for the attempted action.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 314. The computer implemented method of claim 313, wherein the action set comprises actions for expressing basic emotions and / or communicating caregiving needs.
315. The computer implemented method of claim 314, wherein the action set comprises 6 actions or more.
316. The computer implemented method of claim 315, wherein the action set comprises 100 actions or more.
317. The computer implemented method of claim 316, wherein the action set comprises 500 actions or more.
318. The computer implemented method of any of claims 313-317, wherein the subject is limited to two or more action sets for the attempted action.
319. The computer implemented method of claim 318, wherein the subject may switch between action sets.
320. The computer implemented method of any of claims 306-319, wherein each action of the device or system is discretized into one or more action representations.
321. The computer implemented method of claim 320, wherein the action representations are decoded from the recorded brain electrical signal data.
322. The computer implemented method of claim 321, wherein each action of the output device or system is decoded from one or more action representations.
323. The computer implemented method of any of claims 320-322, wherein the discretization is performed using an autoencoder.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 324. The computer implemented method of claim 323, wherein the autoencoder comprises one or more convolutional layers.
325. The computer implemented method of claims 323 or 324, wherein the autoencoder uses one or more rectified linear unit (ReLU) activations.
326. The computer implemented method of claim 325, wherein the autoencoder comprises a vector-quantized variational autoencoder (VQ-VAE).
327. The computer implemented method of any of claims 323-326, wherein the discretization occurs during training of the autoencoder.
328. The computer implemented method of claim 327, wherein the training is self-supervised training.
329. The computer implemented method of any of claims 306-328, wherein the electronic output device or system comprises a visual display and / or a loudspeaker.
330. The computer implemented method of claim 329, wherein the visual display and / or loudspeaker is configured to present a humanoid avatar.
331. The computer implemented method of claim 330, wherein the one or more electronic output device or system actions comprise actions performed by the avatar.
332. The computer implemented method of claim 331, wherein actions performed by the avatar include speech, orofacial movements for speech, and / or non-speech communicative gestures.
333. The method of claim 332, wherein the non-speech communicative gestures include emotional expressions using facial muscles.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 334. The computer implemented method of any of claims 331-333, wherein the method further comprises: receiving reference avatar animations for a plurality of actions; encoding each reference avatar animation into a temporal sequence of discrete action representations using the autoencoder; receiving brain electrical signal data associated with the attempted action by the subject for each one of the plurality of actions; training the decoding machine learning model to predict the most likely discrete action representation associated with a segment of electrical signal data using the action representations derived from the reference avatar animations and the brain electrical signal data associated with the attempted action for each one of the plurality of actions.
335. The computer implemented method of claim 334, wherein separate decoding machine learning models are trained for speech associated actions and non-speech communicative gestures.
336. The computer implemented method of claims 334 or 335, wherein the avatar is used by the subject to communicate with one or more individuals in person.
337. The computer implemented method of any of claims 334-336, wherein the avatar is used by the subject to communicate with one or more individuals in an interactive virtual environment.
338. The computer implemented method of claim 337, wherein the method further comprises generating a virtual environment for the avatar.
339. The computer implemented method of claim 338, wherein the action performed by the avatar is determined using the virtual environment.
340. The computer implemented method of any of claims 331-339, wherein previous actions performed by the avatar are used to determine the action performed by the avatar.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 341. A non-transitory computer-readable medium comprising program instructions that, when executed by a processor, cause the processor to perform the computer implemented method of any one of claims 306-340.
342. A kit comprising the non-transitory computer-readable medium of claim 341 and instructions for decoding brain electrical signal data associated with attempted speech by a subject.
343. A system for controlling an avatar to perform one or more actions using brain electrical signals, the system comprising: a neural recording device comprising an electrode adapted for positioning at a location in a sensorimotor cortex region of the brain of the subject to record brain electrical signal data associated with an attempted action by the subject; a processor programmed to use a machine learning model to decode an avatar animation from the recorded brain electrical signal data; an interface in communication with a computing device, said interface adapted for positioning at a location on the head of the subject, wherein the interface receives the brain electrical signal data from the neural recording device and transmits the brain electrical signal data to the processor; and a display component for displaying the avatar action from the recorded brain electrical signal data.
344. The system of claim 343, wherein the subject has difficulty with said speech because of anarthria, a stroke, a traumatic brain injury, a brain tumor, or amyotrophic lateral sclerosis.
345. The system of claims 343 or 344, wherein the neural recording device is adapted for positioning on a surface of the sensorimotor cortex region or within the sensorimotor cortex region.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 346. The system of any of claims 343-345, wherein the electrode of the neural recording device is adapted for positioning on the pial surface of the sensorimotor cortex.
347. The system of any of claims 343-345, wherein the neural recording device is adapted to be positioned such that the recording device covers regions associated with speech production and language perception.
348. The system of claim 347, wherein the covered regions include a middle portion of the superior and / or middle temporal gyrus, the precentral gyrus, and / or the postcentral gyrus.
349. The system of any of claims 343-348, wherein the neural recording device is adapted to be centered on the central sulcus.
350. The system of any of claims 343-349, wherein the neural recording device comprises an electrocorticography (ECoG) electrode array.
351. The system of claim 350, wherein the ECoG electrode array is a high-density array.
352. The system of claim 351, wherein the high-density array comprises 200 electrodes or more.
353. The system of any of claims 350-352, wherein the electrode array comprises non- penetrating surface electrodes.
354. The system of any of claims 343-353, wherein the interface comprises a percutaneous pedestal connector attached to the subject's cranium.
355. The system of any of claims 343-354, wherein the interface further comprises a headstage connected to the percutaneous pedestal connector.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 356. The system of claim 355, wherein the headstage is configured to process and digitize the brain electrical signal data before transmitting the data to the processor.
357. The system of any of claims 343-356, wherein the electrical signal data comprises high- gamma frequency content features.
358. The system of any of claims 343-357, wherein the electrical signal data comprises neural oscillations in a range from 70 Hz to 150 Hz.
359. The system of any of claims 343-358, wherein the electrical signal data comprises low- frequency signals.
360. The system of claim 359, wherein the electrical signal data comprises neural oscillations in a range from 0.3 Hz to 17 Hz.
361. The system of any of claims 343-360, wherein the brain electrical signal data is recorded from a sensorimotor cortex region selected from the precentral gyrus, postcentral gyrus, superior temporal gyrus, middle temporal gyrus, or any combination thereof.
362. The system of any of claims 343-361, wherein the subject is limited to a specified action set for the attempted action.
363. The system of claim 362, wherein the action set comprises actions for expressing basic emotions and / or communicating caregiving needs.
364. The system of claims 362 or 363, wherein the action set comprises 6 actions or more.
365. The system of claim 364, wherein the action set comprises 100 actions or more.
366. The system of any of claims 362-365, wherein the subject is limited to two or more action sets for the attempted action.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 367. The system of claim 366, wherein the processor is configured to switch between action sets.
368. The system of any of claims 343-367, wherein the decoding machine learning model comprises a neural network.
369. The system of claim 368 wherein the neural network comprises one or more convolutional layers.
370. The system of claims 368 or 369, wherein the neural network is bidirectional.
371. The system of claim 370, wherein the neural network is a recurrent neural network (RNN).
372. The system of claim 371, wherein the RNN comprises gated recurrent units (GRUs) or long-short term memory (LSTM).
373. The system of claim 372, wherein the RNN comprises GRUs.
374. The system of any of claims 343-373, wherein the processor is further programmed to discretize each avatar action into one or more action representations.
375. The system of claim 374, wherein the action representations are decoded from the recorded brain electrical signal data.
376. The system of claim 375, wherein each avatar action is decoded from one or more action representations.
377. The system of any of claims 374-376, wherein the discretization is performed using an autoencoder.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 378. The system of claim 377, wherein the autoencoder comprises one or more convolutional layers.
379. The system of claims 377 or 378, wherein the autoencoder uses one or more rectified linear unit (ReLU) activations.
380. The system of claim 379, wherein the autoencoder comprises a vector-quantized variational autoencoder (VQ-VAE).
381. The system of any of claims 377-380, wherein the discretization occurs during training of the autoencoder.
382. The system of claim 381, wherein the training is self-supervised training.
383. The system of any of claims 343-382, wherein actions performed by the avatar include speech, orofacial movements for speech, and / or non-speech communicative gestures.
384. The system of claim 383, wherein the speech orofacial movements comprise one or more of: a tongue tip raise, tongue retraction, tongue body raise, tongue advance, lip rounding, pinching nostril flare, upper lip pull, lower lip tuck, lower lip push, lower lip pull, lip flare, jaw opening, lip compression, and / or lip adduction.
385. The system of claim 383, wherein the non-speech communicative gestures comprise the abduction, adduction, flexion, extension, and / or circumduction of one or more body parts.
386. The system of claims 383 or 385, wherein the non-speech communicative gestures include emotional expressions using facial muscles.
387. The system of claim 58, wherein the emotional expressions comprise happy, sad, and / or surprised expressions.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 388. The system of any of claims 377-387, wherein the processor is further programmed to: obtain reference avatar animations for a plurality of actions; encode each reference animation into a temporal sequence of discrete action representations using the autoencoder; record brain electrical signal data associated with the attempted action by the subject for each one of the plurality of actions; train the decoding machine learning model to predict the most likely discrete action representation associated with a segment of electrical signal data using the action representations derived from the reference avatar animations and the brain electrical signal data associated with attempted action for each one of the plurality of actions.
389. The system of claim 388, wherein the decoding machine learning model is trained to learn mappings between neural activity patterns of electrical signals in the brain electrical signal data and the discrete action representations.
390. The system of claims 388 or 389, wherein the training uses a CTC loss function.
391. The system of any of claims 377-390, wherein each action of the avatar is decoded from one or more action representations using the autoencoder.
392. The system of any of claims 388-391, wherein the plurality of actions includes at least two of speech, orofacial movements for speech, and / or non-speech communicative gestures.
393. The system of claim 392, wherein the decoding machine learning model is trained to discriminate between actions performed by different regions of the body and / or between speech associated actions and non-speech communicative gestures.
394. The system of any of claims 377-393, wherein separate decoding machine learning models are trained for speech associated actions and non-speech communicative gestures.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 395. The system of claim 394, wherein separate decoding machine learning models are trained for speech and orofacial movements for speech.
396. The system of claims 394 or 395, wherein the avatar is controlled using actions decoded from multiple machine learning models.
397. The system of claim 396, wherein the actions decoded from the multiple machine learning models occur concurrently and the avatar is controlled to perform the actions simultaneously.
398. The system of any of claims 343-397, wherein the visual display comprises a computer monitor, a television, and / or a visual projection device.
399. The system of any of claims 343-398, wherein the visual display comprises a virtual reality headset, goggles, or contacts.
400. The system of any of claims 343-399, wherein the visual display comprises an augmented reality headset, goggles, or contacts.
401. The system of any of the claims 343-400, wherein the processor is further programmed to generate a virtual environment for the avatar.
402. The system of claim 401, wherein the action performed by the avatar is determined using the virtual environment.
403. The system of any of the claims 343-402, wherein previous actions performed by the avatar are used to determine the action performed by the avatar.
404. The system of any of claims 343-403, wherein the processor is provided by a computer, a handheld device, or a headset.Atty Docket No: UCSF-723WO Client Ref No.: SF-2023-165-2-PCT-0 405. The system of claim 404, wherein the handheld device is a cell phone or a tablet.
406. A kit comprising the system of any of claims 343-405 and instructions for decoding brain electrical signal data associated with attempted speech by a subject.