Decoding language from non-invasive brain recordings

JP2025514868A5Pending Publication Date: 2026-03-02BOARD OF RGT THE UNIV OF TEXAS SYST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024549665
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-02-22
Filing Date
2023-02-22
Publication Date
2026-03-02

AI Technical Summary

Technical Problem

Existing language decoding devices using non-invasive brain records are limited in their ability to decipher continuous speech, as they can only decode single words, and invasive methods pose risks and require neurosurgery for maintenance.

Method used

The system uses non-invasive brain records to detect changes in blood oxygen levels associated with neural activity, employing linguistic reconstruction models to predict brain activity measurements and decode them into continuous language sequences.

Benefits of technology

This approach enables the reconstruction of continuous natural language from brain activity, improving communication for individuals with speech impairments without the risks associated with invasive methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Embodiments can obtain brain activity measurements and decode them into continuous language. Embodiments can use non-invasive brain recordings, such as functional magnetic resonance imaging (fMRI) and functional near-infrared spectroscopy (fNIRS), to detect changes in blood oxygen levels associated with neural activity. These brain recordings can be used in language reconstruction models, including neural language models that predict the next word in a sequence, and encoding models that decode the recordings into continuous language or word sequences.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 312,801, filed February 22, 2022, which is incorporated by reference herein in its entirety for all purposes. The present invention relates to decoding language from non-invasive brain recordings. [Background technology]

[0002] Brain-computer interfaces (BCIs), which act as a direct communication pathway between the brain and external devices (such as computers), have opened the door to a wide range of applications. Specifically, language-decoding BCIs can translate what a person hears, reads, or thinks from brain activity into text. This can be useful in helping people who are cognitively normal but unable to speak. For example, this can be useful for locked-in syndrome or motor neuron diseases (such as ALS).

[0003] However, existing language decoding devices have significant problems. Current methods using non-invasive (i.e., non-surgical) recordings can only decode single words, preventing users from engaging in smooth conversations. Decoding continuous speech is possible using invasive recordings that require neurosurgery, but surgically implanting the recording device poses additional risks to the user. The quality of the signal obtained from a surgically implanted recording device can also degrade over time due to scarring, requiring further neurosurgery to replace or maintain the device.

[0004] Embodiments of the present invention address these and other problems, individually and collectively. Summary of the Invention

[0005] Certain embodiments of the present disclosure provide methods and systems that use non-invasive brain recordings to detect changes in blood oxygen levels associated with neural activity. These brain recordings can be used by language reconstruction models, including neural language models that predict the next word in a sequence, and encoding models that decode the recordings into continuous language or word sequences.

[0006] These and other embodiments of the present disclosure are described in more detail below. For example, other embodiments are directed to systems, apparatus, and computer-readable media associated with the methods described herein.

[0007] The nature and advantages of the embodiments of the present disclosure may be better understood with reference to the following detailed description and accompanying drawings. [Brief description of the drawings]

[0008] [Figure 1] Provides a supervised learning flow diagram for training an encoding model to predict brain activity measures for continuous language or word sequences.

[0009] [Diagram 2] An example flow diagram of each step of FIG. 1 is provided.

[0010] [Diagram 3] A flow diagram of language reconstruction that translates brain activity measurements into continuous language, or word sequences, is provided.

[0011] [Figure 4] An example flow diagram of each step of FIG. 3 is provided.

[0012] [Diagram 5] 1 provides a flow chart of a method for performing language reconstruction that results in prediction of continuous language.

[0013] [Figure 6A]A predicted analysis of the text by the language decoder is provided. [Figure 6B] A predicted analysis of the text by the language decoder is provided.

[0014] [Figure 7] A table of linguistic similarity scores is provided.

[0015] [Figure 8A] We provide an analysis of decoding from different cortical language networks. [Figure 8B] We provide an analysis of decoding from different cortical language networks.

[0016] [Figure 9A] Provides uses of language decoders and privacy implications. [Figure 9B] Provides uses of language decoders and their privacy implications.

[0017] [Figure 10] Provides an analysis of the causes of decoding errors.

[0018] [Figure 11] An analysis of the performance of the encoding model and the word-time decoder (ie, the word-rate model) is provided.

[0019] [Figure 12] The results of the analysis of perceived speech discrimination performance and imagined speech discrimination performance are presented.

[0020] [Figure 13A] The present invention provides analytical results of performance evaluation of language decoding device prediction. [Figure 13B] The present invention provides analytical results of performance evaluation of language decoding device prediction.

[0021] [Figure 14] We provide an analysis of language decoding across cortical regions.

[0022] [Figure 15] We provide analytical results comparing language decoding performance across different experiments.

[0023] [Figure 16] We provide analytical results on the performance of the cross-subject encoding model and the word-time decoder.

[0024] [Figure 17] An analysis of decoding performance as a function of training data is presented.

[0025] [Figure 18] We provide an analysis of the decoding performance at lower spatial resolution.

[0026] [Figure 19] Provides analysis results of decoder ablation.

[0027] [Figure 20] Provides analysis of separated encoding models and language model scores.

[0028] [Figure 21] FIG. 1 illustrates a block diagram of an exemplary computer system usable with systems and methods according to embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0029] Previous brain-computer interfaces have demonstrated the ability to decode speech articulatory or other signals from intracranial recordings to restore communication to people who have lost the ability to speak. Although effective, these language decoders require invasive neurosurgery and are therefore unsuitable for most other uses. Language decoders that use noninvasive recordings may be more widely adopted and could be used for both remedial and augmentative applications. Although noninvasive brain recordings can capture many types of linguistic information, previous attempts to decode this information have been limited to identifying one output among a small set of possibilities, and it remains unclear whether current noninvasive recordings have the spatial and temporal resolution necessary to decode continuous language.

[0030] To solve this problem, an embodiment can obtain non-invasive brain recordings and determine a language decoder that uses continuous natural language to reconstruct perceived or imagined stimuli. While methods that detect blood oxygen level changes associated with neural activity (such as fMRI) can have good spatial specificity, the blood oxygen level dependent (BOLD) signal that this method measures is notoriously slow, i.e., impulses of neural activity cause the BOLD to rise and fall over approximately 10 seconds. For naturally spoken English (more than 2 words per second), this means that each brain image may be affected by 20 or more words. Thus, to decode continuous language, a misposed inverse problem must be solved, as there may be many more words to decode than there are brain images. This embodiment can achieve this by generating candidate word sequences, scoring the likelihood that each candidate will elicit the recorded brain response, and then selecting the best candidate.

[0031] To compare the word sequences with the subject's brain responses, this embodiment can use an encoding model that predicts how the subject's brain will respond to natural language. In this embodiment, the brain responses can be recorded while the subject listens to a naturally spoken story. In this embodiment, the encoding model can be trained on this dataset by extracting semantic features that capture the meaning of the stimulus phrases and using linear regression to model how the semantic features affect the brain response (602 in FIG. 6A). Given any word sequence, the encoding model predicts with a fair amount of accuracy how the subject's brain will respond when hearing that sequence (FIG. 11). The encoding model can then score the likelihood that the word sequence elicited the recorded brain response by measuring how well the recorded brain response matches the predicted brain response.

[0032] Theoretically, the most likely stimulus word can be identified by comparing the recorded brain response with the prediction of the encoding model for all possible word sequences. However, the number of possible word sequences is too large to make this approach practical, as most of such sequences may not resemble natural language. To restrict the candidate sequences to well-formed English, this embodiment can use a neural network language model trained on a large dataset of word sequences. Given any word sequence, the language model can predict the likely next word.

[0033] However, even with the constraints imposed by the language model, it is computationally infeasible to generate and score all candidate sequences. To efficiently search for the most likely word sequences, this embodiment can use a beam search algorithm that can generate candidate sequences for each word. In a beam search, the language decoder can maintain a beam (i.e., a hypothesis beam) that contains the k most likely candidate sequences at any time. When a new word is detected based on brain activity in the auditory and speech domains (FIG. 11), the neural language model can generate continuations of each sequence in the beam using previously decoded words as context. The encoding model can then score the likelihood that each continuation elicited the recorded brain response, and the k most likely continuations can be retained in the beam for the next time step (604 in FIG. 6A). This process can continuously approximate the most likely stimulus word over any time.

[0034] In short, in this embodiment, brain activity measurements can be taken and decoded into sequential language. This embodiment can use non-invasive brain recordings such as functional magnetic resonance imaging (fMRI) and functional near-infrared spectroscopy (fNIRS) to detect changes in blood oxygen levels associated with neural activity and perform brain activity measurements. The current time interval between brain activity measurements obtained using FMRI and fNIRS can be 1 to 5 seconds. In some embodiments, electrical signals from neurons or magnetic fields in the brain such as magnetoencephalography (MEG) can be used to record brain measurements of neural processes. The measured brain activity may be used to decode language, but is not necessarily caused by language. For example, a person may watch a silent movie, and brain recordings of the person watching the silent movie can be recorded and decoded into sequential language.

[0035] I. Training the encoding model FIG. 1 shows supervised learning to train an encoding model 111 to predict brain activity measurements of stimuli 102. A brain imaging device 104 measures brain activity while a subject receives stimuli 102. Annotation entities 106 and a neuro-linguistic model 108 are used to convert a word sequence (stimuli 102) into a sequence of word embedding vectors. Then, the actual brain measurements are used for comparison by linear regression 110 to train an encoding model 111 that converts the sequence of word embedding vectors into a brain activity image. FIG. 2 shows examples of different measurements obtained by FIG. 1 to train an encoding model 220. The flow diagram of FIG. 2 has the same process as FIG. 1.

[0036] A. Brain activity measurement In S112, the brain imaging device 104 measures activity in different parts or voxels of the brain while the subject is thinking, reading, or listening to the stimuli 102. The brain imaging device 104 may be an fMRI, fNIRS, etc., capable of measuring the subject's blood oxygen levels. An example of collecting activity in different parts of the brain may include collecting MRI data on a Siemens 3TSkyra scanner using a 64-channel Siemens volume coil. Functional scans may be collected using gradient echo EPI with repetition time (TR)=2.00 seconds, echo time (TE)=30.8 milliseconds, flip angle=71°, multiband factor (simultaneous multislice)=2, voxel size=2.6mmx2.6mmx2.6mm (slice thickness=2.6mm), matrix size=(84, 84), and field of view=220mm. Anatomical data can be collected using a T1-weighted multi-echo MP-RAGE sequence on the same 3T scanner with voxel size = 1mmx1mmx1mm following the Freesurfer morphometry protocol. Anatomical data can be collected on a Siemens 3T TIM Trio scanner with a 32-channel Siemens volume coil using the same sequence.

[0037] The stimuli 102 are measured by the brain imaging device 104 at the precise times at which the annotating entity 106 labels the word sequences in time. For example, when a subject thinks the word sequence "I have a dog," the precise time at which the subject thought each word must be recorded by the annotating entity 106. Thus, the stimuli 102 are typically presented to the subject using a computer that is synchronized with the brain imaging device 104, allowing the annotating entity to time-align the labeled stimuli with the brain activity measurements.

[0038] The labeling of word sequences by the annotating entity 106 and the brain activity measurements made by the brain imaging device 104 do not necessarily have to occur simultaneously. For example, the stimuli 102 and brain activity images may be pre-recorded and the recordings provided to the annotating entity 106 at a later time. The annotating entity 106 may then label the times of words in the recordings at other times.

[0039] Stimulus 102 and step S112 are shown in FIG. 2 with an audio stimulus 202 and a brain 204. The audio stimulus 202 is received by the brain 204. This may include the subject at the brain 204 thinking, reading, or listening to the audio stimulus 202. As the brain 204 receives the audio stimulus 202, a brain imaging device 206 may measure the subject's brain activity. The brain imaging device 206 may measure brain activity for each voxel of the brain, with each voxel eliciting a different measurement. An example of a voxel measurement is shown in brain measurement 208, with up to n voxels, with each voxel having its own measurement over time t.

[0040] In S114, the brain activity measurements of different voxels of the brain are sent to the linear regression 110. The acquisition time for the brain imaging device 104 to acquire the brain activity images may vary depending on the brain imaging device 104. For fMRI, the acquisition time for acquiring the brain activity images may be 2 seconds. The stimuli 102 may be decomposed into word sequences such that each sequence of words is correlated with the number of words perceived / received by the subject within one or more acquisition periods. Even if there is one brain activity image measured per acquisition period and one word sequence uttered per acquisition period, words from previous periods may affect the brain activity measured in subsequent periods. The linear regression 110 uses the brain activity measurements made by the brain imaging device 104 to compare with the word sequences and find an encoding model 111.

[0041] In cases where certain brain regions are more intact or more accessible, this embodiment can target those specific brain regions. For example, brain activity measurements (e.g., whole-brain MRI data) can be divided into three cortical regions: the vocal network, the parieto-temporal-occipital association cortex, and the prefrontal cortex. Examples of how brain activity can be measured are described in more detail below.

[0042] The auditory network can be functionally localized in each subject using auditory and motor localizers. Auditory localizer data can be collected in a single 10-minute scan. Subjects can listen to 10 repetitions of 1-minute auditory stimuli, including 20 seconds of music (e.g., Arcade Fire), speech (e.g., Ira Grass, This American Life), and natural sounds (e.g., babbling brook). To determine whether a voxel responded to the auditory stimuli, we took the average response across the 10 repetitions, subtracted this average response from each single-trial response to calculate a single-trial residual, and quantified the repeatability of the voxel response using the F-statistic, which can be calculated by dividing the variance of the single-trial residual by the variance of the single-trial response. This metric can directly quantify the amount of variance in the voxel response that can be explained by the average response across repetitions. The repeatability map can be used by human annotators to define the auditory cortex (AC). Motor localizer data can be collected in two identical 10-minute scans. Subjects can be cued to perform six different tasks ("hands", "feet", "mouth", "speak", "saccade", and "rest") in a random order in 20-s blocks. For the "speak" cue, subjects can be instructed to self-generate a story without vocalizing. Linear models can be estimated to predict the response of each voxel using the six cues as categorical features. Weight maps of the "speak" features can be used by human annotators to define Broca's area and superior ventral premotor cortex (sPMv) speech regions. There is widespread agreement that these language areas, unlike the parieto-temporal-occipital association cortex and prefrontal cortex, are necessary for speech perception and production. Most existing invasive language decoding devices can record brain activity from these speech regions.

[0043] Freesurfer ROIs allow the location of parietal-temporal-occipital association cortices and prefrontal cortices anatomically for each subject. Parietal-temporal-occipital association cortices can be defined using the following labels: superior parietal, inferior parietal, superior limbic, posterocentral, precuneus, superior temporal, middle temporal, inferior temporal, bankust, fusiform, lateral temporal, entorhinal cortex, temporal pole, parahippocampal, lateral occipital, lingual, cuneate, pericranial, posterior cingulate cortex, and isthmus cingulate cortex. Prefrontal cortices can be defined using the following labels: superior frontal, rostral middle frontal, caudal middle frontal, paraopercular, paratriangular, paracuspal, lateral orbitofrontal, medial orbitofrontal, precentral, paracentral, frontal pole, rostral anterior cingulate cortex, and caudal anterior cingulate cortex. Voxels identified as part of the speech network (AC, Broca's area, and sPMv speech region) can be excluded from the parieto-temporal-occipital and prefrontal cortices. A functional definition can be used for the speech network because previous studies have shown that the anatomical location of the speech network varies across subjects, whereas an anatomical definition can be used because the parieto-temporal-occipital and prefrontal cortices are broad and functionally diverse.

[0044] B. Labeling At S116, the stimuli 102 are labeled by the annotating entity 106. The annotating entity 106 labels the stimuli 102 by labeling the identifier (which word is used) and time of each word thought, read, or heard by the user. For example, if the word sequence is "I have a dog", the annotating entity 106 can label the time of each word in the sequence. The time of each word may be labeled by speaking the stimuli 102 directly to the user and the annotating entity 106, so that the annotating entity 106 may be able to accurately measure the exact time of the word sequence spoken to or heard by the user and synchronize the stimulus annotation with the brain activity measurement. The annotating entity 106 may be any device capable of recognizing words in the stimuli 102, such as automatic speech recognition software, or may be a human being capable of annotating the time. The annotating entity 106 may be performed by human annotation for accurate measurement, but may also be performed by software, such as a natural language processor. In one embodiment, a human annotator manually transcribes the words in the stimulus audio files, software is used to automatically predict word times by aligning the transcribed words with the stimulus audio files, and the human annotator manually verifies the alignment.

[0045] Step S116 is illustrated in FIG 2 by human annotations 210 and stimulus transcript 212. The audio stimulus 202 is labeled by a human with the time of each word. This is represented as human annotations 210. This produces a stimulus transcript 212 in which each word in the transcript is labeled with a time. For example, in the stimulus transcript 212, each word in the word sequence "I grew up in" is labeled with a time: "I" is labeled with 0.1, "grew" with 0.2, "up" with 0.5, and "in" with 0.6.

[0046] At S118, the neuro-language model 108 receives the word sequence of the stimulus 102 that has been labeled with time by the annotating entity 106. The labeled word sequence undergoes quantitative feature transformation using the neuro-language model 108 to modify each labeled word in the word sequence as a word embedding (feature) vector. The neuro-language model 108 extracts a meaningful representation of each word in the sequence rather than predicting the next word.

[0047] The extracted word embeddings of each word in a sequence may represent something about both the meaning (semantics) and form (syntax) of the word. The extracted word embeddings are automatically learned and determined by a neural network in the neural language model 108. For example, if a word sequence is "I have a dog," the neural network in the neural language model 108 may automatically determine three embeddings or features for the word "I," which may represent something about both the meaning and the form. These extracted embeddings or features may be written as vectors of numerical values, with each value representing a different extracted feature of the word.

[0048] An example of a neural language model 108 may be a generative pre-trained transformer (GPT) trained to predict the next word in a sequence given the previous word. GPTs can be used to extract semantic features from linguistic stimuli. To successfully perform the next word prediction task, GPTs can learn how to extract quantitative features that capture the meaning of the input sequence. Given a word sequence S=(s1,s2,...,s n ), the activations of the GPT hidden layer select the most recent word in the context s n It is possible to provide a vector embedding that represents the meaning of

[0049] Step S118 is shown in Figure 2 with the stimulus transcript 212. The stimulus transcript 212 is input into a neuro-linguistic model 214 to generate word embeddings 216. As shown in the word embeddings 216, each word in the sequence "I grew up in" has a word embedding vector with three different features beneath each word.

[0050] C. Synchronization To map extracted stimulus features to brain activity images, we would ideally need to measure brain activity for every word. However, there are typically 4-8 words in a word sequence, ensuring that there is one word sequence per brain activity image in one acquisition time interval. Thus, there are typically 4-8 times as many words as there are brain activity images.

[0051] To match the number of stimulation measurements with the number of brain activity images, the words in the sequence are downsampled to generate a single averaged word embedding vector. This is done by taking the average of the word embedding vectors of the words in the sequence. In some embodiments, the downsampling may be performed by other techniques such as the Lanczos kernel. For example, if the word sequence is "I have a dog", the four word embedding vectors of the word sequence "I have a dog" are averaged to generate a single averaged word embedding vector. This average may weight the word embeddings by the difference between the word time and the brain image acquisition time.

[0052] A word sequence does not have just one immediate response; it has a responsiveness that evolves over time. A word sequence can affect approximately the next 8 seconds of brain activity measurements, or four periods of time when measured using fMRI. For example, if the first word sequence is "I have a dog," that sequence will not only affect the first brain activity image created by the brain imaging device 206, but may also affect the second, third, or fourth brain activity images. Thus, each brain activity image at a given time can be modeled as a function of the words that occurred in the previous four acquisition periods, or the previous four word sequences.

[0053] Because each sequence of words may have a different degree of influence on the brain activity image at different times in the future, the learned convolution kernel is used to assign different coefficients to each of the previous four word sequences, or the previous four averaged word embedding vectors. For example, the oldest word sequence may have less influence on the brain activity image than the most recent word sequence. The convolution kernels with different weights are applied to the previous four averaged word embedding vectors to generate the final word embedding vector. The final word embedding vector is used to predict the brain activity image.

[0054] For example, if the first word sequence is "I have a dog", the second word sequence is "and a puppy", the third word sequence is "that I raised", and the fourth word sequence is "since I was six", then each sequence of words will produce a single averaged word embedding vector, and each of the four averaged vectors will be passed through a convolution kernel to produce the first final word embedding vector, which will be used to match the fourth brain activity image, and the fifth brain activity image will be correlated with the second, third, fourth, and fifth word sequences.

[0055] D. Linear regression Once the final word embedding vector is determined in S120, the final word embedding vector is sent to linear regression 110. Linear regression 110 can take each feature in the final word embedding vector and determine a linear mapping to voxel space. This is done by using regularized linear regression to determine an encoding model 111 that predicts how the final word embedding vector will affect the brain response (the response of each voxel) by fitting each feature of the vector to the voxel's brain activity measurement. Since encoding model 111 is mapping the features of the final word embedding vector to brain voxels, encoding model 111 is a matrix with dimensions J×L, where J is the number of features and L is the number of voxels.

[0056] Linear regression 110 is shown in FIG. 2 along with linear regression 218. The word sequences in word embeddings 216 are transformed, averaged, and convolved into a final word embedding that attempts to map each feature in the final word embedding to each voxel in the brain measurement 208. Linear regression 218 is used to determine an encoding model 220 that maps each feature in word embeddings 216 to each voxel in the brain measurement. If there are three features in word embeddings 216 and n voxels in brain measurement 208, then encoding model 220 has dimensions of 3×n, with each row representing a feature and each column representing a voxel.

[0057] In some embodiments, a word-time decoder can also be trained using linear regression 110. The word-time decoder predicts when a subject thinks, reads, or hears a word. Word times are recorded by the annotating entity 106, but in a real language reconstruction, the time each word is spoken is unknown. The word-time decoder learns a mapping between brain activity measurements and a vector of word rates, i.e., the number of words in a sequence, at each acquisition time. This mapping is estimated for each word sequence in the stimulus 102 using linear regression 110. For example, the input can be the TxA response for the A voxel associated with word timing information, the output can be Tx1, and the learned weights can be Ax1.

[0058] To predict word rate in perceived speech, the brain response can be restricted to auditory cortex. To predict word rate in imagined speech and perceived movies, the brain response can be restricted to Broca's area and sPMv speech area. A separate linear temporal filter with four delays (t+1, t+2, t+3, and t+4) can be fitted to each voxel. For a TR of 2 seconds, this is accomplished by concatenating responses after 2, 4, 6, and 8 seconds to predict the word rate at time t. Given a new brain response, the model can predict the word rate at each acquisition. The word time can then be predicted by evenly dividing the time between successive acquisitions (e.g., 2 seconds) by the predicted word rate (rounded to the nearest nonnegative integer).

[0059] In S122, the encoding model 111 is determined by the linear regression 110 that best converts the final word embedding vectors to actual brain activity measurements. The convolution kernel weights can also be trained in conjunction with the linear regression 110 to estimate the best final word embedding vector that best converts to actual brain activity measurements. The encoding model 111 is adjusted until all brain activity images measured by the brain imaging device 104 have been used to train the encoding model 111. The encoding model 111 is later used by the language reconstruction model of FIG. 3 to derive brain predictions that are compared to the brain measurements.

[0060] II. Reconstructing Language Once the encoding model has been trained, it may be used in a language reconstruction model that takes in a new stimulus (which may be unknown) and reconstructs brain measurements of the new stimulus made by a brain imaging device into a continuous linguistic prediction of the new stimulus.

[0061] FIG. 3 shows a flow diagram of language reconstruction that translates a subject's (user's) brain measurements or images into continuous language or word sequences. In some embodiments, a brain imager 304 outputs brain activity measurements while the subject is receiving stimulation 302. Each brain measurement made by the brain imager 304 is then used by a ranking computer 312 to compare with multiple brain predictions (also called augmented hypotheses) of the continuation to identify the most likely brain prediction. The continuation is constructed by combining each hypothesis (the first initial hypothesis) in the hypothesis beam 306 with a set of continuation words predicted by a neural language model 308. The continuation then passes through an encoding model 310 to convert the continuation into multiple brain predictions. FIG. 4 shows examples of different measurements that FIG. 3 performs to translate brain measurements into word sequences. The flow diagram of FIG. 4 has the same process as FIG. 3.

[0062] According to Bayes' theorem, the distribution P(S|R) of a word sequence (S) given a brain response (R) can be factorized into a prior distribution P(S) of the word sequence and an encoding distribution P(R|S) of the brain response given the word sequence. test Given that, in theory, we can use a language model to evaluate P(S) for all possible word sequences S, and a subject's encoding model to evaluate P(R test |S) to find the most likely word sequence S test However, the combinatorial structure of natural language makes it computationally infeasible to evaluate all possible word sequences. Instead, the most likely word sequences can be predicted using a beam search algorithm that includes the hypothesis beam 306, the neural language model 308, the encoding model 310, and the ranking computer 312 of FIG. 3.

[0063] This embodiment, or language decoder, can maintain a beam containing the k most likely word sequences. The hypothesis beam can be initialized with an empty word sequence. As new words are detected by the word rate model (word time decoder), the neural language model can generate continuations for each candidate S in the beam. The neural language model can generate continuations for the predicted words (s) in the candidates. n-i ,...,s n-1 ) to calculate the distribution of the next word P(s n |s n-i ,...,s n-1) Using kernel sampling, we can identify K words that belong to the top p percent of the probability set and have a probability within a factor r of the most likely word. Each of the K words in the kernel can be added to the candidates to form a continuation C.

[0064] The encoding model estimates the likelihood P(R testEach continuation can be scored by |C). The k most likely continuations across all candidates are kept in a hypothesis beam. After iterating through all predicted word times, the decoder can output the most likely candidate sequence.

[0065] The Bayesian decoder according to this embodiment, or the use of a beam search algorithm, can be unique in two important respects. First, existing Bayesian decoders can typically only collect a large empirical prior distribution of images or videos and calculate the P(R|S) of stimuli in the empirical prior distribution. The decoder's predictions can be obtained by selecting the most likely stimuli or by employing weighted combinations of stimuli. In contrast, this embodiment uses a neural language model in advance that can generate completely novel sequences. Second, existing Bayesian decoders can evaluate all stimuli in the empirical prior distribution. In contrast, this embodiment can use a beam search algorithm to efficiently search the combinatorial space of possible sequences, so that the words evaluated at each time point depend on previously decoded words.

[0066] A. Brain Measurements In S314, the stimulus 302 is sent to the brain imaging device 304. The brain imaging device 304 measures the activity of different parts or voxels of the brain when the subject is thinking, reading or listening to the stimulus 302. The brain imaging device 304 may be an fMRI, fNIRS, etc., capable of measuring the subject's blood oxygen level. The brain activity measurements may be taken over different time periods. A single brain activity measurement may include L brain measurements, i.e., L voxels, of the subject in a single time period. For example, the brain imaging device 304 may measure the brain activity measurement in a current time period. The brain activity measurements taken by the brain imaging device 304 and the successive language predictions by the ranking computer 312 do not necessarily occur simultaneously. The brain activity images of the brain activity measurements may be recorded in advance, and the recordings may be provided to the ranking computer 312 later to generate the successive language predictions.

[0067] Step S314 is illustrated in FIG. 4 by a stimulus 402 and a brain 404. The stimulus 402 is received by the brain 404. This may include an audio stimulus 402, which the user's brain 404 thinks, reads, or hears. Once the brain 404 receives the audio stimulus 402, a brain imaging device 406 may measure the user's brain activity. The brain imaging device 406 may measure brain activity for each voxel of the brain, with each voxel eliciting a different measurement. An example of a voxel measurement is shown in brain measurement 408, where there are up to n voxels, and each voxel has its own measurement over time t.

[0068] In S316, the brain activity measurements of different voxels of the brain are sent to the ranking computer 312. Depending on the brain imaging device 304, the acquisition time for the brain imaging device 304 to acquire the brain activity image may vary. For fMRI, the acquisition time for acquiring the brain activity image may be 2 seconds. The stimuli 302 may be word sequences, and each sequence of words is provided to the subject in one acquisition time. Thus, there is one brain activity image and one word sequence per acquisition period, but the words in one period may affect the measurements in the subsequent period.

[0069] B. Neural Language Models The hypothesis beam 306 has N initial hypotheses. Each hypothesis is composed of a word sequence. The hypothesis beam 306 contains the N best combined word sequences in the current model. Each hypothesis can be composed of the same number of words. At the start of the language reconstruction, the default value of the hypothesis beam 306 can contain N initial hypotheses. The N initial hypotheses can be customized based on different use cases. For example, the N initial hypotheses can be pronouns or the top 50 most commonly used words or phrases as a starting point. In one embodiment, a sentence can start with a pronoun ("I", "we", "they", "she", "he"). The starting words can be customized based on the use case of the decoder.

[0070] In S318, each hypothesis is input to the neural language model 308. The neural language model 308 predicts a corresponding set of K possible continuation words for each hypothesis. For example, if hypothesis 1 has a word sequence such as "we went", the corresponding set of K continuation words would be "to", "hiking", etc. Each hypothesis can be combined with a corresponding set of K continuation words. Since each of the N hypotheses has a set of K continuation words, there are a total of N*K continuations after passing through the neural language model 308. A continuation may be a combination of a hypothesis and a continuation word. Each of the N*K continuations is also converted into a set of word embedding vectors, similar to the changes described in step S120. Each word in the continuation has a word embedding vector, and each continuation has a set of word embedding vectors. Thus, the total of N*K continuations includes N*K sets of word embedding vectors.

[0071] The set of K continuation words must be unique for each parent hypothesis. For example, if the two hypotheses are "i grew up in" and "i grew up around", the continuations of the first hypothesis might include "i grew up in Georgia" and "i grew up in suburbia", while the continuations of the second hypothesis might include "i grew up around doctors" and "i grew up around animals". Thus, the set of K continuation words is unique for each parent hypothesis.

[0072] An example of a neural language model 308 may be a Transformer Large Language Model (LLM) such as a Generative Pre-Trained Transformer (GPT). The LLM may be a 12-layer neural network that uses multi-head self-attention to combine the representation of each word in a sequence with the representation of the previous word. The LLM may be trained on a large corpus of books to generate a representation of a sequence (s1,s 2, ...,s n-1 ) next word s n It is possible to predict the probability distribution over

[0073] LLM can estimate a prior probability distribution P(S) over a word sequence. 2, ...,s n-1 ), LLM can compute the probability of observing S in a natural language by multiplying the probability of each word conditional on the previous word. TIFF2025514868000002.tif2692, where s 1:0 is the empty sequence φ.

[0074] LLMs can also be used to extract semantic features from linguistic stimuli (as shown in Figure 1). To successfully perform the task of next word prediction, LLMs can learn to extract quantitative features that capture the meaning of an input sequence. Given a word sequence S = (s1,s 2, ...,s n-1 ) is given 、 The activations in the hidden layer of the LLM are used to find the most recent word in the context s n A vector embedding is provided that represents the meaning of

[0075] Step S318 is illustrated in FIG. 4 with a hypothesis beam 410 and a neuro-language model 412. The hypothesis beam 410 has N hypotheses, shown as "hypothesis 1...hypothesis n." The N hypotheses are passed through a neuro-language model 412, which predicts K continuation words for each hypothesis. This generates an N*K continuation 414, where the number reflects the Nth hypothesis and the letter reflects the Kth continuation word. Thus, "continuation 1.b" would represent a combination of the first hypothesis and the second continuation word predicted by the neuro-language model 412.

[0076] C. Synchronization Each continuation is composed of a word sequence, and each sequence of words correlates to the number of words perceived by the subject in one acquisition time (i.e., period). However, since the words are not labeled with a time and the words in the continuation are predictions made by the neural language model 308, the number of words that compose each sequence is unknown. To solve this problem, the word-time decoder trained in FIG. 1 predicts the word rate at each acquisition time, i.e., the number of words in a sequence, and estimates the word time by dividing the time between acquisitions equally by the predicted word rate. Thus, the word-time decoder predicts the word rate of each sequence and labels the time of each word. In one example, the decoder can perform a count that determines the number of words in each interval, e.g., 8 words for the interval 0-2 seconds, 4 words for the interval 2-4 seconds, and 5 words for the interval 4-6 seconds. The word-rate decoder can be implemented using linear regression.

[0077] To map the extracted stimulus embeddings or features to brain activity images, brain activity measurements are required for every word. However, there are typically 4-8 words in a word sequence, and one word sequence per brain activity image in one acquisition time interval. Thus, there are typically 4-8 times as many words as there are brain activity images.

[0078] To match the number of words and the number of brain activity images, the words in the sequence are downsampled to generate a single averaged word embedding vector. This is done by taking the average of the word embedding vectors of the words in the sequence. For example, if the word sequence is "I have a dog", the four word embedding vectors of the word sequence "I have a dog" are averaged to generate a single averaged word embedding vector. In some embodiments, the downsampling may be performed by other techniques, such as the Lanczos kernel.

[0079] A word sequence does not have just one immediate response; it has a responsiveness that evolves over time. A word sequence can affect approximately the next 8 seconds of brain activity measurement, or four periods of time when measured using fMRI. For example, if the first word sequence is "I have a dog," that sequence will not only affect the first brain activity image created by the brain imaging device 304, but may also affect the second, third, or fourth brain activity image. Thus, each brain activity image at a given time will be a function of the words that occurred in the previous four acquisition periods, or the previous four word sequences.

[0080] Since each sequence of words may have a different degree of influence on the brain activity image at different time points in the future, the learned convolution kernel may be used to assign different coefficients to each of the previous four word sequences, or the previous four averaged word embedding vectors. For example, the oldest word sequence may have less influence on the brain activity image than the latest word sequence. Convolution kernels with different weights are applied to the previous four averaged word embedding vectors to generate a final word embedding vector. The final word embedding vector is determined for each N*K continuation before being input to the encoding model 310. The weights of the convolution kernels may be learned in combination with the weights of the stimulus features during the estimation of the encoding model. Thus, the convolution kernels may include different weights, each weight being determined based on the degree of influence of the averaged word embedding vector and the previous averaged word embedding vector on the subject's brain activity measurement in the current time period.

[0081] D. Prediction of brain activity images Once the N*K final word embedding vectors are determined in S320, the N*K final word embedding vectors are sent to the encoding model 310. In the encoding model 310, each final word embedding vector is transformed into a brain prediction (i.e., a predicted brain response). The encoding model 310 is a feature by voxel matrix that transforms the features of the final word embedding vectors into L voxels of brain activity measurements (i.e., L candidate response values). The encoding model can be further specified in the following example.

[0082] Voxel-wise modeling allows for the extraction of quantitative features from stimulus words and uses regularized linear regression to estimate a set of weights that predict how each feature will affect the BOLD signal in each voxel.

[0083] A stimulus matrix can be constructed from stimuli 102 (e.g., training stories). i ,t i ), for a word sequence s i-5 ,s i-4 ,...,s i-1 ,s i ) to the GPT language model, and the ninth layer gives s i Previous studies have shown that intermediate layers of language models extract optimal semantic features for predicting brain responses to natural language. This allows vector-time pairs (M i , t i ), where M i s i These vectors can be resampled using a 3-row Brancheau filter at a time corresponding to the fMRI acquisition.

[0084] A linearized finite impulse response (FIR) model can be fitted to every cortical voxel in each subject's brain. A separate linear temporal filter with four lags (t-1, t-2, t-3, and t-4 time points) can be fitted to each of the 768 features, resulting in a total of 3,072 features. For a TR of 2 seconds, this is achieved by concatenating feature vectors from 2, 4, 6, and 8 seconds ago to predict the response at time t. Taking the dot product of this concatenated feature space and a set of linear weights is functionally equivalent to convolving the original stimulus vector with a linear temporal kernel with nonzero entries for 1, 2, 3, and 4 time point lags. Before running the regression, each feature channel across the training matrix can be Z-scored. This can be performed to match features with the Z-scored fMRI response within each scan.

[0085] The 3,072 weights for each voxel can be estimated using L2 regularized linear regression. The regression procedure has a single free parameter that allows controlling the degree of regularization. This regularization factor can be found for each voxel for each subject by repeating the regression and cross-validation procedure 50 times. In each iteration, approximately one-fifth of the time points are removed from the model training dataset and reserved for validation. Model weights can then be estimated at the remaining time points for each of 10 possible regularization factors (logarithmic between 10 and 1,000). These weights are used to predict the response for the reserved time points and the R between the actual and predicted responses is calculated. 2 can be used to calculate. For each voxel, the regularization factor can be selected as the value that yielded the best performance, averaged across bootstraps, at the reserved time point. The 10,000 cortical voxels with the highest cross-validated performance were used for decoding.

[0086] The encoding model 310 can estimate a function R̂ that maps a word sequence S from its semantic features to a predicted brain response R̂(S). Assuming that the brain signal measured by fMRI, the Brain Oxygen Level Dependent (BOLD) signal, is affected by Gaussian additive noise, the likelihood of observing a brain response R given semantic features S is, on average, TIFF2025514868000003.tif1129 and covariance It can be modeled as a multivariate Gaussian distribution P(R|S) with

[0087] Previous studies have estimated the noise covariance Σ using the residuals between predicted and actual responses for the training dataset. However, this underestimates the actual noise covariance because the encoding model learns to predict some of the noise in the training dataset during model estimation. To get around this issue, a bootstrap procedure can be used to estimate the covariance Σ. Each story is excluded from the model training dataset, and the remaining data can be used to estimate the encoding model. The bootstrap noise covariance matrix for the excluded stories can be calculated using the residuals between the predicted and actual responses for the excluded stories. The covariance Σ can be estimated by averaging the bootstrap noise covariance matrix over the excluded stories.

[0088] The encoding model 310 is illustrated in FIG. 4 by encoding model 416. The continuation 414 undergoes variance, averaging, and convolution before being input to the encoding model 416, which variates the sequence features in the continuation 414 into L voxels of brain activity measurements (i.e., L candidate response values) and outputs a number of brain predictions 418 (i.e., a number of predicted brain responses). Each candidate response value can correspond to a measurement of a different voxel (i.e., a portion of the brain) or a different sensor. An example of a brain prediction is illustrated in brain prediction 418, which has n voxels similar to brain measurement 408.

[0089] In S322, the plurality of brain predictions having N*K brain predictions are compared to brain activity measurements made at the same acquisition time at which the predictions were made. For example, the brain activity measurements for the current period can be compared to the N*K brain predictions made for the current period. The brain activity measurements can include L brain measurements, or L voxel measurements including responses from L different voxels of the brain, with each voxel eliciting a different measurement. Each brain prediction is scored by taking a linear function (e.g., a dot product) of the difference between the L voxels (i.e., L candidate response values) of the brain activity measurements of the brain prediction and the true voxel responses (i.e., L voxel measurements) of the brain measurements, thereby determining an N*K score. The score can be determined by measuring the distance between the predicted response value and the true response value. Once the N*K scores are determined for the entire N*K plurality of brain predictions, the best N brain predictions, or N continuations, are identified from the scored N*K brain measurements.

[0090] The best N brain predictions are narrowed down from a list of N*K brain predictions. The best N*K brain predictions are narrowed down to prevent a combinatorial explosion of hypotheses. If K new continuation words are added to each of the N hypotheses and the N*K word sequences are not narrowed down, the K word sequences will grow exponentially with each cycle, making each new sequence completely intractable due to their large scale.

[0091] Step S322 is illustrated in Figure 4 by multiple brain predictions 418. Multiple brain predictions 418 are made for N*K successions, and the multiple brain predictions 418 are compared to the brain measurements 408 taken at the same acquisition time. Each prediction in the multiple brain predictions 418 and the brain measurements are then compared and given a score. The scores are then sorted and the best N predicted images are selected.

[0092] In S324, the N brain predictions, or word sequences, are fed back to the hypothesis beam 306, also represented at 420. Note that the size of the hypothesis beam 306 does not change as it always holds the N word sequences. This process is repeated until the end of the last brain measurement or stimulus 302. When the process is finished, the best hypothesis that contains the word sequence is selected from the list of N hypotheses in the hypothesis beam 306. The selected hypothesis becomes the final prediction for the stimulus 302.

[0093] III. Method A computer system can train an encoding model and use the encoding model to reconstruct language or word sequences from non-invasive brain activity recordings. The computer can train the encoding model through supervised learning, which trains the encoding model to predict brain activity measurements of stimuli. Once the encoding model is trained, the computer system can use the trained encoding model to translate brain measurements or images of a subject (user) into language.

[0094] A. Training the encoding model The computer system may receive a stimulus including a word sequence. Each word in the word sequence may be associated with one acquisition time (i.e., one time period). The computer system may annotate each word in the word sequence with a time label. For example, if the word sequence is "I have a dog," the annotating entity 106 may label the time of each word in the sequence. The computer system may annotate the word sequence with the time labels using speech recognition software or a human annotator.

[0095] The computer system can use a neural language model to transform a word sequence into a set of word embedding vectors. Each word in the word sequence can undergo quantitative feature transformations that transform each word in the word sequence into a word embedding vector. Each word embedding vector for a word can represent the semantics or syntax of the word.

[0096] The computer system may use the word embedding vectors to determine a final word embedding vector. The computer system may determine the final word embedding vector by downsampling the set of word embedding vectors to generate an averaged word embedding vector. For example, if the word sequence is "I have a dog", then the four word embedding vectors for the word sequence "I have a dog" are averaged to generate a single averaged word embedding vector. The computer system may apply a convolution kernel to the averaged word embedding vector and the previous averaged word embedding vector to generate the final word embedding vector. The convolution kernel is applied because a word sequence does not have just one immediate response. It has a responsiveness that evolves over time. A word sequence may affect approximately the next 8 seconds for brain activity measurements, or 4 periods for measurements using fMRI. For example, if the first word sequence is "I have a dog", then the sequence may not only affect the first brain activity image created by the brain imaging device 206, but may also affect the second, third, or fourth brain activity images.

[0097] The computer system can receive brain activity measurements including L brain measurements from a brain imaging device, where each of the L brain measurements corresponds to a measurement of a different voxel of the brain. The brain measurements can be taken using the brain imaging device while the subject is thinking, reading, or listening. The brain activity measurements can correspond to a time or period of acquisition.

[0098] Once the computer system receives the brain activity measurements and determines the final word embedding vector, it can determine a linear mapping of the final word embedding vector to a voxel space of the brain activity measurements, including the L brain activity measurements, to determine the encoding model. The linear mapping can be determined using linear regression to predict how the final word embedding vector will affect brain responses by fitting each feature of the vector to the voxel's brain activity measurements.

[0099] Brain activity measurements and final word embedding vectors can be determined for every period and compared to determine a linear mapping to train an encoding model. Brain activity measurements can be compared to the final word embedding vectors corresponding to that period.

[0100] B. Language reconstruction 5 shows a flow chart of a method for reconstructing language from non-invasive brain recordings by a computer system. The method includes steps 510 to 580.

[0101] Step 510 includes receiving updated N hypotheses, each consisting of the best predicted word sequences up to the current period. Each hypothesis may contain the same number of words. The N hypotheses are the N best combined word sequences in the current model. For example, if the actual words of the stimulus were "I walked a dog to the park outside during afternoon", then in the current period of language reconstruction, only "I walked a dog" may have been reconstructed, and the N hypotheses may include "I walked a cat", "I walked many dogs", "I ran with dogs", etc. Here, each hypothesis is the best sequence of words reconstructed in that period.

[0102] Step 520 involves predicting K continuation words for each hypothesis in the hypothesis beam (N hypotheses) by the neuro-language model. For example, if a hypothesis contains a word like "I ran", then the K possible continuation words are "over", "my", "fast", etc. The K continuation words are the best predicted words specific to the parent hypothesis. For example, if the two hypotheses are "I ran" and "I ate", then the continuation words for the first hypothesis will be "over", "my", etc., and the continuation words for the second hypothesis will be "burgers", "food", etc. Thus, the k continuation words are specific to each parent hypothesis. This step is similar to S318 in FIG. 3.

[0103] Step 530 includes combining the K continuation words with the N hypotheses to generate N*K continuations. A continuation may be a combination of hypotheses and continuation words. Each continuation is divided according to different sequences based on the predicted word times assigned by the word rate decoder, and each sequence is transformed into a set of word embedding vectors. The word embedding vectors in each sequence are averaged to generate one averaged word embedding vector per sequence. Then, a final word embedding vector for the current sequence is determined by assigning a different coefficient to each of the previous four sequences of the averaged word embedding using a predefined convolution kernel. The final word embedding vector can be used to predict brain responses. A detailed description of the transformation is provided in Section II.C.

[0104] Step 540 involves converting each continuation or final word embedding vector into L candidate values ​​of predicted brain responses. An encoding model trained before language reconstruction is used to map the continuations, i.e., final word embedding vectors, to predicted brain responses. The encoding model maps different features of the final word embedding vectors to candidate values ​​of predicted brain responses. There are L candidate values ​​of predicted brain responses, where each candidate response value corresponds to measurements of a different part or voxel of the brain, or a different sensor measuring a different part.

[0105] Step 550 includes receiving brain activity measurements, including L brain measurements by a brain imaging device, for the current time period or current sequence. Each of the L brain measurements corresponds to measurements of different voxels of the brain. The brain measurements can be taken with the brain imaging device while the subject is thinking, reading, or listening. The L brain measurements are taken per acquisition time or per sequence. Each of the L brain measurements is mapped to a predicted brain response for each of the L candidate values ​​computed using the final word embedding vector.

[0106] Step 560 includes comparing the brain activity measurements, including the L brain measurements, to predicted brain responses of the L candidate values ​​of the N*K predicted brain responses. Each predicted brain response in the N*K predicted brain responses is compared to the brain activity measurements made by the brain imaging device in step 550 and scored based on how similar it is to the brain activity measurement. Once the N*K scores for the N*K predicted brain responses or continuations are determined, the N*K scores are ranked according to how similar each continuation is. The continuations may be ranked according to the similarity scores.

[0107] Step 570 involves identifying the top N continuations based on the ranked N*K scores. If there are more brain measurements that the language reconstruction model needs to reconstruct, the selected top N continuations return to step 510 and start from the beginning of the reception of the N hypotheses. If there are no more measurements to reconstruct, proceed to step 580.

[0108] Step 580 includes outputting the word sequence corresponding to the brain activity measurement based on a selection of the best predicted brain response from among the top N sequences, where the selected predicted brain response becomes the final prediction made by the model predicting the best continuing language for the brain measurement performed by the brain imaging device.

[0109] IV. Model performance The performance of the language decoder can be analyzed through different assessments, tests and functions to determine how the sequence of decoded words matches the stimuli.

[0110] A. Decoder parameters The language decoder can be equipped with several parameters that can affect the performance of the model. The beam search algorithm can be parameterized by the beam width k. The encoding model can be parameterized by the number of context words provided when extracting the GPT embedding. The noise model can be parameterized by a shrinkage factor a that normalizes the covariance Σ. The language model parameters include the length of the input context, the kernel mass p and ratio r, and the set of possible output words.

[0111] Preliminary analysis shows that decoding performance can improve with beamwidth, but plateaus beyond k=200. Thus, a beamwidth of 200 sequences can be used for the analysis. All other parameters can be adjusted manually with grid search, but subjects must listen to a calibration story (e.g., "From Boyhood to Fatherhood" by Jonathan Ames on The Moth Radio Hour) separately from the training and test stories. The calibration story can be decoded using each configuration of parameters. The best-performing parameter values ​​can be verified and adjusted through qualitative analysis of the decoder predictions. The parameters that have the greatest impact on decoding performance are the kernel ratio r and the noise model shrinkage α. Setting r too small may reduce the linguistic coherence of the decoder, while setting r too large may reduce the semantic accuracy of the decoder. Setting α too small may overestimate the actual noise covariance, while setting α too large may underestimate the actual noise covariance. Both reduce the semantic accuracy of the decoder. The parameter values ​​used in this study can provide a default decoder configuration, but in practice can also be individually and continuously adjusted for each subject to improve performance.

[0112] To ensure that results generalize to new subjects and stimuli, all pilot analyses can be restricted to data collected while subjects are listening to the test stories. All pilot analyses on the test stories can be qualitative. The analysis pipeline can be frozen before reviewing the results of the remaining subjects, stimuli, and experiments.

[0113] B. Linguistic similarity index A wide range of automated metrics for assessing linguistic similarity can be used to compare the decoded word sequence to a reference word sequence. Word Error Rate (WER) calculates the number of edits (word insertions, deletions, or substitutions) required to change the predicted sequence to the reference sequence. Bilingual Evaluation Lower (BLEU) calculates the number of predicted n-grams that occur in the reference sequence (precision). A unigram variant, BLEU-1, can be used. Evaluation Metric for Transactions with Explicit Ordering (METEOR) combines the number of predicted unigrams that occur in the reference sequence (precision) with the number of reference unigrams that occur in the predicted sequence (recall), and takes into account synonyms and stemming using an external database. Bidirectional Encoder Representation with Transformer Score (BERTScore) uses a bidirectional transformer language model to represent each word in the predicted and reference sequences as a contextualized embedding and computes a matching score against the predicted and reference embeddings. BERT score can be used for all analyses where no linguistic similarity metric is specified.

[0114] In experiments of perceived speech, multiple talkers, and decoder resistance, the stimulus transcriptions can be used as reference sequences. In imagined speech experiments, subjects are asked to speak individual story segments out loud outside the scanner, and the speech is recorded and manually transcribed to provide reference sequences. In perceived movie experiments, audio descriptions can be manually transcribed to provide reference sequences. To compare word sequences decoded from different cortical regions (808 in Figure 8A), each sequence was scored using the other as a reference, and the scores were averaged (predicted similarity).

[0115] The predicted and reference words within a 20-second window can be scored for approximately every second of the stimulus (window similarity). The scores are averaged across the window to quantify how accurately the decoder predicted the complete stimulus (story similarity).

[0116] To estimate the upper bound of each metric, a perceived speech test story (e.g., "Where There is Smoke") can be translated into Mandarin by a professional translator. The translator was instructed to preserve all details of the story in the correct order. The story was then translated into English using a state-of-the-art machine translation system. The similarity between the words of the original story and the output of the machine translation system can be scored. These scores provide an upper bound on decoding performance, since modern machine translation systems can be trained with large amounts of paired data, and the Mandarin translation contains virtually the same information as the words of the original story.

[0117] To test whether decoder predictions can be used to identify the time instants of perceived speech, a post-hoc discriminant analysis can be performed using the similarity scores between the predicted and reference sequences. M reflects the similarity between the i-th predicted window and the j-th reference window. ij For each time point i, we can sort all reference windows by their similarity to the i-th predicted window and score time points by the percentile rank of the i-th reference window. The average percentile rank of the complete stimulus can be obtained by averaging the percentile ranks across time points.

[0118] To test whether the decoder predictions can be used to identify imagined speech scans, a post-hoc discriminant analysis using the similarity scores between the predicted and reference sequences can be performed. For each scan, the similarity scores between the decoder predictions and the five reference transcriptions can be normalized to probability. Top-1 accuracy can be calculated by assessing whether the decoder prediction for each scan is most similar to the correct transcription. A 100% top-1 accuracy for each subject can be observed. The cross-entropy for each scan can be calculated by taking the negative logarithm (base 2) of the probability of a correct transcription. It can be observed that the average cross-entropy is 0.25-0.82 bits. The cross-entropy of a perfect decoder is 0 bits, and the cross-entropy of a chance-level decoder is log2(5)=2.32 bits.

[0119] C. Statistical Tests To test the statistical significance of the word rate model (i.e., the word-time decoder), we calculated the linear correlation between the predicted and actual word rate vectors across the test stories and randomly shuffled 10-TR segments of the actual word rate vectors to generate 2,000 correlations. We could compare the observed linear correlations to a null distribution by one-sided permutation tests. p-values ​​were calculated as the proportion of shuffles with linear correlations equal to or greater than the observed linear correlation.

[0120] To test the statistical significance of the decoding scores, null sequences can be generated by sampling from the language model without using any brain data other than predicting the word times. The word rate model and the decoding scores can be evaluated separately, since the linguistic similarity metric used to calculate the decoding score is sensitive to the number of words in the predicted sequence. This test isolates the ability of the decoder to extract semantic information from the brain data by generating null sequences with the same word times as the predicted sequences. To generate null sequences, the same beam search procedure as the real language decoder can be followed. The null model maintains a beam of 10 candidate sequences and generates continuations from the kernels of the language model for each predicted word time. The only difference between the real decoder and the null model is that the null model randomly assigns a probability to each continuation, rather than ranking the continuations by the likelihood of the fMRI data. After iterating through all the predicted word times, the null model outputs the most likely candidate sequence. By repeating this process 200 times, 200 null sequences can be generated. This process is as similar as possible to a real decoding device without using brain data for word selection, so these sequences reflect the null hypothesis that the decoding device does not recover meaningful information about the stimuli from the brain data. Null sequences relative to the reference sequences can be scored to generate a null distribution of decoding scores. Observed decoding scores can be scored against this null distribution using a one-sided nonparametric test. p-values ​​were calculated as the proportion of null sequences with decoding scores equal to or greater than the observed decoding scores.

[0121] To verify that the null scores are not too low, the similarity scores between the reference sequence and the 200 null sequences can be compared to the similarity scores between the reference sequence and the other 62 story transcripts. The average similarity between the reference sequence and the null sequences is found to be higher than the average similarity between the reference sequence and the other story transcripts, indicating that the null scores are not too low.

[0122] To test the statistical significance of the post-hoc discriminant analysis, blocks of 10 rows of the similarity matrix M can be randomly shuffled before computing the mean percentile rank. 2,000 shuffles can be evaluated to obtain a null distribution of the mean percentile ranks. The observed mean percentile ranks can be compared to this null distribution using a one-sided permutation test. p-values ​​were calculated as the proportion of shuffles with a mean percentile rank equal to or greater than the observed mean percentile rank.

[0123] Unless otherwise stated, all tests were performed within each subject and then could be replicated across all subjects (n=7 for the between-subjects decoding analysis shown in 910 in Figure 9B, n=3 for all other analyses). All tests could be corrected for multiple comparisons using false discovery rates (FDR) where necessary. For all quantitative results, ranges across subjects were reported.

[0124] D. Motion Understanding Evaluation To assess understanding of the decoder predictions, an online behavioral experiment can be conducted to test whether subjects can answer multiple-choice questions about the stimulus stories posed by others using only the decoder predictions (Figures 13A and 13B). 80-second segments of perceived speech test stories can be selected on the basis that they are relatively independent in content. For each segment, four multiple-choice questions about the actual stimuli can be written without reference to the decoder predictions. To further ensure that the questions are not biased by the decoder predictions, the multiple-choice answers are written by a separate researcher who has never seen the decoder predictions.

[0125] The experiment was presented as a Qualtrics survey. One hundred online subjects (50 women, 49 men, and 1 non-binary) on Prolific were recruited and randomly assigned to experimental and control groups. For each segment, subjects in the experimental group were shown the words decoded from the subject, while subjects in the control group were shown the actual stimulus words. The words and corresponding multiple-choice questions from each segment were presented together on one page in the Qualtrics survey. The segments were presented in story order. The functionality of the "back" button was disabled, so subjects could not change their answers from a previous segment after viewing a new segment. The experimental protocol was approved by the Institutional Review Board of the University of Texas at Austin. Informed consent was obtained from all subjects.

[0126] E. Causes of Decryption Errors To test whether decoding performance is limited by the size of the training dataset, the decoder can be trained on different amounts of data. The decoding score appeared to increase linearly every time the size of the training dataset was doubled. To test whether the diminishing returns from adding training data is due to the decoder being trained on overlapping data samples, simulations were used to compare how the decoder behaves when trained on non-overlapping and overlapping data samples. Also used are real encoding models and real noise models to simulate the brain response to the 36 sessions of training stories. Non-overlapping samples of 3, 7, 11, and 15 sessions were obtained by taking sessions 1–3, 4–10, 11–21, and 22–36. Overlapping samples of 3, 7, 11, and 15 sessions were obtained by taking sessions 1–3, 1–7, 1–11, and 1–15. We were able to train decoders on these simulated datasets and found that the relationship between decoding score and number of training sessions was very similar for non-overlapping and overlapping datasets (Figure 11), suggesting that the diminishing returns to additional training observed is not due to the decoders being trained on overlapping sample data.

[0127] To test whether decoding performance depends on the high spatial resolution of fMRI, the fMRI data can be spatially smoothed by convolving each image with a 3-dimensional Gaussian kernel (Figure 18). Gaussian kernels with standard deviations of 1, 2, 3, 4, and 5 voxels can be tested, corresponding to full width at half maximum (FWHM) 6.1, 12.2, 18.4, 24.5, and 30.6 mm. An encoding model, noise model, and word rate model are estimated based on the spatially smoothed perceptual speech training data, and a decoding device is evaluated based on the spatially smoothed perceptual speech test data.

[0128] To test whether decoding performance was limited by noise in the test data, the signal-to-noise ratio of the test responses was artificially increased by averaging across repetitions of the test story.

[0129] To test whether decoding performance is limited by model misspecification, word-level decoding performance can be quantified by using a 300-dimensional GloVe embedding to represent words. A 10-s window centered on each stimulus word was considered. The maximum linear correlation between the stimulus word and the predicted words in the window was calculated. Then, for each of the 200 null sequences, the maximum linear correlation between the stimulus word and the null word in the window was calculated. The match score for a stimulus word was defined as the number of null sequences whose maximum correlation was lower than the maximum correlation of the predicted sequence. A match score above 100 indicates decoding performance higher than expected by chance, and a match score below 100 indicates decoding performance lower than expected by chance. The match scores were averaged over all occurrences of the word in the six test stories. Word-level match scores were compared to behavioral ratings of emotional valence (pleasantness), arousal (intensity of emotion), dominance (degree of control exerted), and concreteness (degree of sensory or motor experience). Each set of behavioral ratings was linearly rescaled to be between 0 and 1. Word-level match scores were also compared to the length of the word in the test dataset, the probability of the language model in the test dataset (corresponding to the information that the word conveys), the word frequency in the test dataset, and the word frequency in the training dataset.

[0130] F. Decoder Ablation When the word rate model (i.e., the word-time decoder) detects a new word, the language model proposes continuations using previously predicted words as an autoregressive context, and the encoding model ranks the continuations using the fMRI data. To understand the relative contributions of the autoregressive context and the fMRI data to decoding performance, the decoder can be evaluated on the perceived speech data in the absence of each component (Figure 19). A standard decoding approach was performed up to a cutoff point in the perceived speech test story. After the cutoff, the autoregressive context was reset or the fMRI data was removed. To reset the autoregressive context, all candidate sequences were discarded and the beam with an empty sequence was reinitialized. A standard decoding approach was then performed for the remainder of the scan. To remove the fMRI data, continuations for the remainder of the scan were assigned random probabilities rather than the encoding model probabilities.

[0131] G. Separate Encoding Model and Language Model Scores In effect, the decoder uses previously predicted words to predict the next word. This use of autoregressive context makes it difficult to attribute the source of error to any one component, as errors propagate between the encoding model and the language model. To isolate the errors introduced by each component, the decoder components can be evaluated separately on perceived speech test stories using actual rather than predicted stimulus words as context (Figure 20). At each word time t, the encoding model and the language model can be provided with the actual stimulus word and 100 randomly sampled distractor words.

[0132] To assess how well the encoding model can be used to decode a word at time t, the encoding model can be used to rank the actual stimulus word and 100 distractor words based on the likelihood of the recorded response. A separate encoding model score can be calculated based on the number of distractor words that received attention ranked lower than the actual word. Because the encoding model score is independent of the language model error and the autoregressive context, it provides an upper bound on how well each word can be decoded from the fMRI data.

[0133] To evaluate how well a word at time t can be generated using the language model, the language model can be used to rank the actual stimulus word and 100 distractor words based on their probability given previous stimulus words. A separate language model score can be calculated based on the number of distractor words that received attention ranked lower than the actual word. The language model score is independent of the errors in the encoding model and the autoregressive context, providing an upper bound on how well each word can be generated by the language model.

[0134] For both the separated encoding model and language model scores, 100 indicates perfect performance and 50 indicates chance-level performance. Separate encoding model and language scores were calculated for each word. For comparison with the 610 perfect decoding scores in Figure 6B, word-level scores were averaged over 20-second windows of stimuli.

[0135] H. Anatomical Alignment To test whether it was possible to estimate the decoder without training data from the target subjects, we used volume- and surface-based methods to anatomically align training data from individual source subjects to the volumetric space of the target subjects.

[0136] For volume matching, a linear map can be calculated from each source subject's volume space to the MNI template space using the get_mnixfm function in a computational function (e.g., pycortex). This map was applied to the recorded brain responses of each training story using the transform_to_mni function in pycortex. The transform_mni_to_subject function can then be used in pycortex to map the responses in MNI152 space to the target subject's volume space. Response time courses can be Z-scored for each voxel in the target subject's volume space.

[0137] For surface-based matching, the get_mri_surf2surf_matrix function in pycortex can be used to compute a map from each source subject's surface vertices to the target subject's surface vertices. This map was applied to the brain responses recorded in each training story. The target subject's surface vertices can then be mapped into the target subject's volumetric space using pycortex's line-nearest-neighbor scheme. Response time courses can be Z-scored for each voxel in the target subject's volumetric space.

[0138] A bootstrap procedure can be used to sample five sets of source subjects against the target subjects. Each source subject independently generated aligned responses against the target subjects. To estimate the encoding and word rate models, the aligned responses across source subjects can be averaged. For the word rate model, the target subject's phonetic network can be localized by anatomically aligning the source subject's phonetic network. To estimate the noise model Σ, the aligned responses from a single randomly sampled source subject can be used to compute the bootstrap noise covariance matrix for each left-out training story. The across-subject decoder was evaluated with the actual responses recorded from the target subjects.

[0139] V. Results Figures 6-10 are results illustrating embodiments of the present invention. Figures 6A, 6B, and 7 show decoding results of language generated using a language decoder. Figures 8A and 8B show decoding across different cortical regions. Figures 9A and 9B show applications of the language decoder and privacy implications. Figure 10 shows possible sources of decoding errors.

[0140] A. Results of the language decoder The language decoder was trained on three subjects, and each subject's decoder was assessed on separate single-trial brain responses recorded while the subject listened to a novel test story that was not used to train the model. Because the language decoder represents language using semantic features rather than motor or auditory features, the predictions of the language decoder need to capture the meaning of the stimuli. Results show that the decoded word sequences capture not only the meaning of the stimuli but often even exact words and phrases, demonstrating that detailed semantic information can be recovered from the BOLD signal (see 606 in Figure 6A and Figure 7).

[0141] To quantify decoding performance, several linguistic similarity metrics can be used to compare the decoded sequences and the real word sequences for one test story (1,839 words). Standard metrics such as Word Error Rate (WER), Bilingual Learning Environment (BLEU), and Explicitly Ordered Transactional Evaluation Metrics (METEOR) measure the number of words shared between two sequences. However, because different words can convey the same meaning, such as "we were busy" and "we had a lot of work," a new method that uses machine learning to quantify whether two sequences share meaning, Bidirectional Encoder Representation with Transformer Score (BERTScore), is used. Story decoding performance was notably higher than expected by chance for each metric, but especially for BERT score. (q(FDR)<0.05 for one-sided nonparametric tests; 608 in Figure 6; Figure 7). Most time points within the story (72-82%) had BERT scores significantly higher than expected by chance (610 in Figure 6B) and could be distinguished from other time points based on the BERT score similarity between the decoded words and the actual words (mean percentile rank = 0.85-0.91) (612 in Figure 6B; 802 in Figure 8A). Whether the decoded words captured the original meaning of the story was also tested using a behavioral experiment, which showed that subjects who only read the decoded words were able to answer 9 out of 16 reading comprehension questions.

[0142] 6A and 6B show the results of an analysis of the predicted segments produced by the language decoder. The results of different statistical tests performed on the predicted segments are outlined and the results of the tests are analyzed.

[0143] For 602 model training, brain oxygen level-dependent (BOLD) responses were recorded using fMRI while subjects listened to verbally told stories. A language model (LM) was used to extract quantitative features of each stimulus word. An encoding model (EM) was estimated to predict the BOLD response from word features.

[0144] In language reconstruction 604, to reconstruct language from novel brain recordings, the language decoder can maintain a set of candidate word sequences. When new words are detected, a language model (LM) proposes continuations for each sequence, and an encoding model scores the likelihood of the recorded brain response under each continuation. The most likely continuations can be retained.

[0145] In 606, the decoder was evaluated on single-trial brain responses recorded while subjects listened to test stories that were not used in model training. Segments of the four test stories can be viewed side-by-side with the decoder predictions for one subject. Examples were manually selected and annotated to demonstrate typical language decoder behavior. The language decoder accurately reproduces some words and phrases and captures the gist of many more words and phrases.

[0146] In 608, the language decoder's predictions of the test stories were significantly more similar to the actual stimulus words than expected by chance, based on a wide range of linguistic similarity indices (* indicates q(FDR)<0.05 for one-tailed nonparametric tests for all subjects). To compare across indices, results are displayed as standard deviations from the mean of a null distribution (see Methods). Boxes indicate the interquartile range of the null distribution. Whiskers indicate the 5th and 95th percentiles.

[0147] At most time points, 610, decoding scores were significantly higher than expected by chance under the BERTScore index (one-sided nonparametric test q(FDR)<0.05).

[0148] In Figure 612, discrimination accuracy for one subject is shown. The color (i,j) reflects the similarity between the i th second of the prediction and the j th second of the actual stimulus. Discrimination accuracy was significantly higher than expected by chance (one-tailed permutation test q<0.05).

[0149] Figure 7 shows the linguistic similarity scores between the decoded words and the actual word sequences using several linguistic similarity metrics. Standard metrics such as Word Error Rate (WER), Bilingual Unequal Equation (BLEU), Explicitly Ordered Transactional Evaluation Metric (METEOR), and Bidirectional Encoder Representation with Transformer Score (BERTScore) can be shown in Figure 7.

[0150] In Figure 7, decoder predictions of perceived stories were compared to actual stimulus words using a wide range of linguistic similarity metrics. The lower bound of each metric was calculated by scoring the average similarity between the actual stimulus words and 200 null sequences generated from a language model without using brain data. The upper bound of each metric was calculated by manually translating the actual stimulus words into Mandarin Chinese, automatically translating the words into English using a state-of-the-art machine translation system, and scoring the similarity between the actual stimulus words and the machine translation system's output. On the BERT score metric, the decoder, trained with much fewer paired data and using much noisier inputs, performed about 20% as well as the machine translation system on the floor.

[0151] B. Decoding across cortical regions The decoding results shown in Figures 6A and 6B used responses from multiple cortical regions to achieve good performance. The decoder can be used to understand how language is represented within each of these regions. Previous studies have shown that large parts of the cortex are active during language processing, but it is unclear which regions represent language at the granularity of words and phrases, which regions are consistently involved in language processing, and whether different regions encode complementary or redundant language representations. To answer these questions, we partitioned the brain data into three large-scale cortical regions previously shown to be active during language processing (speech network, parieto-temporal-occipital association regions, and prefrontal cortex), each of which can be decoded separately from regions of the hemisphere. (802 in Figure 8A; 1402 in Figure 14).

[0152] To test whether regions encode semantic information at the granularity of words and phrases, we can evaluate the decoder's predictions from regions using similarity indices across languages. Previous studies have decoded semantic features from the BOLD responses of different regions, but the distributed nature of semantic features and the low temporal resolution of the BOLD signal make it difficult to assess whether regions represent fine-grained words or coarser-grained categories. Because the decoder produces interpretable word sequences, we have direct access to how accurately each region represents the stimulus words (806 in Figure 8A). Under the WER and BERTScore indices, the decoder's predictions were significantly more similar to the actual stimulus words than chance predictions in all regions (q(FDR)<0.05 in one-sided nonparametric tests). Under the BLEU and METEOR indices, the decoder's predictions were significantly more similar to the actual stimulus words than chance predictions in all regions except the right hemisphere speech network (q(FDR)<0.05 in one-sided nonparametric tests). These results indicate that multiple cortical regions represent language at the granularity of individual words and phrases.

[0153] Previous analyses quantified how well a region represented a stimulus as a whole, but did not identify whether the region was consistently engaged throughout the entire stimulus or active only at specific times. To identify regions that are consistently engaged in language processing, the percentage of time points saliently decoded from each region can be calculated. Results show that most of the time points saliently decoded from the whole brain could be individually decoded from the association regions (80–86%) and prefrontal regions (46–77%) (806 in Figure 8A; 1404 in Figure 14), suggesting that these regions consistently represent the meaning of words and phrases in language. Notably, only 28–59% of the time points saliently decoded from the whole brain could be decoded from the phonetic network. This is likely a result of the decoding framework. Although the phonetic network is known to be consistently engaged in language processing, the phonetic network tends to represent low-level articulatory and acoustic features, whereas the language decoder operates based on high-level semantic features of entire word sequences.

[0154] Finally, we can assess the relationship between linguistic representations encoded in different regions. One reason for successful decoding from multiple regions is that different regions encode complementary representations in a modular organization, such as different parts of speech. If this is the case, different aspects of the stimulus may be decodable from individual regions, but the complete stimulus should only be decodable from the whole brain. Alternatively, different regions may encode redundant representations of the complete stimulus. If this is the case, the same information may be decoded separately from multiple individual regions. To distinguish between these possibilities, we directly compared decoded word sequences across regions and hemispheres and found that the similarity between each pair of predictions was significantly higher than expected by chance (q(FDR) < 0.05 in one-tailed nonparametric tests; 808 in Figure 8A). This suggests that different cortical regions encode redundant word-level linguistic representations. However, it is possible that the same word may be encoded using different features in different regions.

[0155] This result indicates that word sequences that can be decoded from the whole brain can also be consistently decoded from multiple individual regions (810 in Figure 8B). The practical implication of this redundant coding is that future brain-computer interfaces may be able to selectively record from the most accessible or intact regions and still achieve good performance.

[0156] Figures 8A and 8B show the results of an analysis of how language is represented across different cortical networks. We show decoder predictions from different cortical language networks and analyze the predictions made by the different cortical language networks.

[0157] In Figure 802, partitioned brain data of three large-scale cortical networks are shown. The brain data was partitioned into a classically localized language network, a parieto-temporal-occipital association network, and a prefrontal network.

[0158] In 804, decoder predictions from each region in each hemisphere were significantly more similar to the actual stimulus words than chance predictions in most metrics (* indicates q(FDR)<0.05 for one-sided nonparametric tests for all subjects). Error bars indicate standard error of the mean (n=3 subjects). Boxes indicate the interquartile range of the null distribution. Whiskers indicate the 5th and 95th percentiles.

[0159] In 806, we decode the time course of performance from each region for one subject. Horizontal lines indicate when decoding performance was significantly higher than expected by chance, as measured by the BERTScore metric (one-sided nonparametric test q(FDR)<0.05). Most of the time points significantly decoded from the whole brain were also significantly decoded from association regions and prefrontal cortex.

[0160] In 808, the predictions of the decoder were compared across regions. Word sequences decoded from each pair of regions were significantly more similar than expected by chance (2-tailed nonparametric tests q(FDR)<0.05).

[0161] In Fig. 810, decoder predictions from each network in each hemisphere are shown. Predictions were similar across networks and captured the meaning of the stimuli.

[0162] A. Decryption Uses and Privacy Impact In the previous analysis, a language decoder was trained and tested on brain responses to perceived speech. Next, to demonstrate the potential range of applications of a semantic language decoder, we can evaluate whether a language decoder trained on brain responses to perceived speech can be used to decode brain responses to other tasks.

[0163] 1. Imagined speech decoding A key task of the brain-computer interface is to decode covert imagined speech in the absence of external stimuli. To test whether the language decoder can be used to decode imagined speech, subjects imagined telling five 1-min stories while being recorded with fMRI and also told the same stories independently outside the scanner to provide a reference transcription. For each 1-min scan, we could accurately identify the story the subject had imagined by decoding the scan, normalizing the similarity scores between the decoder predictions and the reference transcripts to probability, and selecting the most likely transcript (identification accuracy 100%; 902 in Figure 9A; 1204 in Figure 12). Across the stories, the decoder predictions were significantly more similar to the corresponding transcripts than expected by chance (p<0.05 with a one-tailed nonparametric test). Qualitative analysis shows that the decoder is able to recover the meaning of the imagined stimuli (904 in Figure 9A; Supplementary Table 2).

[0164] In order for the decoding device to transition between tasks, the target task may share representations with the training task. Because the encoding model is trained to predict how the subject's brain will respond to perceived speech, the explicit goal of the decoding device is to generate words that elicit the recorded brain response when the subject hears them. The decoding device successfully transitions to imagined speech because the semantic representations that are activated when the subject imagines a story are similar to the semantic representations that would be activated if the subject heard the story. Nevertheless, the decoding performance of imagined speech is lower than that of perceived speech (1502 in Figure 15), which is consistent with previous findings that speech production and speech perception involve partially overlapping brain regions. More accurate decoding of imagined speech may be achieved by replacing the encoding model trained on perceived speech data with an encoding model trained on attempted or imagined speech data. This gives the decoding device an explicit goal of generating words that elicit the recorded brain response when the subject imagines them.

[0165] 2. Cross-modal Decoding Semantic representations are also shared between language perception and a wide range of other perceptual and conceptual processes, suggesting that unlike previous language decoders that primarily used motor or auditory signals, a language decoder may be able to reconstruct linguistic descriptions from brain responses to non-linguistic tasks. To test this, subjects watched four short films recorded with fMRI without sound, and the recorded responses were decoded using a semantic language decoder. The decoded word sequences can be compared to the linguistic descriptions of the films for blind subjects. The results show that they are significantly more similar than expected by chance (q(FDR)<0.05 in one-tailed non-parametric tests; 1502 in Figure 15). Qualitatively, the decoded sequences accurately described events from the films (906 in Figure 9A). This suggests that a single semantic decoder trained during language perception may be used to decode a wide range of semantic tasks.

[0166] 3. Effects of attention on decoding Because semantic representations are modulated by attention, the semantic decoder needs to selectively reconstruct attended stimuli. To test the effect of attention on decoding, subjects listened to two repetitions of multi-talker stimuli constructed by temporally overlapping a set of stories told by a female and a male talker. In each presentation, subjects were instructed to attend to a different talker. The predictions of the decoder were significantly more similar to attended stories than to unattended stories (q(FDR) < 0.05 in one-tailed paired t-tests across subjects), indicating that the decoder selectively reconstructs attended stimuli (908 in Figure 9B; 1504 in Figure 15). These results suggest that the semantic decoder can operate well in complex environments with multiple information sources. Moreover, these results indicate that subjects consciously control the output of the decoder, suggesting that the semantic decoder can reconstruct only those to which subjects actively attend.

[0167] 4. Privacy Impact An important ethical consideration with semantic decoding is the potential for violating mental privacy. To test whether it is possible to train a decoder without human cooperation, perceived speech from each subject was decoded using a decoder trained on data from the other subjects. In this analysis, seven subjects collected data by listening to 5 h of stories. These data were anatomically aligned across subjects using volume- and surface-based methods. The decoder trained on the between-subject data (Figure 16) performed slightly above chance and was significantly worse than the decoder trained on the within-subject data (q(FDR)<0.05 in a two-tailed t-test). This suggests that training the decoder still requires subject cooperation (910 in Figure 9B; 1508 in Figure 15).

[0168] To test whether a decoding device trained with human assistance could later be resisted consciously, subjects silently performed three cognitive tasks, namely arithmetic ("count in multiples of seven"), semantic memory ("name and imagine animals"), and imagined speech ("tell different stories"), while listening to narrative segments. Performing the semantic memory and imagined speech tasks significantly degraded decoding performance compared to the baseline of passive listening in each cortical region (q(FDR) < 0.05 in one-tailed paired t-tests across subjects). This indicates that in adversarial scenarios, semantic decoding can be consciously resisted, and that this resistance cannot be overcome by focusing the decoding device exclusively on specific brain regions (912 in Figure 9B; 1508 in Figure 15).

[0169] In Figures 9A and 9B, some practical considerations are presented for how a language decoder could be deployed as a brain-computer interface. Results from imagining sounds and watching a short film are described and analyzed.

[0170] In 902, to test whether the language decoder could transfer to imagined speech, subjects imagined speaking five 1-minute test stories twice. Single-trial brain responses were decoded and compared to reference transcripts recorded independently from the same subjects. Identification accuracy is shown for one subject. Each row corresponds to a scan, and colors reflect the similarity between the decoder predictions and all five reference transcripts, normalized to probability. For each scan, the decoder prediction was most similar to the reference transcript of the correct story (100% identification accuracy).

[0171] At 904, the reference transcript is displayed alongside the decoder's predictions for three imagined stories about a single subject.

[0172] In 906, to test whether the language decoder could be transferred to different modalities, subjects watched four silent short films. Single-trial brain responses were decoded using a language decoder. Decoder predictions were significantly associated with the films (q(FDR)<0.05 in one-sided nonparametric tests) and often accurately described film events. Frames from two scenes were presented side-by-side with the decoder predictions for one subject.

[0173] In 908, to test whether the decoding mechanism is modulated by attention, subjects listened to multi-talker stimuli in which one talker told a female talker and the other talker told a male talker while attending to either one. The predictions of the decoding mechanism were significantly more similar to the attended stories than to the unattended stories (* indicates q(FDR)<0.05 in a one-tailed paired t-test across n=3 subjects; t(2)=6.15 for female talkers, t(2)=6.45 for male talkers). Markers indicate individual subjects.

[0174] In 910, to test whether decoding could be successful without training data from a particular subject, we trained a decoder on brain responses from five sets of other subjects (indicated by markers) matched using volume- and surface-based methods. The between-subject decoder performed slightly above chance and significantly worse than the within-subject decoder (* indicates q(FDR)<0.05 in a two-tailed t-test), suggesting that the within-subject training data is important.

[0175] In 912, to test whether subjects could consciously resist decoding, subjects silently performed three resistance strategies: counting, naming animals, and telling different stories. Decoding performance in each condition was compared to the passive listening condition (* indicates q(FDR)<0.05 across n=3 subjects). Naming animals (t(2)=6.95 for whole brain, t(2)=4.93 for speech network, t(2)=6.93 for association regions, t(2)=4.70 for prefrontal cortex) and telling different stories (t(2)=4.79 for whole brain, t(2)=4.25 for speech network, t(2)=3.75 for association regions, t(2)=5.73 for prefrontal cortex) significantly impaired performance in each cortical area, indicating that they may impede decoding. Markers indicate individual subjects. Because story decoding scores depend on stimulus length, different experiments (perceived speech, imagined speech, perceived film, multiple talkers, decoding device resistance) cannot be compared based on story decoding scores.

[0176] 5. Causes of decryption errors To identify potential avenues of improvement, decoding errors during speech perception, which reflect limitations in the fMRI recordings, the model, or both, can be assessed (1002 in Figure 10).

[0177] BOLD fMRI recordings typically have a low signal-to-noise ratio (SNR). During model estimation, the effect of noise in the training data can be mitigated by increasing the size of the dataset. To assess whether decoding performance is limited by the size of the training dataset, we can train the language decoder using different amounts of data. Although decoding scores were significantly higher than expected by chance in a single session of training data, more training data was required to consistently decode different parts of the test stories (Figure 17). Decoding scores appeared to increase by the same amount each time the size of the training dataset was doubled (1004 in Figure 10). This suggests that training with more data improves decoding performance, albeit with diminishing gains for each successive scanning session.

[0178] A low SNR in the test data may also limit the amount of information that can be decoded. To assess whether future improvements in the SNR of single-trial fMRI would improve decoding performance, we artificially increase the SNR by averaging brain responses collected during different repetitions of the test story. Decoding performance increases slightly with the number of responses averaged (1006 in Figure 10), suggesting that some component of the decoding error reflects noise in the test data.

[0179] Another limitation of fMRI is that current scanners are too large and expensive for most practical decoding applications. Portable techniques such as functional near-infrared spectroscopy (fNIRS) measure the same hemodynamics as fMRI, albeit at a lower spatial resolution. To test whether decoding relies on the high spatial resolution of fMRI, we were able to smooth the fMRI data down to the estimated spatial resolution of current fNIRS systems and found that approximately 50% of the stimulation time points could still be decoded (Figure 18). This suggests that the decoding approach may eventually be applicable to portable systems as well.

[0180] Finally, to assess whether decoding performance is limited by model misspecification, such as using suboptimal features to represent the linguistic stimuli, we can test whether decoding errors follow systematic patterns. We can score how well individual words were decoded across the six test stories (see Methods) and compare the scores on behavioral word ratings and dataset statistics. If decoding errors were caused solely by noise in the test data, all words should be similarly affected. However, we found that decoding performance was significantly correlated with behavioral ratings of word concreteness (rank correlation ρ = 0.15-0.28, q(FDR) < 0.05), suggesting that decoders are poor at recovering words with certain semantic properties (1008 in Figure 10). Notably, decoding performance was not significantly correlated with word frequency in the training stimuli, suggesting that model misspecification is not primarily caused by noise in the training data (1010 in Figure 10).

[0181] The results indicate that model misspecification is a major source of decoding error, apart from random noise in the training and test data. Assessing how different components of the decoder contribute to this misspecification, we find that the decoder continually relies on the encoding model to achieve good performance (Figure 19), and moments of poor decoding tend to reflect errors in the encoding model (Figure 20). Computational advances that reduce encoding model misspecification, such as the development of better semantic feature extractors, can be expected to result in notable improvements in decoding performance.

[0182] Figure 10 shows the results of the analysis of decoding errors. Potential factors limiting decoding performance were tested to identify directions for improvement. By taking into account the bias and fMRI noise of the model in Figure 10 and comparing, better predictions were come up with.

[0183] 1002 lists three possible causes of decoding errors: Decoding errors can arise from misspecification of features in the language model, insufficient data to train the encoding model, or noise from fMRI to measure brain activity.

[0184] In 1004, we trained a decoder on different amounts of data to test whether decoding performance was limited by the size of the training data set. The decoding score appeared to increase by the same amount each time the size of the training data set was doubled.

[0185] In 1006, to test whether decoding performance was limited by noise in the test data, the signal-to-noise ratio of the test responses was artificially increased by averaging across repetitions of the test story. Decoding performance improved slightly depending on the number of responses averaged.

[0186] In 1008, word-level decoding scores were compared to behavioral assessments and dataset statistics to test whether decoding performance was limited by model misspecification (* indicates q(FDR)<0.05 in two-tailed permutation tests for all subjects). Markers indicate individual subjects.

[0187] In 1010, decoding performance was significantly correlated with word concreteness, suggesting that model misspecification contributes to decoding errors, but was uncorrelated with word frequency in the training stimuli. This suggests that model misspecification is not caused by noise in the training data. For all results, black lines indicate the mean across subjects and error bars indicate the standard error of the mean (n=3).

[0188] VI. Supplementary Results 11-20 are enhancements of the experimental data shown in FIGS. 6-10.

[0189] FIG. 11 may show the performance of two language decoder components that interface with the fMRI data: the encoding model and the word rate model.

[0190] In 1102, the encoding model was evaluated by predicting brain responses to perceived audio test stories and calculating linear correlations between the predicted and actual single-trial responses. Correlations for subject S3 were projected onto a cortical flat map. The encoding model was successful in predicting brain responses in most cortical areas except for primary sensory and motor areas.

[0191] In 1104, the encoding models were trained on different amounts of data. To summarize the performance of the encoding models across the cortex, we averaged the correlations across the 10,000 voxels used for decoding. The performance of the encoding models improved depending on the amount of training data collected from each subject.

[0192] In 1106, the encoding model was tested on brain responses averaged across different repetitions of a perceived speech test story to artificially increase the signal-to-noise ratio (SNR). The performance of the encoding model improved depending on the number of responses averaged.

[0193] In 1108, word rate models were trained with different amounts of data. The word rate models were evaluated by predicting word rates for test stories and computing the linear correlation between the predicted and actual word rate vectors. The performance of the word rate models improved slightly depending on the amount of training data collected from each subject.

[0194] 1110, for brain responses to perceived speech, word-rate models fitted to auditory cortex significantly outperformed word-rate models fitted to prefrontal speech-producing regions or randomly sampled voxels (* indicates q(FDR)<0.05 in a two-tailed paired t-test across n=3 subjects).

[0195] In 1112, there was no significant difference in the performance of word rate models fitted to different cortical regions for brain responses to imagined speech. For all results, the black line indicates the mean across subjects and the error bars indicate the standard error of the mean (n=3).

[0196] Perceived speech discrimination performance and imagined speech discrimination performance may be shown in Figure 12. For subjects S1 and S2, a language decoder was trained based on fMRI responses recorded while the subjects listened to the story.

[0197] In 1202, the decoder was assessed on single-trial fMRI responses recorded while subjects listened to perceived speech test stories. The color of (i,j) reflects the degree of similarity of the BERT score between the i-th second of the decoder prediction and the j-th second of the actual stimulus. Identification accuracy was significantly higher than expected by chance (q<0.05 in a one-tailed permutation test). The corresponding results for subject S3 are shown in Figure 6 in the main text.

[0198] In 1204, the decoder was assessed on single-trial fMRI responses recorded while subjects imagined telling five 1-min test stories twice. The decoder predictions were compared to reference transcripts recorded independently from the same subjects. Each row corresponds to a scan, and colors reflect the degree of similarity between the decoder predictions and all five reference transcripts, normalized to probability. For each scan, the decoder prediction was most similar to the reference transcript of the correct story (100% discrimination accuracy). The corresponding results for subject S3 are shown in 902 of Figure 9A of the main text.

[0199] 13A and 13B may illustrate the performance evaluation of the language decoder predictions. 80-second segments from the perceived speech test stories were selected. For each segment, four multiple-choice questions were created based on the actual stimulus words without seeing the decoder predictions. One hundred subjects were recruited for the online behavioral experiment and randomly assigned to the experimental and control groups. For each segment, subjects in the experimental group answered the questions after reading the decoded words from subject S3, while subjects in the control group answered the questions after reading the actual stimulus words.

[0200] In 1302, the experimental group's scores were significantly higher than expected by chance on 9 of 16 questions (* indicates q(FDR)<0.05 in a two-tailed binomial test).

[0201] At 1304, the decoded words of the segment and the actual stimulus words are shown.

[0202] In 1308, multiple choice questions cover different aspects of the stimulus story.

[0203] FIG. 14 may show decoding performance across cortical regions.

[0204] In 1402, the cortical regions of subjects S1 and S2 are shown. The brain data used for decoding (colored regions) were partitioned into the auditory network, parieto-temporal-occipital association regions, and the prefrontal cortex (PFC).

[0205] In 1404 we show the time course of decoding performance for the perceived speech test stories from each domain. Horizontal lines indicate when the decoder's predictions were significantly more similar to the actual stimulus words than would be expected by chance, as measured by the BERTScore index (one-sided nonparametric tests with q(FDR)<0.05). The corresponding results for subject S3 are shown in Figure 8A in the main text at 802 and 804.

[0206] Figure 15 shows a comparison of decoding performance across experiments. Decoder predictions from different experiments were compared based on the proportion of prominently decoded time points with a BERTScore metric (q(FDR)<0.05). The proportion of prominently decoded time points was used because it is independent of the stimulus length.

[0207] In 1502, the decoder was successful in recovering 72-82% of the time points in the perceived speech, 33-73% of the time points in the imagined speech, and 16-45% of the time points in the perceived movie.

[0208] In 1504, during multiple-talker stimulation, the decoder was successful in recovering 36–73% of the time points where the female talker spoke when the subject was attending to the female talker, 0–1% of the time points where the female talker spoke when the subject was attending to the male talker, 60–76% of the time points where the male talker spoke when the subject was attending to the male talker, and 0–3% of the time points where the male talker spoke when the subject was attending to the female talker.

[0209] In 1506, during the perceived story, the within-subject decoder succeeded in recovering 65–82% of the time points, the volume-based between-subject decoder succeeded in recovering 1–2% of the time points, and the surface-based between-subject decoder succeeded in recovering 1–5% of the time points.

[0210] In 1508, during the perceived story, the within-subject decoder was successful in recovering 52–57% of the time when subjects passively listened, 4–50% of the time when subjects resisted by counting in multiples of seven, 0–3% of the time when subjects resisted by naming animals, and 1–26% of the time when subjects resisted by imagining a different story.

[0211] Figure 16 shows the performance of the between-subject encoding and word rate models. For each subject, the encoding and word rate models were trained based on brain responses from five other subjects (indicated by markers) matched using volume- and surface-based methods. The models were evaluated on within-subject single-trial responses to the perceived speech test stories.

[0212] In 1602, the between-subject encoding model performed significantly worse than the within-subject encoding model (* indicates q(FDR)<0.05 in a two-tailed t-test).

[0213] In 1604, the between-subject word rate model performed significantly worse than the within-subject word rate model (* indicates q(FDR)<0.05 in a two-tailed t-test).

[0214] Figure 17 shows the decoding performance as a function of training data. Decoders were trained with different amounts of data and evaluated on perceived speech test stories.

[0215] In 1702, the proportion of saliently decoded time points increased with the amount of training data collected from each subject, but plateaued after 7 scanning sessions (7.5 hours) and did not increase substantially until 15 sessions (16 hours). The substantial increase up to 7 scanning sessions suggests that although the decoder can recover certain semantic concepts after training on a small amount of data, more training data is needed to achieve consistently good performance across test stories.

[0216] In 1704, the mean discrimination percentile rank increased with the amount of training data collected from each subject, but plateaued after 7 scanning sessions (7.5 h) and did not increase substantially until 15 sessions (16 h). For all results, the black line indicates the mean across subjects and the error bars indicate the standard error of the mean (n=3).

[0217] Figure 18 may show decoding performance at lower spatial resolution. Although fMRI offers high spatial resolution, current MRI scanners are too large and expensive for most practical decoding device applications. Portable alternatives such as functional near-infrared spectroscopy (fNIRS) measure the same hemodynamics as fMRI, albeit at lower spatial resolution. To simulate how the decoding device would perform at lower spatial resolution, the fMRI data were spatially smoothed using Gaussian kernels with standard deviations of 1, 2, 3, 4, and 5 voxels. This corresponds to full width at half maximum (FWHM) 6.1, 12.2, 18.4, 24.5, and 30.6 mm. An encoding model, noise model, and word rate model were estimated based on the spatially smoothed training data, and the decoding device was evaluated on spatially smoothed responses to perceived speech test stories.

[0218] In 1802, each subject's fMRI images were spatially smoothed using progressively larger Gaussian kernels.

[0219] In 1804, story similarity decreased as the data was spatially smoothed, but remained high at moderate levels of smoothing.

[0220] In 1806, the proportion of significantly decoded time points decreased as the data were spatially smoothed, but remained high at moderate levels of smoothing.

[0221] In 1808, the predictive performance of the encoding model improved as the data was spatially smoothed, indicating that decoding performance and encoding model performance are not fully coupled. Spatial smoothing reduces information, making stimuli harder to decode, but it also reduces noise, making responses easier to predict. For all results, the black line shows the average across subjects, and the error bars show the standard error of the mean (n=3). The gray dashed line shows the estimated spatial resolution of the current portable system. These results indicate that approximately 50% of the stimulus time points can still be decoded with the estimated spatial resolution of the current portable system, providing a benchmark for how much the portable system needs to improve to reach different levels of decoding performance.

[0222] Figure 19 may illustrate the ablation of the decoder. To decode new words, the decoder uses both the autoregressive context (i.e., previously decoded words) and the fMRI data. To understand the relative contributions of the autoregressive context and the fMRI data, the decoder was evaluated in the absence of each component. A standard decoding approach was performed up to a cutoff point in the perceived speech test story. After the cutoff, the autoregressive context was reset or the fMRI data was removed. To reset the autoregressive context, all candidate sequences were discarded and the beam was reinitialized with an empty sequence. A standard decoding approach was then performed for the remainder of the scan. To remove the fMRI data, random probabilities were assigned to continuations rather than the encoding model probabilities for the remainder of the scan.

[0223] In 1902, the cutoff point was defined as 300 s after stimulus onset for one subject. When the autoregressive context was reset, decoding performance decreased but quickly recovered. When the fMRI data were removed, decoding performance rapidly decreased to chance levels. The grey shaded region indicates the 5th to 95th percentiles of the null distribution.

[0224] In 1904, ablation was repeated at the cutoff point every 50 seconds of stimulation. The difference in performance between the original and ablated decoders was averaged across cutoff points and subjects to provide a profile of how decoding performance changes after each component is ablated. Blue and purple shaded areas indicate standard error of the mean (n=27 trials). These results indicate that the decoder continually relies on the encoding model and fMRI data to achieve good performance and does not require a good initial context. In these figures, each time point is scored based on a 20 second window ending at that time point, whereas in all other figures, each time point is scored based on a 20 second window centered at that time point. This shifted indexing scheme highlights how decoding performance changes after the cutoff. Grey dashed lines indicate cutoff points

[0225] Figure 20 shows the separated encoding model and language model scores. The encoding model and language model were evaluated separately on the perceived speech test stories to separate their contributions to decoding errors. At each word time t, the encoding model and language model were provided with the actual stimulus word and 100 randomly sampled distractor words. The encoding model ranks the words based on their likelihood in the recorded fMRI response, and the language model ranks the words based on their probability given the previous stimulus word. The separated encoding model and language model scores were calculated based on the number of distractor words that were ranked below the actual words. A score of 100 indicates perfect performance, and a score of 50 indicates chance-level performance. To compare with the 610 perfect decoding scores in Figure 1, the word-level encoding model and language model scores were averaged over 20-second windows of the stimuli.

[0226] In 2002, encoding model scores were significantly correlated with full decoding scores (linear correlation r = 0.22 to 0.58, p < 0.05), suggesting that many of the poorly decoded time points in Figure 6B are inherently more difficult to decode, even using the encoding model.

[0227] In 2004, language model scores did not significantly correlate with full decoding scores.

[0228] In 2006, for each word, we compared the encoding model scores from 10 randomly sampled sets of distractors to chance level 50 using two-tailed t-tests. Most stimulus words with statistically significant encoding model scores across the brain (q(FDR)<0.05, two-tailed t-tests) also had statistically significant encoding model scores in the speech network (80–87%), association regions (88–92%), and prefrontal cortex (82–85%). This suggests that the 806 results in Figure 8A are not primarily due to language models. Word-level encoding model scores were significantly correlated across each pair of regions (q(FDR)<0.05, two-tailed permutation tests), suggesting that the 808 results in Figure 8A are not primarily due to language models.

[0229] In 2008, to characterize bias in the decoder component, word-level encoding and language model scores were correlated against the word properties tested in 1008 of Figure 10. (* indicates q(FDR)<0.05 in two-tailed permutation tests for all subjects). The encoding and language models were biased for some word properties. These effects may have balanced out in the full decoder, resulting in no observed correlation between word properties and full decoding scores (1008 of Figure 10).

[0230] VII. Computer Systems Any of the computer systems referred to herein may utilize any suitable number of subsystems. Examples of such subsystems are shown in FIG. 21 of computer system 10. In some embodiments, a computer system includes a single computer device and the subsystems may be components of the computer device. In other embodiments, a computer system may include multiple computer devices, each of which is a subsystem with internal components. Computer systems may include desktop and laptop computers, tablets, mobile phones, and other mobile devices.

[0231] The subsystems shown in FIG. 21 are interconnected via a system bus 75. Additional subsystems are shown, such as a printer 74 in conjunction with a display adapter 82, a keyboard 78, a storage device 79, and a monitor 76 (e.g., a display screen such as LEDs). Peripherals and input / output (I / O) devices coupled to the I / O controller 71 can be connected to the computer system by any means known in the art, such as an input / output (I / O) port 77 (e.g., USB, FireWire). For example, the I / O port 77 or an external interface 81 (e.g., Ethernet, Wi-Fi, etc.) can be used to connect the computer system 10 to a wide area network, such as the Internet, a mouse input device, or a scanner. The interconnection via the system bus 75 allows the central processor 73 to communicate with each subsystem and control the execution of a number of instructions from the system memory 72 or storage device(s) 79 (e.g., a fixed disk such as a hard drive or an optical disk) and also exchange information between the subsystems. The system memory 72 and / or the storage device(s) 79 may embody a computer-readable medium. Another subsystem is a data collection device 85 such as a camera, microphone, accelerometer, etc. Any of the data described herein can be output from one component to another, or to a user.

[0232] A computer system may include multiple identical components or subsystems connected to each other, for example, by an external interface 81, by an internal interface, or through storage devices that can be connected and disconnected from one component to another. In some embodiments, the computer systems, subsystems, or devices may communicate over a network. In such a case, one computer may be considered a client and another computer a server, each of which may be part of the same computer system. The client and server may each include multiple systems, subsystems, or components.

[0233] Aspects of the embodiments can be implemented in the form of control logic using hardware circuits (e.g., application specific integrated circuits or field programmable gate arrays) and / or computer software with a processor that is generally programmable in a modular or integrated manner. A processor as used herein can include dedicated hardware as well as single-core processors, multi-core processors on the same integrated chip, or multiple processing units on a single circuit board or over a network. Based on the disclosure and teachings provided herein, one of ordinary skill in the art may know and appreciate other ways and / or methods of implementing embodiments of the present disclosure using hardware and combinations of hardware and software.

[0234] Any software components or functions described in this application may be implemented as software code executed by a processor using any suitable computer language, such as, for example, Java, C, C++, C#, Objective-C, Swift, or script, or scripting languages, such as, for example, Perl or Python, using conventional or object-oriented techniques. The software code may be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission. Suitable non-transitory computer-readable media may include random access memory (RAM), read-only memory (ROM), magnetic media, such as hard drives or floppy disks, or optical media, such as compact discs (CDs), digital versatile discs (DVDs), or Blu-ray discs, flash memory, and the like. The computer-readable medium may be any combination of such devices. Additionally, the order of operations may be rearranged. A process may terminate when its operations are completed, but may include additional steps not included in the figures. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, and the like. If a process corresponds to a function, its termination may correspond to a return of the function to the calling function or to the main function.

[0235] Such programs may also be encoded and transmitted using carrier signals adapted for transmission over wired, optical, and / or wireless networks conforming to various protocols, including the Internet. Thus, computer-readable media may be created using data signals encoded with such programs. Computer-readable media encoded with program code may be packaged in a compatible device or provided separately from other devices (such as by Internet download). Any such computer-readable media may be present on or within a single computer product (such as a hard drive, CD, or an entire computer system) or may be present on or within different computer products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing a user with any of the results referred to herein.

[0236] Any of the methods described herein may be performed in whole or in part on a computer system including one or more processors that may be configured to perform the steps. Thus, embodiments may be directed to a computer system configured to perform any of the steps of the methods described herein, potentially with different components performing each step or group of steps. Although shown as numbered steps, steps of the methods herein may be performed simultaneously or at different times or in different orders. Furthermore, some of these steps may also be used with some of the other steps of other methods. Also, all or some of the steps may be optional. Furthermore, any steps of any method may be performed using modules, units, circuits, or other means in a system for performing these steps.

[0237] The specific details of the particular embodiments may be combined in any suitable manner without departing from the spirit and scope of the embodiments of the present disclosure, however, other embodiments of the present disclosure may be directed to particular embodiments relating to each individual aspect or particular combinations of these individual aspects.

[0238] The above description of the exemplary embodiments of the disclosure has been presented for purposes of illustration and description, it is not intended to be exhaustive or to limit the disclosure to the precise form described, and many modifications and variations are possible in light of the above teaching.

[0239] The words "a," "an," or "the" are intended to mean "one or more," unless specifically indicated to the contrary. The use of "or" means "inclusive or," and not "exclusive or," unless specifically indicated to the contrary. Reference to a "first" component does not necessarily require that a second component be provided. Furthermore, reference to a "first" or "second" component does not limit the referenced component to a particular location, unless expressly stated. The term "based on" is intended to mean "based at least in part on." When Markouche groups or other groupings are used herein, it is intended that all individual members of the group, and all possible combinations and subcombinations of the group, are individually included in the disclosure.

[0240] All patents, patent applications, publications, and descriptions referred to herein are incorporated by reference in their entirety for all purposes. None are admitted to be prior art. In the event of a conflict between this application and any references provided herein, this application shall control.

Claims

1. 1. A method executed by a computer system, comprising: (a) receiving N hypotheses each containing the same number of words; (b) predicting a set of K consecutive words for each hypothesis using a language model; (c) combining each hypothesis with a corresponding set of K consecutive words to obtain N*K consecutive words; (d) converting each successive word into a predicted brain response of L candidate response values ​​by an encoding model, each candidate response value corresponding to a measurement of a different part of the brain or a different sensor; (e) receiving brain activity measurements, including L brain measurements, of the subject for the current time period; (f) comparing the brain activity measures comprising the L brain measures with the predicted brain responses of the L candidate response values ​​for each of the N*K consecutive words to obtain N*K scores; (g) identifying the top N consecutive words based on the N*K scores; (h) repeating (a) through (g); and (i) outputting a series of words corresponding to the brain activity measure for the subject based on the top N consecutive words.

2. 10. The method of claim 1, wherein the brain activity measurements are obtained using functional magnetic resonance imaging (fMRI), and the current duration of the fMRI is between 1 second and 5 seconds.

3. 10. The method of claim 1, wherein the brain activity measurements are obtained using functional near-infrared spectroscopy (fNIRS), and the current period of fNIRS is between 1 second and 5 seconds.

4. 10. The method of claim 1, wherein the brain activity measures of the subject are measured by a brain imaging device of the computer system as the subject receives stimulation.

5. 5. The method of claim 4, wherein the subject receives the stimulus by thinking, reading, or hearing the stimulus.

6. The method of claim 1 , wherein the brain activity measurements are pre-recorded.

7. The method of claim 1 , wherein the N hypotheses have customizable default values ​​based on different use cases.

8. 10. The method of claim 1, wherein each successive word is composed of a word sequence, and each word sequence correlates with the number of words perceived by the subject at acquisition time.

9. 9. The method of claim 8, wherein the number of words perceived by the subject at the acquisition time is predicted using a word time decoder.

10. said converting each successive word into a predicted brain response of L candidate response values; Transforming the subsequent words into a set of word embedding vectors, each word embedding vector of the set of word embedding vectors correlating with each word in the word sequence; downsampling the set of word embedding vectors to generate an averaged word embedding vector; applying a convolution kernel to the averaged word embedding vector and the previous averaged word embedding vector to generate a final word embedding vector; and varying features of the final word embedding vectors to the predicted brain responses of L candidate response values.

11. The method of claim 1 , wherein the L brain measurements are L voxel measurements comprising responses from L different voxels of the brain, each voxel eliciting a different measurement.

12. 11. The method of claim 10, wherein the encoding model has dimensions JxL, where J represents the number of features in the final word embedding vector and L represents the number of candidate response values.

13. 11. The method of claim 10, wherein the convolution kernel includes different weights, each weight being determined based on the degree of influence of the averaged word embedding vector and the previous averaged word embedding vector on the brain activity measure of the subject in the current time period.

14. 1. A method executed by a computer system, comprising: (a) annotating each word in a word sequence with a time label, the word sequence being associated with one acquisition time; (b) transforming the word sequence into a set of word embedding vectors using a neural language model, where each word in the word sequence corresponds to a word embedding vector; and (c) using the word embedding vectors to determine final word embedding vectors; and (d) receiving brain activity measurements, the brain activity measurements comprising L brain measurements of the subject; (e) determining a linear mapping of the final word embedding vectors to a voxel space of the brain activity measurements comprising L brain measurements to determine an encoding model.

15. The method of claim 14 , wherein the annotating the word sequence with the time labels is performed by speech recognition software or a human annotator.

16. 15. The method of claim 14, further comprising: determining a word time decoder that can predict the number of words in a single acquisition time by mapping a vector of the brain activity measures and word rates in the word sequence.

17. The method of claim 14 , wherein each word embedding vector of the word represents the semantics and / or syntax of the word.

18. determining the final word embedding vectors using the word embedding vectors, downsampling the set of word embedding vectors to generate an averaged word embedding vector; and applying a convolution kernel to the averaged word embedding vector and a previous averaged word embedding vector to generate the final word embedding vector.

19. A computer program comprising instructions which, when executed, control a computer system to carry out the method of any one of claims 1 to 18.

20. a computer readable medium storing instructions for the computer program of claim 19; and one or more processors for executing instructions stored on the computer-readable medium.

21. A system comprising one or more processors configured to perform the method of any one of claims 1 to 18.