Systems and methods involving decoding / processing imagined sentences and / or discrete language via non-invasive brain / neural interfaces and / or other features
The fNIRS-based system addresses the limitations of existing neuroimaging modalities by achieving high accuracy in decoding imagined speech, facilitating direct human-AI communication and improving interaction for individuals with communication disorders.
Patent Information
- Application Number
- PCT/US2025/034854
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-21
- Filing Date
- 2025-06-23
- Publication Date
- 2025-12-26
AI Technical Summary
Existing neuroimaging modalities like EEG are sensitive to motion artifacts and lack spatial resolution, limiting their effectiveness in decoding imagined speech, while fNIRS has not been successfully implemented for this purpose despite its potential advantages.
A novel fNIRS-based system combining advanced signal processing techniques and machine learning algorithms to classify imagined sentences from brain activity patterns, using high-density fNIRS systems and decoding algorithms to achieve above-chance accuracy in detecting imagined speech.
Demonstrates the feasibility of fNIRS in decoding imagined sentences with high accuracy, enabling direct interaction with an LLM for near-real-time human-AI communication, enhancing human-AI interaction and providing an alternative for individuals with communication disorders.
Smart Images

Figure US2025034854_26122025_PF_FP_ABST
Abstract
Description
[0001] SYSTEMS AND METHODS INVOLVING DECODING / PROCESSING IMAGINED SENTENCES AND / OR DISCRETE LANGUAGE VIA NON- INVASIVE BRAIN / NEURAL INTERFACES AND / OR OTHER FEATURES
[0002] Cross-Reference to Related Application^) Information
[0003] [1] This application claims benefit / priority under the provisions of the Patent Cooperation Treaty, and all other applicable National Stage provisions, of U.S. Provisional Application No. 63 / 663,053, filed June 21 , 2024, the contents of which are incorporated herein by reference and inclusion in their entiret .
[0004] Background
[0005] Field
[0006] [2] The disclosed technology relates, inter alia, to the field of human-AI interact on(s) via implementation of sy stems and methods involving detection, processing, and / or decoding of brain data.
[0007] Description of Related Information
[0008] [3] In the coming decade, artificial intelligence systems are improving exponentially and are set to revolutionize ever}’ industry and facet of human life. Building communication systems that enable seamless and symbiotic communication between humans and Al agents is increasingly important. Human-AI interaction and human-computer interaction could be radically enhanced if humans had the direct capability to transmit their thoughts to artificial intelligence. The concept of humans imagining language in their mind without any physical involvement is sometimes referred to as imagined speech. Therefore, decoding imagined speech using an Al model trained on brain data of participants imagining speech offers the potential for a natural and intuitive means of communication. An imagined speech system could make life better for people who have difficulty communicating and also revolutionise consumer technology and interaction with future Al assistants.
[0009] Overview of Aspects of the Disclosed Technology
[0010] [4] One or more aspects of exemplary systems and methods of the disclosed technologymay involve decoding imagined speech using brain data, such as that obtained via high- density functional near-infrared spectroscopy (fNIRS), EEG, MEG, and / or other brain data modalities. As set forth below aspects of the disclosed technology may include or involve various thought-to-LLM (large language model) systems and / or methods. As set forth in more detail below, various aspects of the disclosed technology are referred to herein as "MindGPTy overall. [5] One or more aspects of the present disclosure generally relate to improved computer- based systems, algorithms, devices (including wearable devices, etc.), methods, platforms and / or user interfaces, and / or combinations thereof, and particularly to, improved computer- based systems, methods, devices, hardware architectures, platforms and / or user interfaces associated with mind / brain-computer interfaces that are driven by various technical features and functionality set forth below, and may involve various related aspects such as laser / optical-based brain signal acquisition, decoding modalities, encoding modalities, and brain-computer interfacing, among other features set forth herein.
[0011] [6] Aspects of the disclosed technology and platforms here may comprise and / or involve processes of collecting and processing brain activity data, such as those associated with the use of a brain-computer interface that enables, for example, decoding and / or encoding a user’s brain / neural activities / activity patterns associated with language and semantic based information. Systems and methods herein may include and / or involve the leveraging of innovative brain-computer interface aspects and / or non-invasive wearable or portable devices to facilitate and enhance user interactions that provide technical outputs, solutions and results, such as those required for or associated with next generation wearable and / or Al devices, controllers, and / or other computing components based on human thought / brain / mind signal and sentence detection and / or related computer device processing and interaction.
[0012] [7] In some embodiments, the systems described herein may include, involve and / or receive information from a non-invasive brain-interface system that enables various imagined sentences, discrete language and the like to be decoded directly from the brain of a user, including various use and / or implementation of such information as commands and / or for interactions with a full breadth of computing devices, software, Al-enabled devices or tools, augmented and virtual reality and metaverse-related devices or content, among others.
[0013] Brief Description of the Drawings
[0014] [8] Various embodiments of the present disclosure may7be further explained with reference to the attached drawings, wherein like structures are referred to by like numerals throughout the several views. The drawings shown are not necessarily to scale, with emphasis instead generally being placed upon illustrating the principles of the present disclosure. Therefore, specific structural and functional details disclosed herein are not to be interpreted as limiting, but merely as a representative basis for teaching one skilled in the art to variously employ one or more illustrative embodiments. [9] FIG. 1 depicts an example of a whole head arrangement of an exemplary high-density fNIRS montage and / or set of channels, consistent with various exemplary aspects of one or more implementations of the disclosed technology.
[0015]
[0010] FIG. 2 depicts an example brain-computer interface in use by a user, consistent with various exemplary aspects of one or more implementations of the disclosed technology.
[0016]
[0011] FIG. 3 depicts an example method for processing imagined sentences, consistent with various exemplary aspects of one or more implementations of the disclosed technology.
[0017]
[0012] FIG. 4 depicts aspects of an exemplar}7application of an illustrative implementation of the disclosed technology7processing imagined sentences, consistent with various exemplary aspects of one or more implementations of the disclosed technology.
[0018]
[0013] FIG. 5 depicts an illustrative summary of representative classifier model performance in classifying imagined speech versus rest for certain exemplary brain / fNIRS data, consistent with various exemplary aspects of one or more implementations of the disclosed technology.
[0019]
[0014] FIG. 6 depicts illustrative surface cortex plots of HbO activations of the contrast between imagined speech and rest condition conditions for individuals studied using the present inventions, consistent with various exemplary aspects of one or more implementations of the disclosed technology.
[0020]
[0015] FIGS. 7A-7C depicts aspects of an exemplary7application of an illustrative implementation of the disclosed technology performing processing and / or interactions with a large language model via a brain-computer interface, consistent with various exemplary aspects of one or more implementations of the disclosed technology7.
[0021]
[0016] FIGS.. 8A-8B depicts additional exemplary7interactions involving a large language model via a brain-computer interface, consistent with various exemplary aspects of one or more implementations of the disclosed technology7.
[0022]
[0017] FIG. 9 depicts an exemplary flow diagram and associated method involving creation and / or use of a model that may decode sentences, consistent with various exemplary aspects of one or more implementations of the disclosed technology7.
[0023] Detailed Description of Certain Illustrative Implementations
[0024]
[0018] Systems and methods associated with mind / brain-computer interfaces are disclosed. Embodiments herein include features related to one or more of optical-based brain signal acquisition, decoding modalities, encoding modalities, brain-computer interfacing, and human- Al interaction, among other features set forth herein. Certain implementations may include or involve processes of collecting and processing brain activity data, such as those associated with the use of a brain-computer interface that enables, for example, decoding and / or encoding imagined sentences, discrete language, a user’s brain functioning, neural activities, and / or activity patterns associated with thoughts, including sensory -based thoughts. Further, the present systems and methods may be configured to leverage brain-computer interface and / or non-invasive wearable device aspects to provide enhanced user interactions for Al applications, next-generation wearable devices, controllers, and / or other computing components based on the human thoughts, brain signals, and / or mind activity, such as imagined sentences and discrete language, that are detected and processed.
[0025]
[0019] According to various exemplary embodiments herein, the present systems and methods may enhance human- Al communication by utilizing brain / brain imaging data to develop an Al model (i.e., MindGPT) capable of decoding imagined speech. In various example embodiments disclosed herein, the brain / brain imaging data is disclosed as being acquired via an illustrative fMRI modality7. However, implementations herein may also utilize other types of NIRS setup, EEG setups, etc., as well as fMRI and other neuroimaging modalities, if desired. Note that using fNIRS on its own is just one embodiment; it is also feasible that the methodology can be leveraged with data different from fNIRS, such as EEG data of imagined sentences or fMRI data of imagined sentences, using the same data collection setup in terms of the user imagining sentences. For example, fNIRS can be used together with EEG and the data collected in concert using the methodology described in more detail below.
[0026]
[0020] Consistent with the disclosed technology7, hemodynamic responses representing neural activity7may be collected from various participants, e.g., participants instructed to imagine three different sentences in one exemplary instance detailed herein. In some implementations, an Extra Trees Classifier (XTC) model may be employed to decode neural patterns and differentiate imagined speech from rest conditions, with the systems and methods herein achieving an average accuracy of -66% across participants, with the best average accuracy at 71%. To further decode neural signals associated wi th specific imagined sentences, a convolutional neural network (CNN) and a ridge regression model may be used as decoders in some embodiments. Among other things, the CNN model demonstrated an advantage with minimally preprocessed optical density data, outperforming the ridge regression model in this task.
[0027]
[0021] According to exemplary implementations herein, output results showed significant decoding accuracy for imagined speech, with the ridge regression model achieving a best accuracy of 57% for one part icipant (chance level: 33%, p-value < 0.001) and the CNN model achieving 47% (chance level: 33%. p-value < 0.001).
[0028]
[0022] In one embodiment, the systems described herein may include a near real-time Al communication system using brain data, such as, e.g., via fNIRS technology, etc., and suitable application software, e.g., a web framework or micro web framework, such as a Flask application, etc., enabling early-stage thought-based communication between participants and an LLM (e.g.. the OpenAI GPT-4 API). Further, the advantages of this direct communication channel extend across various fields, yielding significant improvements for human- Al interaction.
[0029]
[0023] Use functional near-infrared spectroscopy (fNIRS) and / or other types of brain imaging data (e.g.. EEG. etc.), as set forth in connections with the technology herein, is shown to lend itself to innovative systems and methods of neuroimaging for BCI applications involving imagined speech decoding. fNIRS is a non-invasive technique that measures blood oxygenation changes in the brain as an indirect measure of neural activity', through monitoring the associated vascular responses (slow signal), similar to that of fMRI. In contrast to certain other techniques such as EEG, fNIRS provides better spatial resolution and is less sensitive to motion artefacts, allowing it to be applied in more naturalistic settings. fNIRS, being portable, harmless and cost-effective, as well as not requiring conductive gels or any elaborate set-up procedures, provides a number of advantages. High-density fNIRS systems, and high-density diffuse optical tomography (HD-DOT) as an extension, can provide better coverage and have been shown to report similar neuroimaging capabilities to fMRI with better temporal resolution given the higher sampling rate.
[0030]
[0024] Despite the potential advantages of fNIRS, its application in imagined speech decoding has not yet been successfully implemented, hence remains underexplored compared to other neuroimaging modalities. The systems described herein include a novel fNIRS-based system that combines advanced signal processing techniques and machine learning algorithms to classify imagined sentences with different semantic content from a predefined set, based on the analysis of brain activity patterns. In some examples, the systems described herein may achieve achieve above-chance decoding accuracy in detecting imagined speech, demonstrating the potential of fNIRS as a viable modality for imagined sentence decoding.
[0031]
[0025] In summary, the disclosed technology implements novel systems and methods for decoding imagined speech using high-density fNIRS and, moreover, shows the first demonstration of direct imagined sentences interaction with an LLM. In some embodiments, the systems described herein may include the following innovations: (i) Demonstrating the feasibility of fNIRS recording using off-the-shelf commercially available headgears for collecting high signal-to-noise ratio (SNR) signals during imagined speech, compared to rest brain function; (ii) Implementing decoding algorithms that can decipher and classify imagined sentences from a limited dictionary with relatively high accuracy, and (iii) Establishing an thought-based communication channel as a platform for new possibilities in human-AI interaction and synergy’. Moreover, the BCI innovations herein may be further optimized and refined for intriguing applications involving human-AI interaction, beyond imagined speech decoding.
[0032] Exemplary Systems, Methodologies and Setups for Gathering Training Data
[0033]
[0026] In some embodiments, the systems described herein may employ a Continuous-Wave (CW) high-density 48x48 fNIRS system that provides full-head coverage. An exemplary commercially available fNIRS system may include of 48 sources and 47 detectors (the extra detector is used for the short-distance channels). The system consists of a total of 388 channels (194 source wavelength at 760 nm and 194 source wavelength at 850 nm); sampling rate: 5.9 Hz; channel distances ranging from ~21 mm to ~42 mm), and 8 short-distance channels (channel distances: < 10 mm) providing high-density full-head coverage, to monitor changes in oxygenated blood levels in the brain as a proxy for neural activity’. One exemplary’ embodiment of such implementation is shown in Figure 1.
[0034]
[0027] According to various embodiments establishing successful implementation of the disclosed technology’, participants were asked to sit in front of a computer screen, with their hands resting on the table, and the fNIRS fibre bundles arranged in a ponytail attached to the main fNIRS box. Here, for example, an illustrative setup is shown in FIG. 2. In one example test sequence, participants memorised three sentences before the onset of the experiment. Participants were reminded using a beeping metronome tone to imagine the sentences at a set pace, while their brain was recorded with fNIRS. Using a metronome ticking at 100 beats per minute as a reference pace for the imagined sentences ensured consistent timing and speed with the intention for participants to imagine the sentences at 100 words per minute.
[0035]
[0028] To avoid an effect of sentence duration on decoding performance, the sentences were carefully designed to have similar lengths of approximately 25 seconds when imagined. While a set sentence length was chosen for testing purposes, in non-testing applications, the systems described herein may decode sentences of any duration and / or multi-sentence passages. Moreover, the selected sentences (examples in Table 1 below) were crafted to have distinct semantic content (BERT score calculation between sentences was used as a measure of semantic similarity). To obtain a robust dataset, each sentence was imagined repeatedly by participants in a randomised order. The total number of imagined speech trials collected from participants were: 423 for participant 1, 378 for participant 2, 423 for participant 3, and 419 for participant 4. These repetitions were collected across 11 to 16 separate experimental sessions for each participant. All imagined speech trials included an equivalent number of the three imagined sentences (e.g., the 423 imagined speech trials collected from participant 1 included 141 trials for each of the three imagined sentences). Rest condition trials were also collected (n = 25 for participant 1, n = 126 for participant 2, n = 141 for participant 3, and n = 140 for participant 4) with the same duration (25s) as imagined sentence trials.
[0036] Table 1. Imagined sentences with distinct semantic content selected for the 3-class classification
[0037] Illustrative fNIRS Data Preprocessing
[0038]
[0029] As needed for certain implementations and consistent with one example implementation of the disclosed technology utilizing fNIRS, different levels of preprocessing of raw fNIRS data may be implemented to identify which data preparation method would lead to best model performance. Such preprocessing may be varied according to systems and methods herein, e.g., to balance between a lengthy preprocessing of fNIRS data to extract relevant information to decode imagined speech and the complexify of the models used. For example, while CNNs are more complex models compared to a ridge regression model, CNNs may be utilized to extract features from minimally preprocessed data, while a higher level of preprocessing may be needed for the latter. As such, consistent with certain embodiments of the disclosed technology, raw or minimally preprocessed data (e.g., optical density, etc.) achieves higher performance when selecting the CNN decoder, while fully preprocessed data may be processed via ridge regression, or other such simpler model(s). In some embodiments, the brain data collected during sentence imagining may be streamed through NIRx acquisition software and saved in the XDF format. The saved data may then be loaded and subjected to a series of preprocessing techniques. Specifically, the systems described herein may train models using (a) raw data (e.g., intensity only), (b) optical density' data (e.g., minimally preprocessed data which underwent conversion of raw signals to optical density data), and / or (c) fully preprocessed data (e.g., data that underwent conversion of raw signals to optical density', detrending, short channel regression correction, motion artefact correction, conversion to haemoglobin concentration using a partial pathlength factor (ppf) of 6, and bandpass filtering between 0.01 and 0.7 Hz). Finally, regardless of the level of preprocessing, the data may be trimmed to a fixed shape (e.g., 145 time points by 388 channels and / or 194 oxy-haemoglobin and 194 deoxy-haemoglobin channels) and saved as HF5 files, with the filenames corresponding to the respective sentence names.
[0039] Imagined Speech Detection Decoder
[0040]
[0030] In one embodiment, the systems described herein may apply an Extra Trees Classifier (XTC) model to fully preprocessed brain signals to decode imagined speech. In one example, this model may be implemented using the ExtraTreesClassifier from the scikit-leam with the following parameters: bootstrap = False; criterion = "entropy"; max_features = 0.2; min_samples_leaf = 6 min_samples_split = 7; n_estimators = 100. The XTC is an ensemble learning method based on the random forest algorithm, which fits multiple decision trees on various sub-samples of the dataset and uses averaging to improve predictive accuracy and control over-fitting.
[0041]
[0031] With regarding to various exemplary techniques utilized in some embodiments of the disclosed technology, a stratified k-fold cross-validation with 3 folds maybe be employed and tests may be conducted over 5 different seeds to ensure the model's robustness and versatility capability. Stratified sampling may maintain the class distribution across folds. Decoder performances may be assessed using decoding accuracy as the metric of success, which is defined as the accuracy of classification of the predicted test set over many trials. To assess the significance of the classification results, a p-value may be calculated using the cumulative distribution function (CDF) of the binomial distribution. The p-values from each fold were combined using Chi-squared distribution to obtain a single p-value representing the overall statistical significance of the results. Class-wise accuracy distribution was also tracked to analyse the model's performance across different classes (imagined speech vs rest condition). In one example, all models were trained on single-subjects to determine best accuracy per subject, although results from subjects may be averaged to identify overall performance of models in successfully completing imagined speech detection.
[0042] Imagined Sentences Decoder
[0043]
[0032] Certain embodiments of the systems described herein may include one or both of two decoding models. In one embodiment, the first model may include a one-dimensional Convolutional Neural Network (1D-CNN) architecture to analyse time-series fNIRS data. In addition, the systems described herein may use a ridge regression model for imagined sentence classification as the baseline. Both models can be used in the final near real-time demonstration.
[0044]
[0033] In one embodiment, the systems described herein may develop a 1D-CNN with multiple layers for each participant (subject-specific) for the analysis of time series fNIRS data. In some embodiments, a 1D-CNN may represent a good balance between model complexity and effectiveness at extracting relevant features for imagined speech decoding, thus being preferable to the more complex 2D- or 3D-CNNs. In one embodiment, regularization techniques, including dropout and fully connected layers, may be incorporated into the model.
[0045]
[0034] Consistent with certain implementations, decoder performance and / or decoding accuracy results may be utilized as the metric of success, which is defined as the accuracy of classification of predicted test set over many trials. Further, imagined sentence and rest condition labels may be selected as the ground truth, e.g., in the example described herein. Here, for example, a 3-fold cross validation may be run a plurality of times (e.g., 5 times, etc.) with different random seeds. Average and best fold accuracy may then be determine by seed and participant, and such values may be combined to determine a single value for participants’ average and best accuracy across different seeds. In some instances, for each fold, a p-value may be calculated using the cumulative distribution function (CDF) of the binomial distribution. The combined p-value of all 3 folds across seeds may then be calculated using a test, such as Fisher’s test.
[0046] 1D-CNN Model / Training
[0047]
[0035] According to embodiments herein, a 1D-CNN with multiple layers may be utilized for each participant (subject-specific) for the analysis of time series fNIRS data. Here, for example, direct application of such 1D-CNN models may be utilized to decode imagined speech, and represents a good balance between model complexity and effectiveness at extracting relevant features for imagined speech decoding. Hence, in many instance, such 1D-CNN models are more preferable to the more complex 2D- or 3D-CNNs. Regularisation techniques, including dropout and fully connected layers, may also be incorporated into such models / modelling. Further, according to certain implementations, Xavier initialization may be employed for weight initialization, and following softmax activation (for two-class classification for imagined speech vs rest condition) or sigmoid activation (for three-class classification for the three imagined sentences), the final layer's output may be determined.
[0048]
[0036] The recorded fNIRS data may be partitioned into training, validation, and test subsets. The scikit-leam’s Robust scaler may be fitted with training data and applied to validation and test data separately. The model may be trained and subsequently evaluated on the validation set before final performance assessment on the test set. The Cross Entropy Loss may be utilised as the loss function, while the Adam optimizer may be employed for optimization steps. The learning rate may be optimised using the ReduceLROnPlateau scheduler. Furthermore. LI and L2 regularisation, early stopping, and cross-validation may be implemented, along with a search for LI and L2 coefficients. Model performance may be evaluated using the accuracy of the predicted test set, where imagined sentence and rest condition labels may be selected as ground truth. This same procedure may be followed for all data, regardless of the preprocessing steps taken to prepare the data, and for all participants.
[0049] Ridge Regression Model
[0050]
[0037] Additionally or alternatively, a cross-validated ridge regression model may be implemented for each participant (subject-specific) as the baseline model to convert fNIRS brain signals to sentence embeddings. For example, scikit-leam’s RidgeCV may be used with the alpha range set to 10e'3- 10e3where the best alpha value that determines regularisation strength is determined with cross validation. The sentence embeddings may be created using SentenceTransformer, where each sentence is encoded into 768-element ID vectors. The base model all-mpnet-base-v2' may be used. The predicted embeddings from ridge regression may then be fed into a logistic regression classifier to output the predicted identity of the sentence.
[0051]
[0038] For pre-processing, the fNIRS data may be scaled with scikit-leam’s Robust scaler. The first 10 samples of each trial may be averaged to obtain the baseline, which may be then subtracted from the rest of the signal for removing the baseline. Delta Hb / HbO value over 16 (pMol for haemo data) may be clamped in order to remove outliers. The scaler may be fitted with training data and applied to validation and test data separately. This same procedure may be followed for all data, regardless of the preprocessing steps taken to prepare the data, and for all participants.
[0052] Imagined Speech Related Brain Activations
[0053]
[0039] In addition to applying decoding models directly to the imagined speech data, analyses involving haemodynamic response modelling and GLM-based statistical testing may be conducted to identify imagined speech-related activations in the brain. To identify semantic representation in the brain, brain activations during imagined speech and rest condition may be compared (e.g., fully preprocessed fNIRS data may be used). This may be repeated for all participants.
[0054]
[0040] In some examples, the fNIRS data may be fully preprocessed using the steps described above. A first-level design matrix may be constructed to model the hemodynamic response associated with neural activity, for example using the python package mne-nirs (version 0.6.0). Processed haemo data may be obtained by converting raw fNIRS signals to optical density signals, and then into haemo data using the Beer Lambert law with ppf=6. Channels (distance > 10mm) are used in the general linear model with a cosine function to model and correct for low-frequency drift in the signal, and a high-pass filter of 0.005Hz to remove slow signal variations not contributed to neuronal activities. The adopted hemodynamic response function (HRF) model is based on the Statistical Parametric Mapping (SPM) approach, a standard model for estimating the brain's vascular response to neural activity. Lastly, different stimulus durations may be considered to assess the temporal progression of brain activations throughout the 25s interval (durations of 5s, 10s, 15s, 20s, and 25s). Full brain activation results were assessed and recorded, and the most characteristic activation for each participant was noted in the results section. In some examples, short channel data (distance < 10mm) are included as nuisance regressors in the design matrix. Conditions ‘Imagined speech’ and ‘rest condition’ are specified and the GML parameters are estimated. The contrast 'Imagined speech > rest condition’ is then estimated from GLM theta values, and z-scores are calculated. Surface plots are generated with the estimated z-scores.
[0055] Illustrative Application of the Inventive Systems and Methods (MindGPT)
[0056]
[0041] Consistent with the disclosed technology', certain fNIRS-based systems herein may be implemented such that the BCI provides human- Al thought-based communication (MindGPT). In some embodiments, the systems described herein may include an application (e.g.. based on a micro-framework such as Flask, etc.) that automates sending decoded human thoughts to an LLM such as ChatGPT. The core of such innovative systems herein involve integration of brain activity data, captured during sentence imagery tasks, with the capabilities of the OpenAI GPT4 API for a direct mind-to-OpenAI communication. Here, for example, brain data may be stored in the SNIRF format, with the raw data being preprocessed via a series of steps as described above.
[0057]
[0042] In one embodiment, the 1D-CNN model may serve as the foundation for classification tasks. This setup may enable dynamic retrieval of corresponding texts through a dictionary lookup mechanism, with the decoded sentence being stored, subsequently fetched by a server, and presented within a web interface. Interactions with the OpenAI GPT4 API may be driven by these queries, with the system capable of receiving and displaying responses to users in near real-time, as illustrated in FIG. 3.
[0058]
[0043] In one example confirmatory experiment, as illustrated in FIG. 3, a user first imagines their preferred sentences from a predefined set of 3 sentences, while their brain data is being recorded using fNIRS. In this example, the participant is imagining the sentence related to going to the ‘Restaurant". Then, the participant’s brain data is provided as input into a 1D-CNN decoder which processes fNIRS data and classifies the data into one of the 3 sentence options (‘Call’, ‘Restaurant’, ‘Venus’). The text associated with the classified sentence is then retrieved from a look up table and sent as a prompt to ChatGPT in the application. ChatGPT then generates an answer based on the user’s imagined input, attempting this way to create a coherent dialogue. In some examples, this process happens in near real-time.
[0059]
[0044] In some examples, the application may function in ‘near real-time’ as a slight delay is present between new data collection from a user and the decoded output, resulting from the need to start and stop the data recording system between decoding attempts. This limitation is driven by the NIRx acquisition software used to collect fNIRS data, but embodiments utilizing different software may achieve a traditional real-time application. GPT4 may be instructed to provide useful suggestions based on the user’s imagined input, in an attempt to create a coherent dialogue.
[0060]
[0045] In another example, as illustrated in FIG. 4, a user first imagines their preferred sentences from a predefined set of 3 sentences, while their brain data is being recorded using fNIRS. In this example, the participant is imagining the sentence related to going to the ‘Restaurant’. Then, the participant’s brain data is provided as input into a Id-CNN decoder which processes fNIRS data and classifies the data into one of the 3 sentence options (‘Call’, ‘Restaurant’. ‘Venus'). The closest matching ChatGPT-generated question from a list of 10 questions to the classified imagined sentence (i.e., Restaurant in this case), measured by cosine similarity of sentence embeddings, is then picked from a lookup table and sent to GPT4 to create a coherent dialogue. These steps are repeated multiple times to allow for a continued conversation.
[0061]
[0046] In some embodiments, for each imagined sentence, 10 related questions to the specific topic may be generated using an LLM tself prior to the MindGPT experiment, each of which may be also converted to embeddings using SentenceTransformer. After decoding the first sentence and sending it to GPT4, the next round of dialogue may be introduced by comparing the ridge regression-generated sentence embeddings. The closest matching LLM-generated question to the imagined sentence, measured by cosine similarity of sentence embeddings, is then picked from a lookup table and sent to GPT4 again to continue the conversation. This is similar to a zero-shot learning approach, where new test cases are unseen by the trained model.
[0062]
[0047] In one embodiment, the latency of MindGPT thought-based communication may be 27.62 seconds (e.g., using hardware such as CPU: AMD Ryzen 9 5900X 12-Core Processor; GPU: NVIDIA GeForce RTX 3080 Ti; RAM: 64.0 GB). The mam bottleneck lies in the extended time required to imagine a sentence (25 seconds), while decoding is fast once the decoder is trained (e.g., 2.62 seconds - loading fNIRS data file: 0.66 seconds; preprocessing fNIRS data: 0.61 seconds; imagined sentence decoding: 1.34 seconds; trigger sent to ChatGPT: 0.01 seconds). Therefore, the long latency of near real-time MindGPT is a result of task (imagining sentences) limitations, rather than decoding approach limitations (e.g., long preprocessing needed or long decoding time).
[0063] Experimental Proofs / Results
[0064] Imagined Speech Decoding vs Rest Condition
[0065]
[0048] Validation of the innovations herein may be seen in the performance of using such XTC model(s) in detecting imagined speech from the rest condition when classifying neurovascular signals associated with the two conditions (see Figure 5). With regard to the 4- participant exemplary trial discussed herein, a total of 162 imagined speech and 162 rest condition trials were included for participants 2 and 3, while participant 4 counted 123 imagined speech and 123 rest condition trials due to time constraints. As mentioned previously, participant 1 was excluded from this analysis since not enough rest condition trials were collected from participant 1 to train and test a decoder (n = 25 rest condition trials only). Overall, our XTC model achieved an average accuracy of -66% (p-value < 0.001) when considering averaged accuracies across folds in the 3 subjects included in this test. Our best participant (participant 1) reported a best average accuracy across the 3 folds of -71% (p-value < 0.001).
[0049] FIG. 5 illustrates example MindGPT decoding of different imagined sentences (pretrained three imagined sentences) from brain data collected from the user at different instances of imagined speech. Figure 5 is a summan' of XTC model performance (accuracy %) in classifying imagined speech vs the rest condition when using fully preprocessed fNIRS data. Data, standard deviation and p-values are reported for participants 2, 3, and 4. Comparison to chance (50%) is also reported.
[0066] Imagined Sentence Decoding
[0067]
[0050] According to different embodiments, a comparison of the performance of such two exemplary models in decoding brain data into one of the 3 predefined imagined sentence classes across participants also helps validate the innovations herein. Further, comparison of the accuracy of the disclosed 1D-CNN and ridge regression-based language models, when using different preprocessing steps for our raw fNIRS data, may be performed in order to identify' which level of preprocessing leads to the best performance. Table 2 shows that the disclosed models were able to decode all participants' fNIRS neural signals associated with different imagined sentences and identify which of the three predefined sentences the participants were imagining to a significant level (chance level: 33%). In the example testing discussed herein, best accuracy was achieved in participant 2 with the language model-based decoder, when using fully preprocessed data (-57.4% accuracy, chance: 33%, p < 0.001). When comparing results more in detail across participants, the 1D-CNN was seen to outperform the ridge regression-based language model-based decoder in 2 out of 4 participants (participants 1 and 4), whereas the opposite was valid for participants 2 and 3. Further, in certain implementations, highest accuracy using the 1D-CNN may be achieved when data underwent minimal preprocessing, i.e., conversion of raw signals to optical density data. In contrast, highest accuracy using the Ridge Regression was obtained when fully preprocessed data was used. No preprocessing achieved worse performance in both models, compared to both minimally and fully preprocessed data. See Table 2.
[0068] Table 2. Summary of 1D-CNN and language model decoder performance in decoding imagined brain data into one of the three predefined sentences. Averaged accuracies across different folds and seeds were reported for each participant. These are presented by (a) model used - 1D-CNN or Ridge Regression-based language model; and (b) data type - Raw: no preprocessing, OD: optical density (or minimal preprocessed data), Preproc: preprocessed data (or fully preprocessed data with pipeline described). Standard deviation and statistical testing are reported in Supplementary Tables S2-S7. The highest accuracy level achieved using each of the two models (1D-CNN and Ridge Regression-based language model) is shown in bold-face numbers for each participant. Finally, the winning model as reported for each participant, indicating the model that achieved highest accuracy for that specific participant.
[0069] Brain Areas Underlying Imagined Speech
[0070]
[0051] Figure 6 illustrates the surface cortex plots of HbO activations of the contrast between imagined speech and rest condition conditions for all four participants described in the example study. While different stimulus durations were considered to assess the temporal progression of brain activation throughout the 25s during which participants imagined each sentence (durations of 5s, 10s, 15s, 20s, and 25s), this show's the most characteristic activation for each participant. This analysis identified a recurrent brain region recruited across 3 out of 4 participants during imagined speech, i.e., the dorsolateral prefrontal cortex (DLPFC). Interestingly, one of our participants (participant 4) showed a decrease in activation in the dorsolateral prefrontal cortex during imagined speech, which is an unexpected result. Additional regions differentially recruited between participants included the lateral temporal cortex, visual related regions near MT+ complex, auditory' cortex, and early sensorimotor cortex. Accordingly, based on such activations, the fNIRS-based BCI systems and methods herein are capable of capturing imagined speech processing brain activation.
[0071] Illustrative MindGPT Application(s)
[0072]
[0052] With regard to an illustrative application of the innovative fNIRS-based BCI disclosed herein, consider a scenario where such BCI enables early -stage human- Al thoughtbased communication (see Figure 3. section 2. 7 for a schematic of the mindGPT application). Further, Figures 7A-7C illustrate an example of such technology in action, i.e., a near realtime fNIRS-based BCI which enables telepathy-like human-AI communication. Brain signals are collected and processed in near real-time while the user imagines one of the three predefined sentences (see section 2.7). The data is then decoded via the innovative models herein (see section 2.5) and classified into one of the three imagined sentences upon which the models were trained. The decoded sentence is then utilized as a prompt for a language model (in this case ChatGPT), enabling telepathy-like communication between the user and ChatGPT. In some embodiments, the latency of the present thought-based communication is 27.62 seconds (CPU: AMD Ryzen 9 5900X 12-Core Processor; GPU: NVIDIA GeForce RTX 3080 Ti; RAM: 64.0 GB). The main bottleneck lies in the extended time required to imagine a sentence (25 seconds), while decoding is pretty' fast once the decoder is trained (2.62 seconds - loading INIRS data file: 0.66 seconds; preprocessing fNIRS data: 0.61 seconds; imagined sentence decoding: 1.34 seconds; trigger sent to ChatGPT: 0.01 seconds). Therefore, the long latency of the present near real-time systems / technology is a result of task (imagining sentences) limitations, rather than decoding approach limitations (e.g., long preprocessing needed or long decoding time).
[0073]
[0053] Again, Figures 7A-7C shows the disclosed technology’ successful decoding different imagined sentences (pretrained three imagined sentences) from brain data collected from the user at different instances of imagined speech (Figure 7A, 7B and 7C show successful decoding of sentences ‘Restaurant’, ‘Call’, ‘Venus’, respectively).
[0074]
[0054] Using the disclosed technology’, a continuous thought-based conversation between a user of the innovative BCI herein and ChatGPT, starting from the same three imagined sentences (Figure 4 was conducted to validate the viability of the present inventions. Here, for example, 10 different potential follow-up prompts were developed per each imagined sentence topic, e.g., 10 prompts associated with the imagined sentence topic “Restaurant” (see Table 3 below for an example of generated questions for follow-up of decoded imagined sentence topic “Restaurant”). These follow-up prompts were generated using chatGPT itself prior to the MindGPT experiment. To do this, we prompted ChatGPT with the 3 imagined sentences and asked it to generate 10 extra sentences related to the imagined sentence topics. During near real-time testing, after decoding the imagined sentences from brain signals, one out of the 10 follow-up prompts belonging to the decoded imagined sentence topic was chosen as a prompt to ChatGPT (see Figures 8A-8B). This was achieved by comparing the ridge regression-generated sentence embedding and selecting the closest matching questions to the generated sentence embedding (cosine similarity) to continue the conversation (see section 2.7 for further details on methodology). This implementation therefore allowed the user to sustain a continued conversation with ChatGPT on the same topic. In addition, hard- coding the input of one of the language model-generated sentences (based on the user’s imagined sentence) (see Table 3 for example follow-up generated questions) also meant that the user had limited control on how the conversation was continued with ChatGPT, apart from setting the ‘topic’ of conversation by imagining one of the three imagined sentences the decoder was trained on. Regardless of these limitations, this implementation nevertheless exemplifies the first continuous thought-based communication between a human user and Al (ChatGPT).
[0075]
[0055] According to on illustrative implementation, example questions generated for the exemplary imagined sentence “Restaurant” which are selected as follow-up prompts based on cosine similarity of sentence embedding may include questions such as: Are they recognized for any particular beverages such as wines or cocktails? What kind of ambiance or setting is the establishment known for? Is there a signature dish that stands out on their menu?How far in advance are reservations typically required? Do they offer any dishes that are specific to certain seasons? How does its reputation compare to other establishments in the city? Has the restaurant been awarded any culinary accolades? Who is the head chef and what is notable about their culinary background? How is the overall service quality rated by patrons? Is there a recommended dress code for diners? Are they recognized for any particular beverages such as wines or cocktails?
[0076]
[0056] FIGs. 8A-8B illustrates example sessions of continuous MindGPT in action. Different sentences (out of a preset of three sentences) are decoded from brain data collected from the user at different instances of imagined speech and inputted as prompts into ChatGPT to allow telepathy -like communication. To allow for a continuous conversation, after the first decoded imagined sentence, the closest matching question to the generated sentence embedding (cosine similarity) out of ten available questions related to the first decoded imagined sentence is selected as a follow-up prompt to ChatGPT. Example 802 is an example of two correctly decoded imagined sentences in a row (“Venus”) and use of a follow-up questions following the method outline. Example 804 is an example of a correctly decoded imagined sentence (“Restaurant”) followed by an erroneously decoded imagined sentence (“Call”) and a correctly decoded imagined sentence (“Restaurant”).
[0077]
[0057] Consistent with the disclosed technology', the various models implemented (1D-CNN and ridge regression-based language model) perform admirably in generating speech from detected brain data, based on the participant and the level of fNIRS data preprocessing (i.e., raw; optical density, or fully preprocessed data). However, there are subject-specific differences to consider and factor into the determined outputs / results when decoding imagined speech from different participants. Our 1D-CNN model outperformed the simpler ridge regression model used in fMRI pipelines when extracting semantic features from minimally preprocessed brain data using fNIRS. Accordingly, systems and methods herein implemented with the more complex CNN architecture better represent the semantic space during imagined speech without significant preprocessing, in certain embodiments. In certain implementations, other deep learning techniques may be utilized to reduce lengthy preprocessing, and again further confirms the disclosed technology's advantages of using deep learning for extracting features of interest from fNIRS neural data.
[0078]
[0058] In one final illustrative example, as illustrated in FIG. 9, a system 900 may enable user 904 to imagine a sentence 905 that is interpreted by headband 902. At a data acquisition step 910. the systems described herein may measure brain activity related to sentence 905. Next, at a signal processing step 915, the systems described herein may process this brain activity' as described in further detail above. At a step 920, the systems described herein may send the decoded sentence to an LLM.
[0079]
[0059] Consistent with the disclosed technology’, the field of neurotechnology and its societal impact are revolutionized by the development of such fNIRS-based systems and methods for decoding imagined speech. Further, this technology' also provides individuals with communication disorders an alternative means of expressing their thoughts and intentions, ultimately improving their quality of life and social interactions. Additionally, the integration of imagined speech decoding with Al systems leads to more natural and intuitive humanmachine communication, opening up new possibilities in various domains and propelling human evolution forward. The disclosed technology' is a significant step in this direction, implementing a fully -functional direct brain-to-AI communication using fNIRS.
[0080] Overall Implementations of the Disclosed Technology
[0060] While the above disclosure sets forth certain illustrative examples, the present disclosure encompasses multiple other potential arrangements and components that may be utilized to achieve the brain-computer interface innovations of the disclosed technology. Some other such alternative arrangements and / or components may include or involve other optical architectures that provide the desired results, signals, etc., while some such implementations may also enhance performance, correctness, and / or other metrics further. While the above disclosure sets forth certain illustrative examples, such as embodiments utilizing, involving and / or producing fast optical signal (FOS) and haemodynamic (e.g., NIRS, etc.) brain-computer interface features, the present disclosure encompasses multiple other potential arrangements and components that may be utilized to achieve the braininterface innovations of the disclosed technology. Some other such alternative arrangements and / or components may include or involve other optical architectures that provide the desired results, signals, etc. (e.g., pick up NIRS and FOS simultaneously for brain-interfacing, etc.), while some such implementations may also enhance resolution and other metrics further.
[0081]
[0061] Among other aspects, for example, implementations herein may utilize different optical sources, neuroimaging modality altogether, Al system to interact with and Al model to decode the brain data than those set forth, above. Here, for example, such optical sources may include one or more of: semiconductor LEDs, superluminescent diodes or laser light sources with emission wavelengths principally, but not exclusively within ranges consistent with the near infrared wavelength and / or low water absorption loss window (e.g., 700- 950nm, etc.); non-semiconductor emitters; sources chosen to match other wavelength regions where losses and scattering are not prohibitive; here, e.g., in some embodiments, around 1060nm and 1600nm, inter alia; narrow linewidth (coherent) laser sources for interferometric measurements with coherence lengths long compared to the scattering path through the measurement material (here, e.g., (DFB) distributed feedback lasers, (DBR) distributed Bragg reflector lasers, vertical cavity surface emitting lasers (VCSEL) and / or narrow linew idth external cavity lasers; coherent wavelength swept sources (e.g., where the center wavelength of the laser may be swept rapidly at 10-200 KHz or faster without losing its coherence, etc.); multi-wavelength sources where a single element of co packaged device emits a range of wavelengths; modulated sources (e g., such as via direct modulation of the semiconductor current or another means, etc.); and pulsed laser sources (e.g., pulsed laser sources with pulses betw een picoseconds and microseconds, etc.), among others that meet sufficient / proscribed criteria herein.
[0062] Implementations herein may also utilize different optical detectors than those set forth, above. Here, for example, such optical detectors may include one or more of: semiconductor pin diodes; semiconductor avalanche detectors; semiconductor diodes arranged in a high gain configuration, such as transimpedance configuration(s), etc. ; singlephoton avalanche detectors (SPAD); 2-D detector camera arrays, such as those based on CMOS {complementary metal oxide semiconductor} or CCD {charge-coupled device} technologies, e.g., with pixel resolutions of 5x5 to 1000x1000; 2-D single photon avalanche detector (SPAD) array cameras, e.g., with pixel resolutions of 5x5 to 1000x1000; and photomultiplier detectors, among others that meet sufficient / proscribed criteria herein.
[0082]
[0063] Implementations herein may also utilize different optical routing components than those set forth, above. Here, for example, such optical routing components may include one or more of: silica optical fibre routing using single mode, multi-mode, few mode, fibre bundles or crystal fibres; polymer optical fibre routing; polymer waveguide routing; planar optical waveguide routing; slab waveguide I planar routing; free space routing using lenses, micro optics or diffractive elements; and wavelength selective or partial mirrors for light manipulation (e.g. diffractive or holographic elements, etc.), among others that meet sufficient / proscribed criteria herein.
[0083]
[0064] Implementations herein may also utilize other different optical and / or computing elements than those set forth, above. Here, for example, such other optical / computing elements may include one or more of: interferometric, coherent, holographic optical detection elements and / or schemes; interferometric, coherent, and / or holographic lock-in detection schemes, e.g., where a separate reference and source light signal are separated and later combined; lock in detection elements and / or schemes; lock in detection applied to a frequency domain (FD) NIRS; detection of speckle for diffuse correlation spectroscopy to track tissue change, blood flow, etc. using single detectors or preferably 2-D detector arrays; interferometric, coherent, holographic system(s), elements and / or schemes where a wavelength swept laser is used to generate a changing interference patter which may be analyzed; interferometric, coherent, holographic system where interference is detected on, e.g., a 2-D detector, camera array, etc.; interferometric, coherent, holographic system where interference is detected on a single detector; controllable routing optical medium such as a liquid crystal; and fast (electronics) decorrelator to implement diffuse decorrelation spectroscopy, among others that meet sufficient / proscribed criteria herein.
[0084]
[0065] Implementations herein may also utilize other different optical schemes than those set forth, above. Here, for example, such other optical schemes may include one or more of: interferometric, coherent, and / or holographic schemes; diffuse decorrelation spectroscopy via speckle detection; FD-NIRS; and / or diffuse decorrelation spectroscopy combined with TD- NIRS or other variants, among others that meet sufficient / proscribed criteria herein.
[0085]
[0066] Implementations herein may also utilize other multichannel features and / or capabilities than those set forth, above. Here, for example, such other multichannel features and / or capabilities may include one or more of: the sharing of a single light source across multiple channels; the sharing of a single detector (or detector array) across multiple channels; the use of a 2-D detector array to simultaneously receive the signal from multiple channels; multiplexing of light sources via direct switching or by using “fast” attenuators or switches; multiplexing of detector channels on to a single detector (or detector array) via byusing “fast” attenuators or switches in the routing circuit; distinguishing different channels / multiplexing by using different wavelengths of optical source; and distinguishing different channels / multiplexing by modulating the optical sources differently, among others that meet sufficient / proscribed criteria herein.
[0086]
[0067] As disclosed herein, implementations and features of the present inventions may be implemented through computer-hardware, software and / or firmware. For example, the systems and methods disclosed herein may be embodied in various forms including, for example, one or more data processors, such as computer(s), server(s), and the like, and may also include or access at least one database, digital electronic circuitry, firmware, software, or in combinations of them. Further, while some of the disclosed implementations describe specific (e.g., hardware, etc.) components, systems, and methods consistent with the innovations herein may be implemented with any combination of hardware, software and / or firmware. Moreover, the above-noted features and other aspects and principles of the innovations herein may be implemented in various environments. Such environments and related applications may be specially constructed for performing the various processes and operations according to the inventions or they' may include a general-purpose computer or computing platform selectively activated or reconfigured by code to provide the necessary functionality. The processes disclosed herein are not inherently related to any particular computer, network, architecture, environment, or other apparatus, and may be implemented by a suitable combination of hardware, software, and / or firmware. For example, various general-purpose machines may be used with programs written in accordance with teachings of the inventions, or it may be more convenient to construct a specialized apparatus or system to perform the required methods and techniques.
[0068] In the present description, the terms component, module, device, etc. may refer to any type of logical or functional device, process or blocks that may be implemented in a variety of ways. For example, the functions of various blocks may be combined with one another and / or distributed into any other number of modules. Each module may be implemented as a software program stored on a tangible memory (e.g., random access memory, read only memory, CD-ROM memory, hard disk drive) within or associated with the computing elements, sensors, receivers, etc. disclosed above, e.g., to be read by a processing unit to implement the functions of the innovations herein. Also, the modules may be implemented as hardware logic circuitry implementing the functions encompassed by the innovations herein. Finally, the modules may be implemented using special purpose instructions (SIMD instructions), field programmable logic arrays or any mix thereof which provides the desired level performance and cost.
[0087]
[0069] Aspects of the systems and methods described herein may be implemented as functionality programmed into any of a variety' of circuitry7, including programmable logic devices (PLDs), such as field programmable gate arrays (FPGAs), programmable array logic (PAL) devices, electrically programmable logic and memory devices and standard cell-based devices, as well as application specific integrated circuits. Some other possibilities for implementing aspects include: memory7devices, microcontrollers with memory7(such as EEPROM), embedded microprocessors, firmware, software, etc. Furthermore, aspects maybe embodied in microprocessors having software-based circuit emulation, discrete logic (sequential and combinatorial), custom devices, fuzzy7logic, neural networks, other Al (Artificial Intelligence) or machine learning systems, quantum devices, and hybrids of any of the above device types.
[0088]
[0070] Other implementations of the inventions will be apparent to those skilled in the art from consideration of the specification and practice of the innovations disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the inventions being indicated by the present disclosure and various associated principles of related patent doctrine.
[0089]
[0071] As one overview of aspects of the disclosed technology, systems, methods and wearable devices associated with mind / brain-computer interfaces are disclosed. Embodiments herein include features related to one or more of optical-based brain signal acquisition, decoding modalities, encoding modalities, brain-computer interfacing, AR / VR content interaction, brain state assessment, signal to noise ration enhancement, and / or motion artefact reduction, among other features set forth herein. Certain implementations may include or involve processes of collecting and processing brain activity data, such as those associated with the use of a brain-computer interface that enables, for example, decoding and / or encoding a user’s brain functioning, neural activities, and / or activity patterns associated with thoughts, including sensory-based thoughts. Further, the present systems and methods may be configured to leverage brain-computer interface and / or non-invasive wearable device aspects to provide enhanced user interactions for next-generation wearable devices, controllers, and / or other computing components based on the human thoughts, brain signals, and / or mind activity that are detected and processed.
[0090]
[0072] It should also be noted that various logic and / or features disclosed herein may be enabled using any number of combinations of hardware, firmware, and / or as data and / or instructions embodied in various machine-readable or computer-readable media, in terms of their behavioral, register transfer, logic component, and / or other characteristics. Computer- readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, non-volatile storage media in tangible various forms (e.g., optical, magnetic or semiconductor storage media), though do not encompass transitory’ media.
[0091]
[0073] Additionally, while described in terms of specific software approaches above, other software implementations will be apparent to those skilled in the art from consideration of the specification and practice of the innovations disclosed herein. For example, another implementation of the systems described herein may include one or more CNNs, ridge regression models, etc., and / or may use transformers or other neural networks to decode brain data (e.g., imaged speech sentences, etc.). In some embodiments, the systems described herein may incorporate additional and / or alternative neuroimaging methodologies to extra brain data including but not limited to EEG, fNIRS and EEG together, MEG or fMRI, and / or even with more invasive neuro-signal acquisition techniques and methodologies. The Appendices information submitted herewith describe some additional example implementations and technologies that may be implemented with and / or involved in some embodiments, e.g., to provide brain signals, perform signal and / or data processing and / or otherwise be utilized in or with the disclosed technology’.
[0092]
[0074] Other implementations of the disclosed technology / present inventions will be apparent to those skilled in the art from consideration of the specification and practice of the innovations disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the inventions being indicated by the present disclosure and various associated principles of related patent doctrine.
Claims
CLAIMS;1. A computer-implemented method for transforming brain activity corresponding to an imagined sentence into a machine-readable output, the method comprising: a. performing processing to define, for a user, a finite vocabulary of sentences or phrases; b. prompting the user, such as, preferably visually, aurally, or otherwise, to silently imagine one selected sentence from the vocabulary of sentences; c. acquiring, via at least one non-invasive device that utilizes at least one neuro-sensing modality, time-series neural data that temporally overlap the user’s imagination of the sentence; d. inputting the neural data to a decoder implemented by one or more machine-learning models and generating a digital representation that classifies the neural data as the corresponding sentence; and e. outputting the classified sentence (i) as textual data and / or (ii) as a control or command signal to a downstream computing system, the downstream system optionally including a large language model (LLM) or other artificial-intelligence (Al) agent capable of responding to the sentence.
2. The method of claim 1 or the invention of any claim herein, wherein the neuro-sensing modality7comprises functional near-infrared spectroscopy (fNIRS).
3. The method of claim 1, wherein the neuro-sensing modality comprises magnetoencephalography (MEG), electroencephalography (EEG), functional magnetic- resonance imaging (fMRI), and high-density' diffuse optical tomography (HD-DOT).
4. The method of any preceding claim or the invention of any claim herein, further comprising training the decoder by: a. collecting paired sets of neural data and ground-truth sentences from one or more users; and b. optimising a machine-learning architecture comprising one or more of: convolutional neural networks, recurrent or transformer-based networks, tree-based ensembles, regression models, and / or hybrids thereof.
5. The method of any preceding claim or the invention of any claim herein, further comprising of detecting a transition with a secondary classifier between a rest condition and an imagination condition; wherein the transition signal automatically triggering steps acquiring time-series neural data that temporally overlap the user’s imagination of the sentence;inputting the neural data to a decoder implemented by one or more machine-learning models and generating a digital representation that classifies the neural data as the corresponding sentence; and outputting the classified sentence to a downstream computing system.
6. The method of any preceding claim or the invention of any claim herein, wherein the decoder operates in real-time, near real-time, and non-real-time (e.g., in one example, end-to- end latency < 30 s including imagination period, etc.).
7. The method of any preceding claim or the invention of any claim herein, wherein, outputting the classified sentence to a downstream computing system as a natural-language prompt to an LLM that trigger a textual response fonn the LLM back to the user.
8. The method of any preceding claim or the invention of any claim herein, wherein, outputting the classified sentence to a downstream computing system, the classified sentence is mapped to a discrete control code to an external device and / or software application.
9. The method of any preceding claim or the invention of any claim herein, w herein the the decoder automatically selects or blends neural data derived from different preprocessing levels that comprises: raw intensity, optical-density, or fully pre-processed haemodynamic signals.
10. The method of any preceding claim or the invention of any claim herein, wherein the neuro-sensing modality comprises a continue us- wave, high-density fNIRS system providing whole-head coverage, such as. in one illustrative example, with at least 350 channels and a sampling rate of 5 Hz or greater.
11. The method of any preceding claim or the invention of any claim herein, wherein the fNIRS system includes a 48-source, 47-detector array with dedicated short-separation channels for systemic-noise regression.
12. The method of any preceding claim or the invention of any claim herein, wherein neural data are streamed to the decoder over a network interface using a real-time data format and stored to disc in a self-describing container for off-line analysis.
13. The method of any preceding claim or the invention of any claim herein, wherein the decoder’s confidence score is monitored and, when below a threshold, the output is withheld or flagged for operator review.
14. The method of any preceding claim or the invention of any claim herein, implemented by a wearable or portable neuro-sensing apparatus weighing less than 3 kg.
15. A system comprising: one or more computer processors;at least one communication interface operatively coupled to the one or more computer processors, wherein the at least one communication interface transmits a classified sentence or control signal to an Al agent or external device; one or more non-transitory computer readable media, the non-transitory computer readable media including computer-readable program instructions that, upon execution by the one or more computer processors, cause the one or more computer processors to perform operations including: one or more of steps c-e of claim 1.
16. The system of claim 15 or the invention of any claim herein, wherein, as a function of the at least one communication interface transmitting the classified sentence or control signal to the Al agent or external device, a direct thought-to-machine communication channel is established.
17. The system of claim 15 or the invention of any claim herein, further comprising: at least one neuro-sensing apparatus, such as, preferably, a neuro-sensing apparatus configured to perform processing and / or features of any of claims 2, 3, 10 or 11.
18. A system comprising: at least one computer processor; one or more non-transitory computer readable media, the non-transitory' computer readable media including computer-readable program instructions that, upon execution by the at least one computer processor, cause the at least one computer processor to perform operations including: one or more aspects, steps, features, and / or functionality' recited in any claim here or set forth elsewhere in the present disclosure.
19. A computer-implemented method of processing neural data signals derived from brain activity corresponding to at least one imagined sentence, the method comprising: prompting a user to imagine a preferred sentence from a predefined set of options; providing the neural data signals associated with the preferred sentence imagined by the user to one or more processing and / or decoding components; processing, via the one or more processing and / or decoding components, the neural data signals to produce decoded neural data; performing classification of the decoded neural data into sentences; performing processing of the decoded neural data and / or the sentences to provide a machine-readable output associated with the classification into the sentences.
20. The method of claim 19 or the invention of any claim herein, further comprising:performing processing of the decoded neural data and / or the sentences to provide an answer, the performing processing including one or more of: transforming the neural data into prompts; transmitting the prompts to a large language model; and / or generating, via the large language model, the answer based on one or more of the prompts.
21. A computer-implemented method of processing and / or decoding imagined sentences and / or discrete language via one or more non-invasive brain / neural interfaces, the method comprising: providing a predefined set of options to a user prompting the user to imagine a preferred sentence from the predefined set of options; providing a plurality of neural data of the user detected from the one or more non- invasive brain / neural interfaces to one or more processing components; processing, via a decoder, the neural data to produce decoded neural data; classifying the decoded neural data into sentences; performing processing of the decoded neural data and / or the sentences to provide and answer, the performing processing including one or more of: transforming the neural data into prompts; transmitting the prompts to a large language model; and / or generating, via the large language model, the answer based on one or more of the prompts.
22. The method of any claim above or the invention of any claim herein, wherein the predefined set of options are provided to the user via at least one screen.
23. The method of any claim above or the invention of any claim herein, wherein neural data is decoded by the Al model to classify the correct sentence.
24. The method of any claim above or the invention of any claim herein, wherein the correct sentence has been classified, the correct sentence can be sent to the large language model as a prompt.
25. The method of claim 21 wherein the neural data comprises magnetoencephalography (MEG) data.
26. The method of any claim above or the invention of any claim herein, further comprising: utilizing the at least one classified sentence as a command signal.
27. The method of any claim above or the invention of any claim herein, further comprising: decoding sentences as prompts to a ChatGPT and / or Al instance; and sending the classified sentence being sent as a prompt to a software application running ChatGPT.
28. The method of claim 11 wherein the neural data comprises functional near-infrared spectroscopy (fNIRS) data.
29. The method of any of the claims, or the invention of any claim herein, further comprising: developing an Al model utilizing brain and / or mind data and / or brain imaging data, wherein the Al model is capable of decoding imagined speech and / or enhance human- Al communication.
30. The method of any claim above or the invention of any claim herein, wherein the brain and / or mind data and / or brain imaging data is acquired via at least one fMRI modality.
31. The method of any claim above or the invention of any claim herein, further comprising: implementing an extra Trees Classifier (XTC) model to decode neural patterns and / or differentiate imagined speech from rest conditions.
32. The method of any claim above or the invention of any claim herein, wherein the processing, via the decoder, of the neural data includes utilizing a convolutional neural netw ork (CNN) and a ridge regression model to perform decoding of neural signals associated with specific imagined sentences.
33. The method of any claim above or the invention of any claim herein, further comprising: utilizing a real-time Al communication system using brain data, such as, optionally, obtained via fNIRS technology, and / or using suitable application software, such as, optionally, a Web framework or micro Web framework, such as a Flask application, enabling early-stage thought-based communication, such as via a direct communication channel, between participants and a large language model (LLM).
34. The method of any claim above or the invention of any claim herein, further comprising: utilizing a near real-time Al communication system using brain data, such as, optionally, obtained via fNIRS technology, and / or using suitable application software, such as, optionally, a Web framework or micro Web framework, such as a Flask application, enabling early-stage thought-based communication, such as via a direct communication channel, between participants and a large language model (LLM).
35. The method of any claim above or the invention of any claim herein, further comprising: utilizing a non- real-time Al communication system using brain data, such as, optionally, obtained via fNIRS technology, and / or using suitable application software, such as.optionally, a Web framework or micro Web framework, such as a Flask application, enabling early-stage thought-based communication, such as via a direct communication channel, between participants and a large language model (LLM).
36. The method of any claim above or the invention of any claim herein, further comprising: implementing combined processing that includes both: performing advanced signal processing techniques as a function of processing signals received from at least one fNIRS-based system; and performing processing, via one or more machine learning algorithms, to classify imagined sentences with different semantic content from a predefined set, based on the analysis of brain activity patterns.
37. The method of any claim above or the invention of any claim herein, further comprising: collecting high signal-to-noise ratio (SNR) signals via at least one fNIRS-based system during imagined speech; and comparing the high SNR signals compared to rest signals associated with rest brain function.
38. The method of any claim above or the invention of any claim herein, further comprising: implementing decoding algorithms that decipher and classify imagined sentences from a limited dictionary'.
39. The method of any claim above or the invention of any claim herein, further comprising: establishing a thought-based communication channel as a platform for human-AI interaction, such as, optionally, for imagined speech decoding.
40. The method of any claim above or the invention of any claim herein, further comprising: utilizing a Continuous -Wave (CW) high-density fNIRS system, such as, optionally a48x48 fNIRS system, that provides full head coverage, wherein, optionally, the fNIRS system may include one or more of:48 sources and 47 detectors, wherein the extra detector is used for the short-distance channels; a total of between about 360 channels and about 416 channels, such as about 388 channels (e.g., 194 source wavelength at 760 nm and 194 source wavelength at 850 nm, etc); a sampling rate of between about 5.5 Hz and about 6.3 Hz, such as, optionally, 5.9 Hz; channel distances ranging from ~21 mm to ~42 mm); and / or8 short-distance channels (channel distances: < 10 mm) providing high-density fullhead coverage, to monitor changes in oxygenated blood levels in the brain as a proxy for neural activity, such as that of Figure 1.
41. The method of any claim above or the invention of any claim herein, further comprising: utilizing a CNN to extract features from minimally preprocessed data; utilizing a higher level of preprocessing for the minimally preprocessed data.
42. The method of claim 13 or the invention of any claim herein, wherein raw or minimally preprocessed data (e.g., optical density, etc.) achieves higher performance when selecting the CNN decoder, while fully preprocessed data may be processed via ridge regression, or other such simpler model(s).
43. The method of any preceding claim or the invention of any claim herein, further comprising: streaming the brain data collected during sentence imagining through NIRx acquisition software; and saving the data streamed through the NIRx acquisition software in the XDF format; wherein the saved data may then be loaded and subjected to a series of preprocessing techniques.
44. The method of any preceding claim or the invention of any claim herein, further comprising: training models using (a) raw data (e.g., intensity only), (b) optical density data (e.g.. minimally preprocessed data which underwent conversion of raw signals to optical density data), and / or (c) fully preprocessed data (e.g., data that underwent conversion of raw signals to optical density, detrending, short channel regression correction, motion artefact correction, conversion to haemoglobin concentration using a partial pathlength factor (ppf) of 6. and bandpass filtering between 0.01 and 0.7 Hz); and trimming the data to a fixed shape (e.g., 145 time points by 388 channels and / or 194 oxy -haemoglobin and 194 deoxy -haemoglobin channels) and saved as HF5 files, with the filenames corresponding to the respective sentence names.
45. The method of any preceding claim or the invention of any claim herein, wherein the brain data is acquired via a fNIRS sensor electrode system that includes 48 sources and 47 detectors, wherein the extra detector is used for the short-distance channels; and wherein the system consists of a total of 388 channels (194 source wavelength at 760 nm and 194 source wavelength at 850 nm); sampling rate: 5.9 Hz; channel distances ranging from ~21 mm to ~42 mm), and 8 short-distance channels (channel distances: < 10 mm)providing high-density full-head coverage, to monitor changes in oxygenated blood levels in the brain as a proxy for neural activity.
46. A method of directly decoding sentences, such as full sentences, from signals obtained from a user’s mind / brain via one or more non-invasive brain / neural interfaces, the method comprising: providing a predefined set of options to a user prompting the user to imagine a preferred sentence from the predefined set of options; providing a plurality of neural data of the user detected from the one or more non- invasive brain / neural interfaces to one or more processing components; processing, via a decoder, the neural data to produce decoded neural data; classifying the decoded neural data into sentences; performing processing of the decoded neural data and / or the sentences to provide and answer, the performing processing including one or more of: transforming the neural data into prompts; transmitting the prompts to a large language model; and / or generating, via the large language model, the answer based on one or more of the prompts.
47. One or more non-transitory computer readable media including computer-readable program instructions that, upon execution by one or more computer processors, cause the one or more computer processors to perform operations including: one or more aspects, steps, features, and / or functionality recited in any claim here or set forth elsewhere in the present disclosure.