Systems and methods for neural signal decoding and classification

WO2026107504A1PCT designated stage Publication Date: 2026-05-21BOARD OF RGT THE UNIV OF TEXAS SYST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BOARD OF RGT THE UNIV OF TEXAS SYST
Filing Date
2025-11-18
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Conventional Brain-Computer Interfaces (BCIs) lack intuitiveness and have slow communication rates, failing to directly align with users' intended communication due to limitations in neural signal processing and classification.

Method used

Utilizing neuromagnetic sensors, such as MEG sensors with optically-pump magnetometers, to generate spatiotemporal maps like Mel-spectrograms, processed by trained AI models for predictive speech classification, enabling decoding of spontaneous neural speech and biometric authentication, and detecting neurological disorders like ALS.

Benefits of technology

Achieves high accuracy in decoding spontaneous neural speech, biometric authentication, and detecting neurological disorders, providing intuitive and efficient communication solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025055946_21052026_PF_FP_ABST
    Figure US2025055946_21052026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein are systems and methods related to neural signal processing and classification, wherein the method includes: receiving a plurality of neural signals corresponding to spontaneous neural speech of a subject obtained from an array of neuromagnetic sensors; generating a spatiotemporal map from the plurality of neural signals; and determining a predictive speech classification from the spatiotemporal map using a trained artificial intelligence (Al) model.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. 10046-617W018428 WAN SYSTEMS AND METHODS FOR NEURAL SIGNAL DECODING AND CLASSIFICATIONCross-Reference to Related Applications

[0001] This Application claims the benefit of U. S. Provisional Patent Application Ser. No. 63 / 722,062, filed November 18, 2024, which is hereby incorporated by reference in its entirety.Government License Rights

[0002] This invention was made with government support under grant no. R01 DC016621 awarded by the National Institutes of Health. The government has certain rights in the invention.Background

[0003] Brain-Computer Interfaces (BCIs) have recently emerged as a way to facilitate communication between a brain and external devices. For example, BCIs can provide non-muscular means of communication for individuals experiencing difficulties in interacting with the external environment due to physical disabilities or neurodegenerative diseases, such as amyotrophic lateral sclerosis (ALS). In tliis regard, BCIs can interpret the user’s intent by¬ analyzing brain activity induced by specific tasks, including speech imagination / in tention, motor imagery, and other tasks that evoke certain responses. However, these conventional BCI paradigms often lack intuitiveness, may not align directly with the user’s intended communication, and are limited by slow communication rates.

[0004] Thus, there is a benefit to providing improved systems and method for neural signal processing and classification.Summary

[0005] In an aspect, a method for decoding neural speech is disclosed. In some aspects, the method includes: receiving a plurality of neural signals corresponding to spontaneous neural speech of a subject obtained from an array of neuromagnetic sensors; generating a spatiotemporal map from the plurality of neural signals; and determining a predictive speech classification from the spatiotemporal map using a trained artificial intelligence (Al) model.

[0006] In some aspects, the neuromagnetic sensors comprise magnetoencephalography (MEG) sensors.Attorney Docket No. 10046-617W018428 WAN

[0007] In some aspects, the MEG sensors comprise optically-pump magnetometers (OPM).

[0008] In some aspects, the spatiotemporal map comprises a plurality of Mel-spectrograms.

[0009] In some aspects, the spontaneous neural speech comprises uncued speech.

[0010] In some aspects, trained Al model is configured to determine the predictive speech classification from a closed-set of speech units.

[0011] In some aspects, the trained Al model is configured to determine an open-set of speech units from the spatiotemporal map.

[0012] In some aspects, the trained Al model comprises a vocoder configured to generate an audio signal from the spatiotemporal map.

[0013] In another aspect, a method for biometric authentication is disclosed. In some aspects, the method includes: recording a neural signal of a subject from a neuroimaging sensor, wherein the neural signal includes features corresponding to neural activity of the subject performing a cogniti ve task; determining an authentication classification using a trained artificial intelligence (Al) model, wherein the Al model is configured to receive the neural signal and predict an identity of the subject.

[0014] In some aspects, the neuroimaging sensor comprises a non-invasive sensor.

[0015] In some aspects, the neuroimaging sensor comprises a neuromagnetic sensor.

[0016] In some aspects, the neuroimaging sensor comprises a magnetoencephalography (MEG) sensor.

[0017] In some aspects, the MEG sensor comprises an optically-pump magnetometer (OPM).

[0018] In some aspects, the neuroimaging sensor comprises a plurality of sensors.

[0019] In some aspects, the cognitive task comprises overt speech. In some aspects, the cognitive task comprises imagined speech.

[0020] In some aspects, the method further includes prompting the subject to perform the cognitive task prior to recording the neural signal.

[0021] In some aspects, the method includes unlocking a device based on the authentication classification.

[0022] In another aspect, described herein are methods for detecting a neurological disorder in a subject. In various aspects, the method includes: recording a plurality of neural signals of a subject from an array of neuromagnetic sensors, wherein the neural signalsAttorney Docket No. 10046-617W018428 WAN includes features corresponding to neural activity of the subject performing a cognitive task; and determining an disease state of the subject using a trained artificial intelligence (Al) model, wherein the Al model is configured to receive the plurality of neural signals and predict an disease state of the subject based a cortical feature for the neurological disorder.

[0023] In some aspects, the neurological disorder is amyotrophic lateral sclerosis (ALS).

[0024] In some aspects, the cortical feature comprises beta band oscillations.

[0025] In some aspects, the cognitive task comprises overt speech.

[0026] In some aspects, the cognitive task comprises imagined speech.

[0027] In some aspects, the neuromagnetic sensors comprise magnetoencephalography (MEG) sensors.

[0028] In some aspects, the MEG sensors comprise optically-pump magnetometers (0PM).

[0029] In some aspects, the detection of the neurological disorder is a single-trial classification.

[0030] In another aspect, a system is disclosed, having one or more processors; and a memory having instructions stored thereon, wherein the instructions, when executed by the one or more processors, cause the one or more processors to perform any one of the above¬ discussed methods.

[0031] In another aspect, a non -transitory computer-readable medium is disclosed, having instructions stored thereon, wherein the instructions, when executed by one or more processors, cause the one or more processors to perform any one of the above-discussed methods.

[0032] Additional advantages of the invention will be set forth in part in the description which follows, and in part will be clear from the description or may be learned by practice of the invention. The advantages of the invention will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims. It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention, as claimed.Brief Description of the Drawings

[0033] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments and, together with the description, serve to explain the principles of the methods and systems. The patent or application file contains at least one drawing executed in color.Attorney Docket No. 10046-617W018428 WAN

[0034] Fig. 1 depicts a diagram of an example system configured with a neural network model for neural signal processing and classification for decoding spontaneous neural speech in accordance with an illustrative embodiment.

[0035] Fig. 2 depicts a diagram of an example system configured with a neural network model for biometric authentication in accordance with an illustrative embodiment.

[0036] Fig. 3 shows an example method for neural signal processing and classification for decoding spontaneous neural speech.

[0037] Fig. 4 shows an example method for neural signal processing and classification for biometric authentication.

[0038] Fig. 5 shows a classification matrix of the identification of each individual using an exemplary Al model.

[0039] Fig. 6 shows neural features of the individuals using t-SNE (t-distributed Stochastic Neighbor Embedding), which is an unsupervised non-linear dimensionality reduction technique for data exploration and visualizing high-dimensional data.

[0040] Fig. 7 shows accuracy of individual sensors in an array of 2.04 MEG sensors. As combinations of sensors are used, the overall accuracy can be increased.

[0041] Fig. 8. Data summary. (A) Number of spontaneous speech samples (yes or no) uttered by each participant. (B) The distribution of yes or no speech duration across all participants is illustrated in the figure below. Notably, there was no statistically significant difference observed between the durations of yes and no overt speech.

[0042] Fig. 9. Pipeline for spontaneous speech decoding. Magnetoencephalography signals were epoched to yes or no segments corresponding to the yes or no speech. After a baseline linear discriminant analysis decoder training for yes or no classification, a one¬ dimensional (ID) convolutional neural network model was trained that consisted of two blocks of ID convolution layers and Rectified Linear Unit (ReLU) layers, followed by global average max pooling, fully connected layers, and finally a softmax layer that resulted the probability of a given sample belonging to class yes or no. MEG = magnetoencephalography; CNN = convolutional neural network; Conv = convolution; MaxP = Max Pooling; FC = Fully Connected.

[0043] Fig. 10. Spontaneous overt speech decoding. (A) Comparison of spontaneous overt speech decoding performance between baseline EDA decoder and one-dimensional convolutional neural network (ID CNN). Note that ID CNN decoder achieved about 90% decoding accuracy, significantly higher than chance level and baseline decoder performance.Attorney Docket No. 10046-617W018428 WAN (B) Decoding performances for individual participants are shown. Note that even the participant with least decoding accuracy also showed higher than chance level performance. LDA = linear discriminant analysis. *p <.001.

[0044] Fig. 11. Spontaneous intended speech decoding. (A) Magnetoencephalography data prior to the acoustic onset (intended speech segment), from acoustic onset to acoustic offset (overt speech segment), and after acoustic offset (postspeech segment) were used to decode spontaneous speech. (B) Decoding performance with one-dimensional convolutional neural network resulted in about 67% decoding accuracy for intended speech significantly above chance level. Decoding accuracies for overt speech and postspeech was higher than intended speech.

[0045] Fig. 12 shows a timing sequence of a single trial.

[0046] Fig. 13. Architecture of the proposed Mel generator according to one example,

[0047] Fig. 14. Examples of Mel spectrograms generated using the proposed methods.

[0048] Fig. 15. Average Pearson Correlation Coefficients (a) and Root Mean Squared Errors (b).

[0049] Fig. 16. Examples of synthesized audio signals and their pitch curves from participant A3.

[0050] Fig. 17. Histogram distribution of all pair-wise sensor correlations of healthy and ALS participants for each phrase during imagined speech (top) and overt speech (articulation) (bottom) for patients with ALS and healthy controls, respectively.

[0051] Fig. 18. Heatmap of pair- wise sensor correlations for each subject on phrase “Do you understand me” for imagination (top) and articulation (middle), and the distribution of sensor density across the phrase for all participants (bottom). In the heatmaps, each colorful dot / point represents the correlation between a pair of the 196 gradiometers. In the bottom panel, correlation density was calculated as the sum of all absolute correlation values over total number of sensor-pairs across five phrases for each participant.

[0052] Fig. 19. Heatmap of pairwise band-power distances between healthy and ALS sensors for each band during articulation. (A) Delta; (B) Theta; (C) Alpha; (D) Beta; (E) Gamma; (F) High Gamma. In the heatmaps, each colorful dot / point represents the bandpower distance between a pair of the 196 gradiometers. The bottom panels provide the bandpower distances for all frequency bands in speech imagination (G) and articulation task (H), respectively.Attorney Docket No. 10046-617W018428 WAN

[0053] Fig. 20. Heatmap of beta band AEC based functional connectivity across all sensors for participants with ALS (top row) and healthy controls (bottom row) during the overt speech task.

[0054] Fig. 21. Single-trial ALS detection accuracy using band power [Delta: 1-4 Hz; Theta: 4-8 Hz; Alpha: 8—16 Hz; Beta: 16-30 Hz; Gamma: 30-59 Hz; High Gamma: 61- 119 Hz; Broadband: 1-119 Hz],Detailed Description

[0055] Each and every feature described herein, and each and every combination of two or more of such features, is included within the scope of the present invention, provided that the features included in such a combination are not mutually inconsistent.

[0056] Definitions

[0057] As used herein, the term “subject” refers to any animal (e.g., a mammal), including, but not limited to, humans, non -human primates, rodents, and the like.

[0058] All patents and publications mentioned in the specification are indicative of the levels of skill of those skilled in the art to which the invention pertains. References cited herein are incorporated by reference herein in their entirety to indicate the state of the art as of their filing date, and it is intended that this information can be employed herein, if needed, to exclude specific embodiments that are in the prior art.

[0059] Example System

[0060] Fig. 1 shows a diagram of an example system 100 configured with a trained artificial intelligence (Al) model (e.g., a neural network) for decoding neural speech.

[0061] In the example shown in Fig. 1, the system 100 includes an addressable array of neuromagnetic sensors 102 configured to record a plurality of neural signals 104 corresponding to spontaneous neural speech of a subject. In some aspects, the spontaneous neural speech comprises uncued speech. “Uncued speech” refers to speech that occurs without direction from a particular prompt or cue.

[0062] In some aspects, the neuromagnetic sensors comprise magnetoencephalography (MEG) sensors. In some aspects, the MEG sensors comprise optically-pump magnetometers (0PM). Exemplary OPMs suitable for the present system include those described in Boto, et al. Nature, 2018; Borna, et al., PLoS One, 2020; Kominis, et al., Nature, 2003; Zhang et al., Science, 2020, each of which is incorporated by reference in its entirety.

[0063] In some aspects, the neuromagnetic sensors are configured as wearable device. Typically, MEG sensors utilizes superconducting quantum interference devices (SQUIDs),Attorney Docket No. 10046-617W018428 WAN which generally involves cryogenic cooling. Wearable OPM-MEG systems, such as those developed by Cerca Magnetics, LLC may also be used for the present systems and methods.

[0064] The neural signals 104 are used to form a spatiotemporal map 106 including a plurality of spectrograms 108, each denoting neural information from a respective neuromagnetic sensor location. In some aspects, the spatiotemporal map comprises a plurality of Mel-spectrograms. The term “Mel spectrogram” refers to a spectrogram that has been converted to a non-linear Mel scale.

[0065] The spatiotemporal map 106 is provided to a trained neural network 112, where a predictive speech classification 114 can be generated. The predictive speech classification 114 can include, for example, a binary set of units. In some aspects, trained neural network is configured to determine the predictive speech classification 114 from a closed-set of speech units. In some aspects, the trained neural network is configured to determine an open-set of speech units from the spatiotemporal map.

[0066] In some aspects, the trained neural network includes a vocoder configured to generate an audio signal from the spatiotemporal map. An example of a vocoder suitable for the present systems and methods is the BigVGAN neural vocoder model described in S. Lee et al., “BigVGAN: A Universal Neural Vocoder with Large-Scale Training,” in Proc.International Conference on Learning Representations, 2023,” which is herby incorporated by reference in its entirety.

[0067] Fig. 2 shows another diagram of an example system 200 configured with a trained artificial intelligence (Al) model (e.g., a neural network) for biometric authentication in accordance with an illustrative embodiment.

[0068] In the example shown in Fig. 2, the system 200 includes an array of neuromagnetic sensors 202 configured to record a plurality of neural signals 204 corresponding to neural activity of a subject performing a predefined cognitive task. Neural signals can be obtained using any invasive (e.g., ECoG, sEEG, Utah array) or non-invasive (MEG, EEG, fNIR, fMRI) neuroimaging techniques while the person performs any “cognition” tasks (e.g., imaging speaking a short phrase, overtly say something, move fingers, think about a simple match question). Although the system 200 of Fig. 2 includes a plurality of sensors, bio-authentication systems having only a single sensor may also be used. The biometric neural signals 204 are provided to a trained neural network 212, where an authentication classification 214 can be generated. In some aspects, the neural signals 204 are used to form spatiotemporal map to be provided as an input to the trained neural network.Attorney Docket No. 10046-617WO18428 WAN The trained neural network 212 is configured to determine a probability that the neural signal 204 is associated with an authorized individual performing the respective cognitive task. In some aspects, the cognitive task comprises overt speech. In some aspects, the cognitive task comprises imagined speech.

[0069] If the neural network 212 determines that a confidence level exceeds a predefined threshold (e.g., 60% or greater, 65% or greater, 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, 99.9% or greater), an authentication classification confirming the identity of the subject can be issued. The authentication classification 214 can be provided to enable access to a device (or system) 220.

[0070] In some aspects, the neuromagnetic sensors comprise magnetoencephalography (MEG) sensors. In some aspects, the MEG sensors comprise optically-pump magnetometers (0PM). In some aspects, the neuromagnetic sensors are incorporated into a portable device.

[0071] The exemplary system and method can be based on any type of machine learning or artificial intelligence system, including, and not limited to, convolutional neural networks, autoencoders, recombinant neural networks, and long-term short-term memory (I, STM) networks. The system 100, 200 can be implemented via a processing unit (e.g., processor) configured to execute computer-readable instructions. In its most basic configuration, the processing unit includes at least one processing circuit (e.g., core) and system memory. Depending on the exact configuration and type of computing device, system memory may be volatile (such as random-access memory (RAM)), non-volatile (such as read-only memory (ROM), flash memory, etc.), or some combination of the two. The processing unit may be a standard programmable processor that performs arithmetic and logic operations necessary for the operation of the computing device. As used herein, processing unit and processor refers to a physical hardware device that executes encoded instructions for performing functions on inputs and creating outputs, including, for example, but not limited to, microprocessors (MCUs), microcontrollers, graphical processing units (GPUs), and application-specific circuits (ASICs).

[0072] While instructions may be discussed as executed by a processor, the instructions may be executed simultaneously, serially, or otherwise executed by one or multiple processors. The processing unit may also include a bus or other communication mechanism for communicating information among various components of the computing device. In someAttorney Docket No. 10046-617W018428 WAN embodiments, the processing unit is configured with co-processors (e.g., FPGA, ASIC) or Al-processors.

[0073] In some aspects, values determined according to the described systems and methods (e.g., the predictive speech classification) can be further outputted as a report. As used herein, the term “report” is intended to describe a presentation of multidimensional data. A report can include various forms of data (e.g., text, audio, image) and may be considered as an analysis tool that can be used to view, manipulate, and print data.

[0074] In some embodiments, the report is displayed on a graphical user interface for clinical or informative review. The term “graphical user interface,” or GUI, may be used in the singular or the plural to describe one or more graphical user interfaces and each of the displays of a particular graphical user interface. Therefore, a GUI may represent any graphical user interface, including but not limited to, a web browser, a touch screen, or a command line interface (CLI) that processes information and efficientlypresents the information results to the user. In general, a GUI may include a plurality of user interface (UI) elements, some or all associated with a web browser, such as interactive fields, pull-down lists, and buttons operable by the business suite user. These and other UI elements may be related to or represent the functions of the web browser.

[0075] Example Method

[0076] Fig. 3 shows a method 300 for decoding neural speech according to one implementation. Method 300 includes: receiving (302) a plurality of neural signals corresponding to spontaneous neural speech of a subject obtained from an array of neuromagnetic sensors. In some aspects, the neuromagnetic sensors comprise magnetoencephalography (MEG) sensors. In some aspects, the MEG sensors comprise optically-pump magnetometers (0PM).

[0077] Method 300 also includes generating (304) a spatiotemporal map from the plurality of neural signals.

[0078] Method 300 further includes determining (306) a predictive speech classification from the spatiotemporal map using a trained artificial intelligence (Al) model.

[0079] In some aspects, the spatiotemporal map comprises a plurality of Mel-spectrograms.

[0080] In some aspects, the spontaneous neural speech comprises uncued speech.Attorney Docket No. 10046-617W018428 WAN

[0081] In some aspects, trained Al model is configured to determine the predictive speech classification from a closed-set of speech units.

[0082] In some aspects, the trained Al model is configured to determine an open-set of speech units from the spatiotemporal map.

[0083] In some aspects, the trained Al model comprises a vocoder configured to generate an audio signal from the spatiotemporal map.

[0084] Fig. 4 shows a method 400 for biometric authentication according to another implementation. Method 400 includes: recording a neural signal of a subject from a neuroimaging sensor. Neural signals can be obtained using any invasive (e.g., ECoG, sEEG, Utah array) or non-invasive (MEG, EEG, INIR, fMRI) neuroimaging techniques while the person performs any “cognition” tasks (e.g., imaging speaking a short phrase, overtly say something, move fingers, think about a simple match question). In some aspects, the neuroimaging sensor comprises a non-invasive sensor. In some aspects, the neuroimaging sensor comprises a neuromagnetic sensor. In some aspects, the neuroimaging sensor comprises a magnetoencephalography (MEG) sensor. In some aspects, the MEG sensor comprises an optically-pump magnetometer (OPM). In some aspects, the neuroimaging sensor comprises a plurality of sensors. The neural signal includes features corresponding to neural activity of the subject performing a cognitive task.

[0085] Method 400 further includes determining (404) an authentication classification using a trained artificial intelligence (Al) model, wherein the Al model is configured to receive the neural signal and predict an identity of the subject.

[0086] In some aspects, the method further includes prompting the subject to perform the cognitive task prior to recording the neural signal.

[0087] In the example shown in method 400, method 400 also includes unlocking (406) a device or system based on the authentication classification. Non-limiting examples of secure devices / systems that may be unlocked using the described methods include, for example, consumer electronics, security systems (e.g., smart locks, alarm systems, safes), automobiles, and banking / financial systems (e.g., ATMs, banking software). In some aspects, the device is a consumer electronic device (e.g., a desktop, a laptop, a smartphone, TVs).

[0088] In another aspect, disclosed herein is a method for detecting a neurological disorder in a subject. As used herein, the terms “neurological disorder,” or “neurodegenerative disorder” refer to any chronic or acute central nervous system (CNS) or peripheral nervous system (PNS) disease associated with neuronal or glial cell defectsAttorney Docket No. 10046-617W018428 WAN including, but not limited to, neuronal loss, neuronal degeneration, neuronal demyelination, gliosis (including macro- and micro-gliosis), or neuronal or extraneuronal accumulation of aberrant proteins or toxins (e.g., p-amyloid, or a-synuclein). Exemplary neurological disorders include but are not limited to amyotrophic lateral sclerosis (ALS), Parkinson's disease, Alzheimer's disease, multiple sclerosis (MS), and Huntington's disease. In some aspects, the neurological disorder is amyotrophic lateral sclerosis (ALS).

[0089] In various aspects, the method includes: recording a plurality of neural signals of a subject from an array of neuromagnetic sensors, wherein the neural signals includes features corresponding to neural activity of the subject performing a cognitive task; and determining an disease state of the subject using a trained artificial intelligence (Al) model, wherein the Al model is configured to receive the plurality of neural signals and predict an disease state of the subject based a cortical feature for the neurological disorder.

[0090] In some aspects, the cortical feature comprises beta band oscillations.

[0091] In some aspects, the cognitive task comprises overt speech. In some aspects, the cognitive task comprises imagined speech. In some aspects, the cognitive task includes both overt speech and imagined speech.

[0092] In some aspects, the neuromagnetic sensors comprise magnetoencephalography (MEG) sensors. In some aspects, the MEG sensors comprise optically-pump magnetometers (0PM). In some aspects, the detection of the neurological disorder is a single-trial classification.Experimental Results and Examples

[0093] Example #1: Neural Signals-Based User Identification

[0094] Brain activity signals are unique subject-specific biological features that cannot be forged or stolen. Recognizing this trait, the present study tested whether brain waves can be used as a far more secure, sensitive, and confidential biometric approach for user identification. By way of example, neural signal-based user identification can be used in any scenario that has high security requirement (e.g., banking, government, and military). In the future cyborg society, where people are connected to the internet or virtual reality / augmented reality in their daily life. Neural ID can serve as a digital ID in the digital world. Digital ID can be used to unlock devices (VR / AR devices, air fighters, or any devices with high security requirement), used as a person authorization in the cyberworld.Attorney Docket No. 10046-617W018428 WAN

[0095] The present study describes a system and method that use neural signals for security purpose. Neural signals can be obtained using any invasive (e.g., ECoG, sEEG, Utah array) or non-invasive (MEG, EEG, fNIR, fMRI) neuroimaging techniques while the person performs any “cognition” tasks (e.g., imaging speaking a short phrase, overtly say something, move fingers, think about a simple match question), or just in resting-state (relax and without intention to “think” something). In various implementations, the recording can be just a few seconds. The Al model can identify if the person is authorized / identified or not. The present results illustrate data collection using MEG.

[0096] As shown in Figure 5, a classification matrix of person identification on six subjects, providing a response to a particular phrase (e.g., how are you doing?). The overall accuracy was 100%.

[0097] Figure 6 shows neural features of these six persons using t-SNE ( t -distributed Stochastic Neighbor Embedding), which is an unsupervised non-linear dimensionality reduction technique for data exploration and visualizing high-dimensional data. After the high dimension (204) MEG data was reduced to two dimensions, the results clearly demonstrate distinct signals collected from each respective individual.

[0098] Figure 7 shows the accuracy using individual sensors. The study tested a total of 204 MEG sensors. While single sensors were shown to reach a high accuracy already (above 90%), combinations of multiple sensors increased the accuracy 100%. This results shows the high potential of using neural signals collected from a low cost device to distinguishing persons.

[0099] Example #2: Neural Decoding of Spontaneous Overt and Intended Speech

[0100] Eight healthy adult participants (five females, 22- 30 years of age) without any- known neurological conditions participated in the experiment. Participants were comfortably seated within the MEG unit with their arms resting on a table and overtly pronounced the words yes and no randomly at a rate of approximately one word per 4-5 s. Participants were explicitly instructed not to alternate back and forth between the two words but rather randomly sample between them. An example of a sequence would be yes, yes, no, no, no, yes, no, no, yes... no. An experimenter first demonstrated the task, and the participants practiced the randomization prior to starting the experiment. During the experiment, the researcher kept a running tally of each exemplar. The task condition ended when a minimum of 80 repetitions for each word was achieved (see Figure 8, panel A). The study also checked the trial length and observed that yes and no trials are of no significant difference in durationAttorney Docket No. 10046-617W018428 WAN (equivalence testing [Lakens et al., 2018]: p <.05, df = 1198, effect size = 0.16), which eliminated the potential duration effects in this study (see Figure 8, panel B).

[0101] Data Collection. Participants’ neuromagnefic brain activity was recorded during the task using a 306-channel (204 gradiometers and 102 magnetometers) MEGIN TRIUX (MEGIN LLC US) inside a two-layer magnetically shielded room. The MEG signals were recorded at a 4 kHz sampling rate with an online hardware filter of 0.03-1333 Hz. I' wo pairs of bipolar surface electrodes were used to record the electro-cardiogram and the electrooculogram (EOG) signals. Voice data were recorded into the MEG system’s analog to digital channels using a lapel microphone with the battery and transducer sitting outside the shielded room to avoid significant artifacts.

[0102] First, the MEG data were visually inspected for significant artifacts. Eye blinks and muscle artifacts were marked. The data were then low-pass filtered below 250 Hz with a fourth order noncausal Butterworth filter and notch filtered with cutoff frequencies at 60 Hz and its harmonics. One sensor that displayed a flat or noisy response was omitted from the subsequent analysis. Independent component analysis (ICA) was employed to eliminate artifacts arising from eye blinks, eye movements, and heartbeats. The effectiveness of the ICA procedure was validated by inspection of the data at the artifact markings, EOG, and ECG. During tliis processing, data from one participant contained continuous artifacts and were removed from the analysis leaving seven participants for analysis.

[0103] Second, to facilitate MEG data epoching, the speech signal was extracted and denoised using a Wiener filter (Plapous et al., 2006). Voice activity detection (VAD; Sohn et al., 1999) was implemented to pinpoint the acoustic onset and offset of the words. The duration of each speech event was attached to the VAD labels. These labels were manually inspected, and corrections were made to ensure accurate marking. Each trial was categorized into either the yes or no class after listening to the acoustic signals. Finally, the MEG data were epoched into trials around the speech signal. Each trial was then manually segmented into three segments that are of equivalent length, pre-speech (intended speech), (overt) speech, and postspeech. Here, the study defined overt speech period as the time segment during the speech signal, that is, from acoustic onset to acoustic offset, intended speech period as the time segment with the same length as that of overt speech segment before the acoustic onset and postspeech period as the time segment with the same length with that of overt speech segment after the acoustic offset. There was no overlap between postspeech of a previous trial and pre-speech of a current trial. The onset and offset of overt speech segmentsAttorney Docket No. 10046-617W018428 WAN were evaluated based on synchronously collected acoustic data. Only the MEG gradiometers were used for decoding analysis, as it has been shown that gradiometers have more robust information and outperform magnetometers (Dash, Ferrari, Babajani-Feremi, et al., 2021).

[0104] To optimize model training and evaluation, the study split the data allocating the first 70% of trials for training and the last 30% of the trials for testing to preserve the temporal order of the MEG time series data and transferability to online decoding. The training was performed for each participant using a fourfold cross-validation (CV) strategy to find the optimal model configuration with tuned hyper-parameters. The choice of fourfold CV was to ensure a minimum of 40 trials (80 total trials per class x 70% for training x 75% training data for fourfold CV, that is, three-fold for training and onefold for validation in each of the four iterations of fourfold CV) in the training fold based on previous research in determining optimal number of trials for decoding (Dash, Ferrari, Malik, Montillo, et al., 2018). To enhance robustness, the fourfold CV was repeated 10 times through Bayesian bootstrapping (Rubin, 1981). The performance metric was decoding accuracy, calculated as the ratio of correct predictions by the decoders to the total number of predictions. The C V performance was computed by averaging the accuracy across the fourfold and the 10 bootstrap runs. The trained decoder’s predictions on the test data were then reported as the final decoding performance. The study performed yes and no classification using linear discriminant analysis (EDA) and one-dimensional convolutional neural networks (ID CNN) for both intended and overt speech. Figure 9 illustrates the classification pipeline.

[0105] Initially, the study trained an EDA model as the baseline decoder. EDA, a supervised machine learning classifier, identifies directions (“linear discriminants”) that maximize the separation between multiple classes (McLachlan, 2005). In other past neural speech decoding studies, EDA demonstrated performance equivalent to support vector machines and multilayer perceptron classifiers (Dash, Ferrari, Babajani-Feremi, et al., 2021; Dash, Ferrari, & Wang, 2021) and outperformed naive Bayes, decision trees, ensembles, and k-nearest neighbor classifiers (Dash et al., 2020). Due to its relatively fast training and minimal hyperparameter tuning, LDA became the preferred base-line decoder. The study fine-tuned the classifier’s hyperparameters (alpha and beta parameters of the Dirichlet distribution) for each participant using Bayesian optimization that works by considering the previously seen hyperparameter combinations when determining the next set of hyperparameters to evaluate (Wu et al., 2019). Root mean square features extracted from MEG signals were employed to train the LDA decoder, given their proven effectiveness inAttorney Docket No. 10046-617WO18428 WAN MEG-and EEG -based decoding analyses (Dash, Ferrari, Babajani-Feremi, et al., 2021;Sereshkeh el al., 2017; L. Wang et al., 2020). The feature dimension was 203 corresponding to the 203 gradiometer sensors.

[0106] ID CNNs are particularly well suited for tasks involving sequential data, such as time series analysis, speech recognition, and natural language processing (Kiranyaz et al., 2019). ID CNNs consist of convolutional layers (filters or kernels to local regions of the input sequence, capturing paterns and features to learn hierarchical representations of the input data), pooling layers (downsample the spatial dimensions of the input, reducing the computational complexity of the model and aiding in the extraction of important features), a flatten layer (convert the high-dimensional output into a ID vector), which can then be fed into fully connected layers for further processing. Convolutional operations provide a degree of translation invariance, allowing the model to recognize patterns regardless of their position in the input sequence. The use of shared weights in convolutional layers enables the model to learn spatial hierarchies efficiently.

[0107] The study used a ID CNN that consisted of two blocks of ID convolution layer, a ReLU layer, and a normalization layer. The ID convolution layers had 32 and 64 filters of size 5. They were followed by a global average pooling layer and then a flatten layer and a fully connected layer of two nodes. The last layer was a softmax layer with two nodes depicting the output probability of the two classes (yes and no). An Adaptive Moment Estimation optimizer was used to train the model with initial learning rate (0.003-0.006: different for each participant) with 100 maximum number of epochs. A grid search strategy was used to tune the hyper-parameters of the model.

[0108] The decoding of spontaneous overt speech yielded an accuracy of 79.02% ± 3.75% with EDA and 90.40% ± 3.98% with ID CNN (see Figure 10, panel A). Both LDA and ID CNN decoding performances surpassed the chance level of 50% (LDA: t = 21.05, p <.001; ID CNN: t = 22.72, p <.001). Even the participant with the lowest performance (P003: LDA: 61.64%; ID CNN: 67.82%) achieved results significantly higher than chance level. Notably, this participant also had the fewest samples among all participants (refer to Figure 8, panel A). Furthermore, the average performance with ID CNN was significantly higher than that with baseline LDA (t =3.87; p =.008). These results also suggested CNN outperformed LDA in neural speech decoding. The study also found CNN outperformed other machine learning classifiers such as artificial neural network in prior studies (Dash et al., 2020a). Thus, CNN was applied only for intended speech decoding.Attorney Docket No. 10046-617WO18428 WAN

[0109] The mean decoding accuracy for intended speech was 67.19% ± 2.72% using ID CNN, significantly above chance level (t = 24.72, p <.001; Figure 11, panel A). Intended speech decoding performance was lower compared to decoding overt speech as well as decoding postspeech (i.e., with taking data after acoustic offset; one-way ANOVA: F = 9.72, p =.0014; post hoc Tukey tests: intended vs. overt-p -.0012; intended vs. post-p =.0217).

[0110] This study demonstrated the feasibility of decoding spontaneous speech with high accuracy using noninvasive neuromagnetic signals. Importantly, the results revealed that not only spontaneous overt speech but also spontaneous intended speech can be decoded from neural signals.

[0111] Utilizing a spontaneous speech paradigm removed potential contamination by cue-based interference such as task effects or perceptual processing and provided solid support toward being able to interpret genuine speech motor activity from the brain. Much of the existing literature on overt speech decoding involves tasks with auditory or visual stimuli that participants are required to recognize, process, and respond to such as picture naming (X.-J. Wang, 2010) or paragraph reading (Anumanchipalli et al., 2019). In such scenarios, it becomes challenging to discern whether the neural activity being decoded solely corresponds to overt speech processes or incorporates information related to stimulus recognition and processing, leading to ambiguity. Although some studies have addressed this concern by implementing delayed overt speech production protocols introducing a 1-3 s delay between stimulus presentation and speech production (Dash et al., 2020a), the lingering effects of postauditory or visual processing in the brain may persist into the overt speech production stage. In this study, participants spontaneously uttered yes or no without any explicit cue thereby avoiding these effects.

[0112] It is important to acknowledge that overt speech decoding can potentially be influenced by movement artifacts. This poses a twofold problem for decoding. First, high- frequency muscle artifacts corrupt the MEG sensor data, potentially masking neural signals and, secondly, the low-frequency noise components related to jaw and tongue movements potentially create spurious decoding. Importantly, in a previous study that incorporated movement tracking, it was shown that overt speech decoding from MEG signals is clearly a combination of both neural and movement signals (Dash, Ferrari, Malik, & Wang, 2018). Despite the challenges, this study achieved high performance in classifying overt “yes” and “no” speech samples. This success strongly suggests that spontaneous overt speech canAttorney Docket No. 10046-617W018428 WAN indeed be effectively decoded using non-invasive neural signals, showcasing the robustness of this approach.

[0113] When considering patients with locked-in syndrome, the focus shifts to decoding intended or imagined speech for speech-BCI applications. In studies utilizing INIRS (Chaudhary et al., 2.017; Gallegos- Ayala et al., 2014; Hwang et al., 2016; Rezazadeh Sereshkeh et al., 2019), yes or no intended or imagined speech decoding achieved 64%-75% performance. Similarly, with EEG (Balaji et al., 2017; Hashim et al., 2018; Rezazadeh Sereshkeh et al., 2017), yes or no classification resulted in 63%-73% decoding performance. Evidently, decoding intended or imagined speech poses challenges compared to overt speech decoding. In this study, a 67.19% accuracy was achieved in decoding intended speech, on par with current literature. Notably, the intended speech was spontaneous, adding an extra layer of difficulty in contrast to the previous literature that used visual or auditory stimuli for task presentation where cue-based task effects might contribute to decoding performance. This underscores the significance of these results in addressing the challenges associated with decoding spontaneous intended speech. It is noted that, although the spontaneous intended speech decoding avoids the contribution from cue -based task effects, motor artifacts present during the intended speech segment could have contributed to decoding performance.

[0114] This capability of decoding intended speech holds great promise for applications involving decoding intended speech, especially in cases such as ALS or for locked-in patients. The study showed remarkable performance during the overt speech production stage, where the decoder was trained exclusively with data captured between acoustic onset and offset. This stage encapsulates rich information, including speech motor articulation details, auditory feedback, and potential remainder movement artifacts. Although spontaneous intended speech decoding presents additional challenges, its necessity becomes apparent in the trajectory toward real-world speech-BCI applications.

[0115] Last and interestingly, the study also obtained an above-chancel level accuracy for postspeech segment, which provides additional evidence that speech, like other task-based information processes, is not transient but dynamically linked to ongoing neural processes (for up to a few seconds). This further validates the drawback / limitation of using cue-based task designs to develop speech BCIs. To summarize, the study demonstrated the feasibility of decoding spontaneous speech from noninvasive neural activity. Although challenging, spontaneous speech decoding is a step forward in decoding true speech motor activity from the brain, crucial for scalable speech-BCIs.Attorney Docket No. 10046-617W018428 WAN

[0116] Example #3: Direct Speech Synthesis from Non-invasive, Neuromagnetic Signals

[0117] In this study, the feasibility of direct speech synthesis from non-invasive neural signals was investigated. The study used MEG signals from participants while performing overt speech production tasks. The study adopted a two-stage approach. First, the study used a deep-learning-based Mel generator to convert neural activities into a log Mel spectrogram. The disclosed architecture of Mel generator for MEG features was proposed based on Squeezeeformer (Kim et al., 2022), an enhanced version of Transformer architecture known for its excellence in automatic speech recognition (ASR). Additionally, a bidirectional long short term memory (Bi-LSTM) layer was employed to process information from both past and future contexts simultaneously. For the second stage, the study employed BigV GAN, a cutting edge neural vocoder, to synthesize audible speech waveforms from the generated Mel spectrograms.

[0118] A total of five unique, commonly used English phrases were selected from phrase lists typically used in augmentative and alternative communication (ACC) devices. The phrases were: phrase 1: ‘Do you understand me?’, phrase 2: ‘That’s perfect’, phrase 3: ‘How are you?’, phrase 4: ‘Good-bye’, and phrase 5: ‘I need help’. All participants completed 100 trials for each phrase, with each trial comprising four stages: pre-stimulus (rest) of 0.5 seconds, stimulation of 1 second, preparation of 1 second, and production of 2 seconds, as illustrated in Figure 12. During the pre-stimulus stage, participants were at rest. In the stimulation stage, the target phrase was presented on the screen. The order of phrases was pseudo-randomized to mitigate any response suppression due to repeated exposure. The preparation stage began with a fixation cross to cue participants to prepare to produce the designated phrase. When the fixation cross disappeared, the participants were instructed to articulate the phrase aloud at their natural speaking rate. STIM2 software (Neuroscan, Compumedics Limited, Australia) was used to generate visual stimuli, which were presented via a DIP projector positioned 90 cm away from the participants. Trials that contained high amplitude artifacts or non-compliance with the paradigm, such as reading the phrase before the cue, were excluded following a thorough visual inspection. Furthermore, sensors exhibiting high channel noise and artifacts, or other irregularities were excluded for further analysis.

[0119] Neuromagnetic signals were recorded using a 306-chaiinel SQUID MEG device (MEGIN OY, Espoo, Finland), consisting of 204 planar gradiometers and 102 magnetometerAttorney Docket No. 10046-617W018428 WAN sensors, situated in a magnetically shielded room (MSR) to eliminate unwanted environmental magnetic field interferences. Two pairs of bipolar electrodes were used to record the electrocardiogram (ECG) and electrooculogram (EOG) signals, respectively. The audio signals were recorded, using a standard built-in microphone connected to a transducer outside the MSR, and the recorded analog speech signal was digitized by feeding into the MEG analog-to-digital converter (ADC) in real-time as a separate channel. All data were recorded synchronously at the sampling rate of 3-5 kHz (including the audio data). Clinical MEGs typically support only up to 5khz sampling rate. The neural data were processed with an online band-pass filter of 0.1 to 1300 Hz.

[0120] The raw signals were segmented into epochs from -0.25 to 5 s, relative to stimulus onset time (0 s), and then baseline corrected by subtracting the average value between -0.25 s to 0 s for each channel. Here, the study segmented the data 1 s longer than the length of a single trial, 4 s, to mitigate the signal distortion caused by the filtering process.

[0121] The segmented audio signals were normalized to range from -1 to 1, followed by Wiener filtering to reduce the background noise. Subsequently, each trial was further segmented to span 200 ms before the voice onset and 200 ms after the voice offset, which were determined based on noise-reduced audio signals. The study used a pre-trained BigVGAN.24kHZ..100band model to reconstruct the audio signals from Melspectrograms. Therefore, the noise-reduced audio signals were converted into Mel-spectrograms following the same procedure and parameters used in the pre -trained vocoder. The noise reduced audio signals were upsampled to 24 kHz and scaled to range from -0.95 to 0.95. Then they converted to log scale Mel spectrogram with a window size and nFF'T of 1024, hop size of 256, 100 Mel filter banks, and a maximum frequency of 12000 Hz. These audio signals served as the target in this MEG-to-speech synthesis experiment.

[0122] In this study, only the signals from (204) gradiometers were utilized for further analysis due to their effectiveness in noise reduction and representation of the stimuli-based activation. The segmented MEG signals were epoched based on the noise -reduced audio signals and then downsampled to 500 Hz to lower computational costs after anti-aliasing filtering. Bandstop filtering was applied to eliminate power line noises with cut-off frequencies of 60, 120, and 180 Hz, followed by lowpass filtering at a cut-off frequency of 200 Hz, using 6thorder zero-phase Butterworth filters. The filtered signals were first normalized using the z-score normalization to reduce intertrial variability and then scaled to a range from -0.95 to 0.95, to ensure that the signals were on a similar scale.Attorney Docket No. 10046-617W018428 WAN

[0123] The study employed spectrograms as features of MEG signals. To match the lengths of MEG spectrograms to the Mel spectrograms of audio signals, window and hop sizes were set to 21 and 5, respectively. These settings were adjusted from the parameters used in BigVGAN considering the sampling rate of MEG signals. The nFFT was set to 64 considering the number of features. While using a higher nFFT can capture detailed information, it did not work for training the model in the PC environment due to the number of channels (about 200). When the number of sample points varied due to different sampling rates, zero padding was applied to equalize the lengths between the Mel spectrograms and spectrograms of MEG signals. The spectrograms were extracted up to a 200 Hz cutoff from the available 250 Hz, yielding 26 frequency bins, and then stacked, resulting in a dimension of time samples x features (26 x the number of channels).

[0124] In this study, a Mel generator model was provided based on the Squeezeformer, designed to convert stacked spectrograms of MEG signals into Mel spectrograms. The study utilized the Squeezeformer based on its outstanding performance in the field of ASR.Additionally, a bidirectional long short-term memory (Bi-LSTM) layer was employed to process information from both past and future contexts simultaneously. The proposed Mel generator consists of a linear layer, four sequential Squeezeformer blocks, one bidirectional long short-term memory (Bi-LSTM) layer, and one linear layer as illustrated in Figure 13. The dimensions listed under each layer represent the input dimension. The dimension of input data was time samples x features (26 frequency bins x # channels), fed into the first linear layer composed of 2048 hidden units, yielding the output with a dimension of time samples x 2048. It is noteworthy that the dimension of input data should be the same during the training process, but the length of all trials varied, resulting in different lengths of spectrograms and Mel spectrograms for each trial even for the same phrase. To address this issue, the study set the number of time samples to 60 (about 0.6 s data) for training the model because the minimum length of the spectrogram over all participants was 61. Additionally, the study randomly extracted data segments of 60 samples for each training epoch to ensure that the model can be trained using the entire data. All Squeezeformer blocks have an identical architecture with the encoder dimension of 1280, the number of attention heads of 20, and the convolution kernel size of 31. The other hyperparameters were set to the same as proposed by the authors. The Bi-LSTM layer has 2048 hidden units per direction, and the last linear layer comprises 100 hidden units, resulting in the output dimension of 60 x 100, which corresponds to the dimension of target Mel spectrograms.Attorney Docket No. 10046-617W018428 WAN

[0125] The proposed Mel generator was trained to minimize the mean squared error between the target Mel spectrogram and the predicted Mel spectrogram generated from the stacked spectrograms of MEG signals for 1000 epochs with a batch size of 4. AdamW optimizer was employed alongside a cosine annealing with warmup restarts scheduler. The initial and max learning rates were set to 0.000001 and 0.00001, respectively. The cycle and warm-up steps were set to 500 and 100, respectively, and the decreasing rate of max learning rate per cycle was 0.9. All hyperparameters were empirically determined using the data from Al, with different data splits for cross-validation. Ail implementation and training processes were performed using PyTorch on the M2 MAX chip of 32GB RAM MacBook Pro (Apple, USA).

[0126] A stratified 5-fold cross-validation strategy was employed to assess the performance of the proposed model. The dataset was divided into 5 folds, with each fold maintaining proportional data of each phrase. In each cross-validation, 4 folds were used to train the model, and the remaining fold was used as the validation set. This procedure was repeated 5 times with each fold serving as the validation set once. Although the model was trained with input data of fixed 60 time samples, it allows variable-length input sequences. Thus, the data with entire time samples were fed into the model during validation. The Pearson correlation coefficient and root mean squared error (RMSE) between the target and the generated Mel spectrograms of all validation sets were employed as the evaluation metrics.

[0127] The examples of the target and generated Mel spectrograms from participant A3 are illustrated in Figure 14. Here, “Target’ represents the Mel spectrograms derived from preprocessed audio signals, and “Generated” indicates the Mel spectrograms generated from MEG signals using the proposed method. The color bars denote the range of the logarithmic scale. The examples were generated using the validation set and the phrases ‘Good-Bye’. To assess the similarity of the target and generated Mel spectrograms, Pearson correlation coefficients and RMSEs were evaluated for all the trials of the validation set within the 5-fold cross-validation, shown in Figure 15, panels (a) and (b), respectively. The average coefficients were 0.934 + 0.039, 0.965 ± 0.021, 0.963 + 0.025, and 0.953 + 0.033, respectively, and the grand average across all participants was 0.954 ± 0.014 (mean ± standard deviation). The average MSEs were 1.353 ± 0.423, 1.067 ± 0.305, 1.061 ± 0.312, and 1.118 ± 0.324, respectively, and the grand average across all participants was 1.150 ± 0.138. The results imply that the proposed Mel generator can successfully convert MEGAttorney Docket No. 10046-617WO18428 WAN features into target Mel spectrograms of synchronously recorded audio signals, achieving high similarity.

[0128] Figure 16 depicts the examples of synthesized audio and their pitches from phrase 3: ‘How are you?’ of participant Al. “Target” (indicated by blue lines) refers to the audio signals synthesized using the target Mel spectrograms, while “Generated” (indicated by orange lines) denotes the audio signals synthesized using the Mel spectrograms generated from MEG features. Pitch curves were obtained using the ‘pitch’ function in Matlab. As illustrated in the figure, the synthesized audio signals exhibit characteristics similar to those of the reconstructed audio signals, implying their high fidelity.

[0129] This study demonstrated the feasibility of synthesizing audible and intelligible speech directly from non-invasive brain recordings, for the first time. The proposed Mel generator achieved remarkable performance in converting MEG features into Mel spectrograms of audio signals with a grand average correlation coefficient of 0.95, leading to the synthesis of intelligible audio signals. The half-length of the validation data resulted in the half-length of audio, indicating this model did not just classify the phrases, this study further demonstrated that MEG has great potential for non-invasive speech-BCI. While this study has the limitation of using a dataset comprised of only unique five phrases, it potentially served as an important step toward future advancements in non-invasive speech- BCI technologies. This approach can be extended to enable free communication, not limited to a set of words or phrases. Therefore, it is also envisioned to investigate open vocabulary speech synthesis and ultimately on imagined speech. Since the default setting of the clinical MEG machine used in this study captures audio with a low sampling rate, resulting in low- quality speech, additional studies may also include independently recorded speech at a high sampling rate to achieve more natural speech synthesis.

[0130] Example #4: Automatic detection of ALS from single-trial MEG signals during speech tasks

[0131] MEG (Neuromag TRIUX; MEGIN, LCC) was used to collect the neuromagnetic signals from the participants. This device has 306 SQUID sensors (204 gradiometers and 102 magnetometers). A magnetically shielded room (MSIR) housed the MEG machine to restrict external magnetic noise. A digital light processing projector was used to present the visual stimuli approximately 90 cm from the subjects on a back projection screen. The stimuli were generated by a computer running the STIM2 software (Compumedics, Ltd.). Two pairs of bipolar EEG electrodes were used to record the electrocardiogram (EKG) and theAttorney Docket No. 10046-617W018428 WAN electrooculogram (EOG) signals. A custom air-pressure transducer located outside the MSR and connected to the analog input of the MEG system was used to measure jaw displacement during the tasks. An air-bladder was fixed under the subjects’ chin and relayed jaw movement (via pressure on the bladder) to the transducer via tubing connected to the air-inlet on the sensor. Voice data was recorded using a standard built-in microphone connected to a transducer placed outside the MSR. Both voice and jaw movement analog signals were then digitized by feeding into the MEG ADC in real-time as separate channels. Five commonly used phrases were used as stimuli for the speech tasks: 1. Do you understand me; 2. That’s perfect; 3. How are you? Good-bye; 5. 1 need help. The task phrases came from phrase lists commonly used in alternative augmented communication (AAC) devices and were selected to be more familiar to the patients and easier to recite than novel speech (Benkelman et al., 1984; Dash et al., 2020).

[0132] The experiment was designed as a time-locked delayed overt reading task where each trial was time-locked to stimulus onset (display of phrases on the screen). The phrases were individually presented for 1 s in a pseudorandomized order followed by a 1 s fixation cross. The subjects were previously instructed to think of speaking the phrase without mouthing during the fixation and to overtly articulate the phrase at their normal speaking rate and loudness when the fixation disappeared. The subjects had 3 s to perform the articulation before the next stimulus trial. Each participant completed 100 trials per phrase. To overcome potential difficulties verifying the timing of imagined speech (Cooney et al., 2018), the study designed this protocol to collect both speech imagination and speech production consecutively, in the same trial and under time constraints.

[0133] The MEG data were recorded with 4 kHz sampling frequency with an online filter of 0.3-1,330 Hz. The data were low pass filtered to 250 Hz with a 4thorder Butterworth filter and resampled to 1 kHz. Power line noise (60 Hz) and harmonics were removed with a 2ndorder infinite impulse response (HR) notch filter. Only gradiometer sensors were used for analysis. From the 204 gradiometer sensors, it was observed that four sensors exhibited substantial channel noise during the data collection process from various participants.Additionally, in certain cases, one or two additional sensors displayed irregularities resembling artifacts. Consequently, a total of eight sensors were deemed unsuitable and excluded from the analysis. The discarded sensors were the same for both ALS and healthy data. Therefore, the analysis was conducted using data exclusively from 196 sensors.Attorney Docket No. 10046-617W018428 WAN Independent component analysis (ICA) was used to remove artifacts (cardiac activity, eye blinks, and saccades) from the data.

[0134] The continuous MEG signals were epoched into trials from -0.5 to +4.5 s centered at stimulus onset. Covert speech segment was parsed as the data from 1 s to 2 s and overt speech segment was parsed as the data from 2 s to 4.5 s of each trial. By visually inspecting the data, trials were discarded if they contained high-amplitude artifacts or if the participant did not comply with the paradigm timing (e.g., the participant spoke before being provided the cue to articulate). Jaw movement data during the covert speech segment was used to verify that the participants were not moving their articulators during the covert speech task. Jaw movement data were not used for analysis in this study. Following preprocessing, a single participant’s dataset contained only 63 valid trials for a particular phrase. Therefore, to ensure an impartial comparison, the study exclusively considered the initial 60 trials per phrase per participant. 'The preprocessing of the raw MEG data was conducted using FieldTrip (Oostenveld et al., 2011) in MATLAB 2021b.

[0135] Sensor correlation has been used to characterize neurological disorders (Schindler et al., 2007; Bob et al., 2010). Here, the study computed Pearson’s correlation between each pair of gradiometer sensor signals. Analyses were performed for both speech imagination and speech production for each stimulus (phrase) and participant separately. Correlation values were computed at the single-trial level and then averaged across all trials. For this analysis, the study used all the spectral information (0.3-250 Hz) in the signals. Statistical 2-sample t-tests were used to compare the ALS and healthy groups (N = 15: 3 participants x 5 phrases) based on number of sensors showing larger absolute correlation coefficients (r > 0.5) and correlation density (sum of all absolute correlation values over total number of sensor-pairs) for both imagined and overt speech separately.

[0136] Each neural oscillation is associated with a key functional role in the brain and could potentially carry a neural biomarker of a disorder. Beta-band power has traditionally been associated with motor function in the brain (Fisher et al., 2012; Khanna and Carmena, 2015). Thus, for the speech-motor task (overt speech) and the speech-motor imagination task (imagined speech), the study compared power in this band and other canonical bands between the two groups. The study computed the average power of the neuromagnetic signals for each frequency range of interest: delta (1-4 Hz), theta (4-8 Hz), alpha (8-16 Hz), beta (16-30 Hz), gamma (30-59 Hz), and high gamma (61—119 Hz). The study then averaged the band powers (this was completed separately for each band) across both trials and participants. 'TheAttorney Docket No. 10046-617W018428 WAN pairwise Euclidian distances between healthy and ALS band powers were calculated across all sensors for each phrase and after averaging across the 5 phrases. 1-way analysis of variance (ANOVA) and post-hoc Tukey test was conducted with the six bands as independent groups and 5 phrases as different samples for both imagination and articulation.

[0137] Functional connectivity is defined as the statistical dependence among measured neural signals which explains the temporal coincidence of spatially distant neurophysiological events (Friston, 1994). Functional connectivity analysis has become the conventional choice for a better understanding of the in vivo pathology of ALS. In this study, amplitude envelope correlation (AEC) (O’Neill et al., 2015) was used to measure the functional connectivity for each frequency band. For single-trial functional connectivity analysis, the study used a 4thorder Butterworth bandpass filter to first bandpass the gradiometer signals from all 196 sensors at each frequency range of interest, obtained the amplitude envelopes using Hilbert transform, and then computed the pairwise linear correlation of the amplitude envelopes across all sensors for each frequency range of interest separately. Connectivity was defined as the averaged pair-wise correlation across trials. For the individual subject analysis, first, the study temporally concatenated all bandpass-filtered single trials, extracted the envelope, and then computed the correlations. The study performed the AEC -based functional connectivity analysis for each phrase separately during both imagined and overt speech. A 2-sample one-sided t-test was conducted between healthy and ALS samples (3 subjects x 5 phrases — for each group) of functional connectivity density (sum of AEC values over total number of sensor pairs) to check for the hypothesis of whether patients with ALS show greater beta band connectivity than healthy controls.

[0138] The study used power in the six canonical frequency bands of the MEG signals as features to train a linear discriminant analysis (LDA) algorithm and classified ALS and healthy data during both speech imagination and overt speech. The study trained the model separately for each frequency range of interest and separately using a wide frequency range (0.3-250 Hz) which contained spectral information from all the neural oscillations. The choice of the LDA model was inspired by previous experiments on speech decoding for ALS where the LDA model performed equivalently to both support vector machines and multilayer perceptron classifiers (Dash et al., 2020) at classifying 5 phrases.The fitcdiscr function in the Statistical and Machine Learning Toolbox of MATLAB was used for classification. The lower sample size than the feature dimension motivated for a linear type of discriminant. The linear coefficient threshold (‘Delta’) and the amount ofAttorney Docket No. 10046-617W018428 WAN regularization (‘Gamma’) of the model were tuned as the hyperparameters of the model, computed based on the Bayesian optimization search using a 10-fold cross-validation on the training data. All other parameters were set to the default values of the toolbox. The study used a leave-one-pair-out cross-validation strategy where the study trained the model with all trials from 2. healthy and 2 ALS participants and tested using the remaining data from 1 healthy and 1 ALS participant, irrespective of the phrase. This was repeated until each healthy- ALS pair was tested. This led to a training data size of 1,200 trials (4 participants (2 healthy +2 ALS) x 5 phrases x 60 trials) and a test data size of 600 trials (2 participants (1 healthy +1 ALS) x 5 phrases x 60 trials) for each fold. In this manner, the trained decoder was tested with completely unseen new participant data.

[0139] Figure 17 shows the comparative histogram distribution of sensor-level signal correlations for ALS and healthy controls for each phrase (top for imagined speech and bottom for overt speech). A significantly larger number of sensors showed greater correlations for ALS compared to healthy controls across all phrases during both overt (one-sided, 2-sample t-test: t = 3.76, df = 28, p < 0.001) and imagined speech (one-sided, 2- sample t-test: t = 6.01, df = 28, p < 0.001). This is also evident by the higher variance in the distribution of correlations for ALS compared to healthy controls for both imagined and overt speech across all phrases. In other words, the majority of the correlations were near-mean (i.e., zero correlation) for the controls compared to ALS. For imagined speech, 95% (Bayesian analysis based on Monte Carlo simulations) of the correlation values were in a range of -0.5 to 0.5 for healthy participants. The range was between -0.8 to 0.8 for ALS participants. For overt speech, the correlation range for healthy controls was approximately within the range - 0.8 to 0.8, which was greater for the ALS ranging from -1 to 1 A heatmap plot of the correlation distribution for each subject is shown in Figure 18 (Top for imagined speech and middle for overt speech), which depicts stronger correlations across the whole brain for participants with ALS compared to the healthy controls, especially for the first participant with ALS (Al) who also had the lowest speech intelligibility and speaking rate scores. To interpret these correlation heatmaps, correlation density was calculated as the sum of all absolute correlation values over total number of sensor-pairs and shown for each participant in Figure 18 — Bottom panel. Mean correlation density was higher for participants with ALS (Overt: 0.428; Imagination: 0.326) compared to healthy subjects (Overt: 0.292; Imagination: 0.192) averaged across trials, phrases, and participants as well as statistically across all phrases and participants (one-sided, 2-sample t-tests:Attorney Docket No. 10046-617W018428 WAN overt: t= 6.13, df= 23, p< 0.001; imagined: t= 6.68, df = 28, p <0.001). As expected, a stronger correlation for overt speech was observed compared to speech imagination, irrespective of healthy or ALS data.

[0140] Figure 19 shows the mean band-power distances between healthy and ALS for each during overt speech task (articulation), where panels (A) to (F) are for delta, theta, alpha, beta, gamma, and high gamma bands, respectively. For belter visualization, the distances are shown as heatmaps where the color range from blue to red indicates minimum to maximum range of the normalized distance values. Each cell in the heatmap represents a pairwise distance between the band powers of a healthy sensor (y-axis) and an ALS sensor (x-axis) across all the phrases. The distances were significantly greater for the beta band powers than the other canonical bands for both imagination (1-wayANOVA: F = 208.55, p < 0.001; post-hoc Tukey tests: beta vs. rest: p < 0.0019) and articulation (1-way ANOVA: F = 206.59, p < 0.001; post-hoc Tukey tests: beta vs. rest: p < 0.0021, see Figure 19 — Panel G and H). Also, a larger number of pairwise (ALS — healthy) dissimilarities were observed in the beta band. These oscillatory patterns were similar for each phrase and across all phrases for individual and group subject analysis irrespective of speech task, i.e., imagination or production. A couple of sensors showed the highest distance (solid red lines in the heatmap) which could be because those sensors were noisy.

[0141] Figure 20 shows the AEC-based beta-band functional connectivity for both groups during the production of the phrase ‘Do you understand me?’ in the form of heatmaps; showing the correlation range of -1 to 1 (from blue [minimum] to red [maximum]). Greater beta band connectivity was significant for ALS patients compared to healthy subjects (one sided, 2-sample t-test: p < 0.05). This is notably apparent in the first patient (Al), who had more severe bulbar impairment than the other two (A2 and A3). Interestingly, similar patterns of increased connectivity were also prominent during speech imagination. A more diverse connectivity pattern among the 3 patients with ALS compared to the healthy participants can be observed by visualizing the connectivity strengths.

[0142] Figure 21 shows the median single-trial classification accuracy for the healthy versus ALS group for both speech imagination and overt speech tasks. The best performance (median accuracy -98%) was obtained using beta bands and was similar for both speech tasks. The performance using each individual frequency range of interest (excluding delta) was significantly higher than chance level (50%) and was also higher when compared toAttorney Docket No. 10046-617WO18428 WAN performance using all frequency information (ail: 0.3-250 Hz). The performance accuracy was lowest for the first ALS patient (Al) (mean across folds = 65% for overt speech; 83% for imagined speech), likely because this participant’s speech symptoms were severe compared to the other two participants with ALS. Although median performance was highest for beta band, statistically, 1-way ANOVA based comparison did not show a significant difference between the performances of different bands (F = 1.33, p = 0.05), possibly due to the low sample size. For the case of imagined speech, performances obtained with theta and gamma band were comparable to beta band performance.

[0143] The evidence of greater inter-sensor correlation for ALS compared to healthy participants is a clear distinguishable marker between the two groups. This has been previously observed with M / EEG resting state (Proudfoot et al., 2019) and motor imagery studies (Yang et al., 2018). This difference in sensor correlations was apparent across all phrases and participants which further illustrates that this feature is independent of stimuli and an across-subject observation. A stronger correlation during the overt speech task compared to the imagined speech task indicated greater cortical activity for producing overt speech compared to speech imagination, which was true for both healthy and AL S groups and was expected. The signal artifacts introduced by movement during the production of speech gestures could have also contributed to the higher sensor correlation during overt speech production (Dash et al., 2018). Participant (Al) with the most severe symptoms showed the strongest correlation (Figure 18) suggesting that the proposed approach may be useful as a marker of disease progression. Further, sensor correlation differences could also arise from the differences in the head positions inside the scanner. Mapping the sensor data into source space and performing the correlations across parcels / voxels would be a better way to remove these confounds.

[0144] Beta band has been traditionally associated with motor function, and the observed differences in the beta band power during overt speech are consistent with the hypotheses that ALS is associated with cortical hyperexcitability, possibly due to the loss of inhibitory interneuron (Proudfoot et al., 2017). The results show the importance of beta band for identifying ALS during a bulbar motor task (speech). The prominent beta band differences during the speech imagination task (which does not involve motor execution) suggest that the beta band during speech tasks that involve motor planning could be a potential neural biomarker of ALS. Importantly, the pattern of band-power differences was similar for both imagination and overt speech in the beta band, possibly indicating a functional similarityAttorney Docket No. 10046-617WO18428 WAN between the two speech tasks. Clear differences in band power were also observed in the theta and the gamma band; however, they were less prominent compared to the beta band differences.

[0145] This finding of significantly greater beta band connectivity in the ALS group compared to healthy controls follows literature since beta band functional connectivity changes have been previously shown (Verstraete et al., 2010; Agosta et al., 2013; Proudfoot et al., 2017) both during resting state as well as for spinal motor tasks. An increase in beta band functional connectivity has been hypothesized as the result from loss of intracortical inhibitory influence supported in vivo by neurophysiology findings of accentuated cortical beta-desynchronization during movement preparation and diminished post-movement beta-rebound (Proudfoot et al., 2017). This inhibitory influence may lead to compensatory mechanisms in early-stage ALS resulting in higher functional connectivity. Behaviorally, it may be explained as a compensatory mechanism for speech tasks in early-stage ALS due to the recruitment of larger neural networks, supported by tongue kinematic studies (Green et al., 2013; Kuruvilla-Dugdale and Mefferd, 2017; Teplansky et al., 2019). Additionally, disruption of efficient motor control networks in ALS may lead to higher cognitive control demands and attention increases, both of which are known to modulate beta-band oscillation power and connectivity (Cheyne and Ferrari, 2013; Riddle et al., 2021). This study provides the first evidence of increased beta band connectivity during a speech-motor task.Interestingly, increased connectivity was also prominent during speech imagination indicating higher beta band connectivity during speech planning, thereby strengthening the role of beta band as a neural biomarker for ALS. Similar to the observations with band-power differences, beta-band functional connectivity was greatest for the most severe patient (participant Al), providing additional confidence in the specificity of this marker.Accumulating the connectivity strength (i.e., correlation values) of all three subjects and for all 5 phrases showed greater connectivity strength for the ALS group than the control group indicating beta-band connectivity to be an across-subject marker.

[0146] Evidence of a high single-trial ALS detection accuracy with beta-band suggests that the neural mechanisms for ALS could be specific to spectral content, particularly to the beta band during speech tasks. From the previous qualitative analyses (sensor correlation, band power difference, and functional connectivity) beta-band was expected to perform the best for single-trial classification. The median accuracy with beta band was superior when both overt and imagined phrases were considered, although theta and gamma band alsoAttorney Docket No. 10046-617W018428 WAN showed comparable performance with beta band for the case of imagined phrases. A recent study suggested covert speech emphasizes both beta and gamma band (Moon et al., 2022), which may explain why gamma band also obtained high accuracies. In short, greater performance accuracies during the speech imagination task suggest that neural signals derived while imagining speech may be optimal for diagnosing early-onset ALS, whereas overt speech may be more appropriate for evaluating the rate of disease progression.Crucially, this is the first demonstration of ALS detection from single-trial neural signals.

[0147] In terms of behavior, individuals with ALS exhibited larger onset latency and duration in overt speech tasks when compared to their healthy counterparts (2-sample t-tests, t = 3.09, p = 0.002, N = 900 [3 participants x 5 phrases x 60 trials], as one would expect the patients to take longer time to complete the task. It is plausible that these behavioral effects manifest in elevated sensor correlation and functional connectivity strength for ALS patients as opposed to healthy controls. However, there is notable convergence in these behaviors at the single-trial level between the two population groups, with more than 44% overlap in onset time and over 2.2% overlap in duration. The behavioral difference was mostly driven by the first ALS participant (Al) with the lowest speaking rate and speech intelligibility.Consequently, relying solely on behavioral indicators for single-trial detection proves to be inefficient. In addition, similar cortical differences were also observed during the covert speech task, a scenario where these behavioral markers are absent. Further, covert speech segments are immune to movement artifacts that can be present during overt speech and bias the results. Hence, the optimal approach for single-trial ALS detection involves analyzing neural activity during covert speech tasks.

[0148] Other Machine Learning. In addition to the machine learning features described above, the various analysis system can be implemented using one or more artificial intelligence and machine learning operations. The term “artificial intelligence” can include any technique that enables one or more computing devices or computing systems (i.e., a machine) to mimic human intelligence. Artificial intelligence (Al) includes but is not limited to knowledge bases, machine learning, representation learning, and deep learning. The term “machine learning” is defined herein to be a subset of Al that enables a machine to acquire knowledge by extracting patterns from raw data. Machine learning techniques include, but are not limited to, logistic regression, support vector machines (SVMs), decision trees, Naive Bayes classifiers, and artificial neural networks. The term “representation learning” is defined herein to be a subset of machine learning that enables a machine to automatically discoverAttorney Docket No. 10046-617W018428 WAN representations needed for feature detection, prediction, or classification from raw data. Representation learning techniques include, but are not limited to, autoencoders and embeddings. The term “deep learning” is defined herein to be a subset of machine learning that enables a machine to automatically discover representations needed for feature detection, prediction, classification, etc., using layers of processing. Deep learning techniques include but are not limited to artificial neural networks or multilayer perceptron (MLP).

[0149] Machine learning models include supervised, semi-supervised, and unsupervised learning models. In a supervised learning model, the model learns a function that maps an input (also known as feature or features) to an output (also known as target) during training with a labeled data set (or dataset). In an unsupervised learning model, the algorithm discovers patterns among data. In a semi-supervised model, the model learns a function that maps an input (also known as a feature or features) to an output (also known as a target) during training with both labeled and unlabeled data.

[0150] Neural Networks. An artificial neural network (ANN) is a computing system including a plurality of interconnected neurons (e.g., also referred to as “nodes”). This disclosure contemplates that the nodes can be implemented using a computing device (e.g., a processing unit and memory as described herein). The nodes can be arranged in a plurality of layers such as an input layer, an output layer, and optionally one or more hidden layers with different activation functions. An ANN having hidden layers can be referred to as a deep neural network or multilayer perceptron (MLP). Each node is connected to one or more other nodes in the ANN. For example, each layer is made of a plurality of nodes, where each node is connected to all nodes in the previous layer. The nodes in a given layer are not interconnected with one another, i.e., the nodes in a given layer function independently of one another. As used herein, nodes in the input layer receive data from outside of the ANN, nodes in the hidden layer(s) modify the data between the input and output layers, and nodes in the output layer provide the results. Each node is configured to receive an input, implement an activation function (e.g., binary step, linear, sigmoid, tanh, or rectified linear unit (ReLU)), and provide an output in accordance with the activation function. Additionally, each node is associated with a respective weight. ANNs are trained with a dataset to maximize or minimize an objective function. In some implementations, the objective function is a cost function, which is a measure of the ANN’S performance (e.g., error such as LI or L2 loss) during training, and the training algorithm tunes the node weights and / or bias to minimize the cost function. This disclosure contemplates that any algorithm that finds the maximum orAttorney Docket No. 10046-617W018428 WAN minimum of the objective function can be used for training the ANN. Training algorithms for ANNs include but are not limited to backpropagation. It should be understood that an ANN is provided only as an example machine learning model. This disclosure contemplates that the machine learning model can be any supervised learning model, semi-supervised learning model, or unsupervised learning model. Optionally, the machine learning model is a deep learning model. Machine learning models are known in the art and are therefore not described in further detail herein.

[0151] A convolutional neural network (CNN) is a type of deep neural network that has been applied, for example, to image analysis applications. Unlike traditional neural networks, each layer in a CNN has a plurality of nodes arranged in three dimensions (width, height, depth). CNNs can include different types of layers, e.g., convolutional, pooling, and fully- connected (also referred to herein as “dense”) layers. A convolutional layer includes a set of filters and performs the bulk of the computations. A pooling layer is optionally inserted between convolutional layers to reduce the computational power and / or control overfitting (e.g., by downsampling). A fully-connected layer includes neurons, where each neuron is connected to all of the neurons in the previous layer. The layers are stacked similarly to traditional neural networks. GCNNs are CNNs that have been adapted to work on structured datasets such as graphs.

[0152] As used herein, the term “encoder-decoder Al model” describes an architecture of a machine learning system in which input data are encoded or coded in order then to be decoded again immediately afterward. In the middlebetween the encoder and the decoder the necessary data are present as a type of feature vector. During decoding, depending on the training of the machine learning model, specific features in the input data can then be specially highlighted.

[0153] As used herein, the term “U-net” describes an architecture of a machine learning system which is based on a convolutional network architecture. This architecture is particularly well suited to a fast and accurate segmentation of digital imagesin the biological / medical field. A further advantage of such a machine learning system is that it manages with fewer training data and allows a comparatively accurate segmentation.

[0154] Other Supervised Learning Models. A logistic regression (LR) classifier is a supervised classification model that uses the logistic function to predict the probability of a target, which can be used for classification. LR classifiers are trained with a data set (also referred to herein as a “dataset”) to maximize or minimize an objective function, for example,Attorney Docket No. 10046-617W018428 WAN a measure of the LR classifier’s performance (e.g., error such as LI or L2 loss), during training. This disclosure contemplates that any algorithm that finds the minimum of the cost function can be used. LR classifiers are known in the art and are therefore not described in further detail herein.

[0155] A Naive Bayes’ (NB) classifier is a supervised classification model that is based on Bayes’ Theorem, which assumes independence among features (i.e., the presence of one feature in a class is unrelated to the presence of any other features). NB classifiers are trained with a data set by computing the conditional probability distribution of each feature given a label and applying Bayes’ Theorem to compute the conditional probability distribution of a label given an observation. NB classifiers are known in the art and are therefore not described in further detail herein.

[0156] A k-NN classifier is an unsupervised classification model that classifies new data points based on similarity measures (e.g., distance functions). The k-NN classifiers are trained with a data set (also referred to herein as a “dataset”) to maximize or minimize a measure of the k-NN classifier’s performance during training. This disclosure contemplates any algorithm that finds the maximum or minimum. The k-NN classifiers are known in the art and are therefore not described in further detail herein.

[0157] Conclusion

[0158] The methods, systems, and devices discussed above are examples. Various embodiments may omit, substitute, or add various procedures or components as appropriate. For instance, in alternative configurations, the methods described may be performed in an order different from that described, and / or various stages may be added, omitted, and / or combined. Also, features described with respect to certain embodiments may be combined in various other embodiments. Different aspects and elements of the embodiments may be combined in a similar manner. Also, technology evolves and, thus, many of the elements are examples that do not limit the scope of the disclosure to those specific examples.

[0159] The foregoing method descriptions and the process flow diagrams are provided merely as illustrative examples and are not intended to require or imply that the operations of various embodiments must be performed in the order presented. As will be appreciated by one of skill in the art the order of operations in the foregoing embodiments may be performed in any order. Words such as “thereafter,” “then,” “next,” etc. are not intended to limit the order of the operations; these words are simply used to guide the reader through the description of the methods. Further, any reference to claim elements in the singular, forAttorney Docket No. 10046-617W018428 WAN example, using the articles “a,” “an” or “the” is not to be construed as limiting the element to the singular.

[0160] While the terms “first” and “second” are used herein to describe data transmission associated with a subscription and data receiving associated with a different subscription, such identifiers are merely for convenience and are not meant to limit various embodiments to a particular order, sequence, type of network or carrier.

[0161] Various illustrative logical blocks, modules, circuits, and algorithm operations described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and operations have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such embodiment decisions should not be interpreted as causing a departure from the scope of the claims.

[0162] The hardware used to implement various illustrative logics, logical blocks, modules, and circuits described in connection with the embodiments disclosed herein may be implemented or performed with a general -purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing systems (e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Alternatively, some operations or methods may be performed by circuitry that is specific to a given function.

[0163] In one or more example embodiments, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or codes on a non-transitory computer-readable medium or non-transitory processor-readable medium. The operations ofAttorney Docket No. 10046-617W018428 WAN a method or algorithm disclosed herein may be embodied in a processor-executable software module, which may reside on a non -transitory computer-readable or processor-readable storage medium. Non-transitory computer-readable or processor-readable storage media may be any storage media that may be accessed by a computer or a processor. By way of example but not limitation, such non-transitory computer-readable or processor-readable media may include RAM, ROM, EEPROM, FLASH memory, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that may be used to store desired program code in the form of instructions or data structures and that may be accessed by a computer. Disk and disc, as used herein, includes compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks usually reproduce data magnetically, while discs reproduce data optically with lasers.Combinations of the above are also included within the scope of non-transitory computer- readable and processor-readable media. Additionally, the operations of a method or algorithm may reside as one or any combination or set of codes and / or instructions on a non- transitory processor-readable medium and / or computer-readable medium, which may be incorporated into a computer program product.

[0164] Those of skill in the art will appreciate that information and signals used to communicate the messages described herein may be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.

[0165] Terms “and” and “or” as used herein may include a variety of meanings that also are expected to depend at least in part upon the context in which such terms are used.Typically, “or” if used to associate a list, such as A, B, or C, is intended to mean A, B, and C, here used in the inclusive sense, as well as A, B, or C, here used in the exclusive sense. In addition, the term “one or more,” as used herein, may be used to describe any feature, structure, or characteristic in the singular or may be used to describe some combination of features, structures, or characteristics. However, it should be noted that this is merely an illustrative example, and claimed subject matter is not limited to this example. Furthermore, the term “at least one of” if used to associate a list, such as A, B, or C, can be interpreted to mean any combination of A, B, and / or C, such as A, AB, AC, BC, AA, ABC, AAB, AABBCCC, and the like.Attorney Docket No. 10046-617W018428 WAN

[0166] As used herein, the term “is associated with”, as in A is associated with B, means that A refers to B, is B, identifies a feature of B, or indicates that B exists. For example, an image that is associated with a biological state can, by virtue of the data it contains, refer to that biological state, indicate the presence of that biological state, identify a feature of that biological state, or simply indicate that that biological state exists.

[0167] Further, while certain embodiments have been described using a particular combination of hardware and software, it should be recognized that other combinations of hardware and software are also possible. Certain embodiments may be implemented only in hardware, or only in software, or using combinations thereof. In one example, the software may be implemented with a computer program product containing computer program code or instructions executable by one or more processors for performing any or all of the steps, operations, or processes described in this disclosure, where the computer program may be stored on a non -transitory computer-readable medium. The various processes described herein can be implemented on the same processor or different processors in any combination.

[0168] Where devices, systems, components or modules are described as being configured to perform certain operations or functions, such configuration can be accomplished, for example, by designing electronic circuits to perform the operation, by¬ programming programmable electronic circuits (such as microprocessors) to perform the operation such as by executing computer instructions or code, or processors or cores programmed to execute code or instructions stored on a non -transitory memory medium, or any combination thereof. Processes can communicate using a variety of techniques, including, but not limited to, conventional techniques for inter-process communications, and different pairs of processes may use different techniques, or the same pair of processes may use different techniques at different times.

[0169] While the disclosure has been described in detail with reference to exemplary-embodiments, those skilled in the art will appreciate that various modifications and substitutions may be made thereto without departing from the spirit and scope of the disclosure as set forth in the appended claims. For example, elements and / or features of different exemplary embodiments may be combined with each other and / or substituted for each other within the scope of this disclosure and appended claims.

[0170] Example Embodiments

[0171] Embodiment 1: A method for decoding neural speech, the method comprising:receiving a plurality of neural signals corresponding to spontaneous neural speechAttorney Docket No. 10046-617W018428 WAN of a subject obtained from an array of neuromagnetic sensors; generating a spatiotemporal map from the plurality of neural signals; and determining a predictive speech classification from the spatiotemporal map using a trained artificial intelligence (Al) model.

[0172] Embodiment 2: The method of embodiment 1, wherein the neuromagnetic sensors comprise magnetoencephalography (MEG) sensors.

[0173] Embodiment 3: The method of embodiment 2, wherein the MEG sensors comprise optically-pump magnetometers (OPM).

[0174] Embodiment 4: The method of any one of embodiments 1-3, wherein the spatiotemporal map comprises a plurality of Mel-spectrograms.

[0175] Embodiment 5: The method of any one of embodiments 1-4, wherein the spontaneous neural speech comprises uncued speech.

[0176] Embodiment 6: The method of any one of embodiments 1-5, wherein a temporal window of the plurality of neural signals is 10 seconds or less (e.g., 7 seconds or less, 5 seconds or less, or 3 seconds or less).

[0177] Embodiment 7: The method of any one of embodiments 1-6, wherein the trained Al model is configured to determine the predictive speech classification from a closed-set of speech units.

[0178] Embodiment 8: The method of any one of embodiments 1-6, wherein the trained Al model is configured to determine an open-set of speech units from the spatiotemporal map.

[0179] Embodiment 9: The method of any one of embodiments 1-8, wherein the trained Al model comprises a vocoder configured to generate an audio signal from the spatiotemporal map.

[0180] Embodiment 10: A method for biometric authentication comprising: recording a neural signal of a subject from a neuroimaging sensor, wherein the neural signal includes features corresponding to neural activity of the subject performing a cognitive task; and determining an authentication classification using a trained artificial intelligence (Al) model, wherein the Al model is configured to receive the neural signal and predict an identity of the subject.

[0181] Embodiment 11: The method of embodiment 10, wherein the neuroimaging sensor comprises a non-invasive sensor

[0182] Embodiment 12: The method of any one of embodiments 10-11, wherein the neuroimaging sensor comprises a neuromagnetic sensor.Attorney Docket No. 10046-617W018428 WAN

[0183] Embodiment 13: The method of any one of embodiments 10-12, wherein the neuroimaging sensor comprises a magnetoencephalography (MEG) sensor.

[0184] Embodiment 14: The method of embodiment 13, wherein the MEG sensor comprises an optically-pump magnetometer (OPM).

[0185] Embodiment 15: The method of any one of embodiments 10-14, wherein the neuroimaging sensor comprises a plurality of sensors.

[0186] Embodiment 16: The method of any one of embodiments 10-15, wherein the cognitive task comprises overt speech.

[0187] Embodiment 17: The method of any one of embodiments 10-16, wherein the cognitive task comprises imagined speech.

[0188] Embodiment 18: The method of any one of embodiments 10-17, further comprising prompting the subject to perform the cognitive task prior to recording the neural signal.

[0189] Embodiment 19: The method of any one of embodiments 10-18, further comprising unlocking a device based on the authentication classification.

[0190] Embodiment 20: A method for detecting a neurological disorder in a subject, the method comprising: recording a plurality of neural signals of a subject from an array of neuromagnetic sensors, wherein the neural signals includes features corresponding to neural acti vity of the subject performing a cognitive task; and determining an disease state of the subject using a trained artificial intelligence (Al) model, wherein the Al model is configured to receive the plurality of neural signals and predict an disease state of the subject based a cortical feature for the neurological disorder.

[0191] Embodiment 21: The method of embodiment 20, where the neurological disorder is amyotrophic lateral sclerosis (ALS).

[0192] Embodiment 22: The method of any one of embodiments 20-21, wherein the cortical feature comprises beta band oscillations.

[0193] Embodiment 23: The method of any one of embodiments 20-22, wherein the cognitive task comprises overt speech.

[0194] Embodiment 24: The method of any one of embodiments 20-23, wherein the cognitive task comprises imagined speech.

[0195] Embodiment 25: The method of any one of embodiments 20-24, wherein the neuromagnetic sensors comprise magnetoencephalography (MEG) sensors.Attorney Docket No. 10046-617W018428 WAN

[0196] Embodiment 26: The method of embodiment 25, wherein the MEG sensors comprise optical! y-pump magnetometers (OPM).

[0197] Embodiment 27: The method of any one of embodiments 20-26, wherein the detection of the neurological disorder is a single-trial classification.

[0198] Embodiment 28: A system comprising:one or more processors; anda memory having instructions stored thereon, wherein the instructions, when executed by the one or more processors, cause the one or more processors to perform any one of the methods of embodiments 1-27.

[0199] Embodiment 29: A non-transitory computer readable medium having instructions stored thereon, wherein the instuctions, when executed by one or more processors, cause the one or more processors to perform and one of the methods of embodiments 1-27.

[0200] The disclosures of each and every publication cited herein are hereby incorporated by reference in their entirety.[1] U. S Pat. No. 10,795,440 Bl[2] U. S Pat. No. 7,546,158 B2[3] U. S App. Pub. No. 2013 / 0130799 A l[4] U. S App. Pub. No. 2023 / 0210427 Al[5] U. S App. Pub. No. 2024 / 0194179 Al[6] Dash, Debadatta, Paul Ferrari, and Jun Wang. " Neural decoding of spontaneous overt and intended speech." Journal of Speech, Language, and Hearing Research 67.11 (2024): 4216-4225.[7] Kwon, Jinuk, et al. " Direct Speech Synthesis from Non-Invasive, Neuromagnetic Signals." Proc. Interspeech 2024. 2024.[8] Dash, Debadatta, et al. " Automatic detection of AES from single-trial MEG signals during speech tasks: a pilot study." Frontiers in Psychology 15 (2024): 1114811.[9] DA Moses, MK Leonard, JG Makin, EF Chang, “Real-time decoding of question-and-answer speech dialogue using human cortical activity”, Nature communications, 2019.

[0010] GK Anumanchipalli, J Chartier, EF Chang, “Speech synthesis from neural decoding of spoken sentences”, Nature, 2019.Attorney Docket No. 10046-617W018428 WAN

[0011] Tang, Jerry and LeBel, Amanda and Jain, Shailee and Huth, Alexander G (2023). “Semantic reconstruction of continuous language from non-invasive brain recordings.” Nature Neuroscience. doi:10.1038 / s41593-023-01304-9.

Claims

Attorney Docket No. 10046-617W018428 WAN CLAIMSWhat is claimed is:

1. A method for decoding neural speech, the method comprising:receiving a plurality of neural signals corresponding to spontaneous neural speech of a subject obtained from an array of neuromagnetic sensors;generating a spatiotemporal map from the plurality of neural signals; and determining a predictive speech classification from the spatiotemporal map using a trained artificial intelligence (Al) model.

2. The method of claim 1, wherein the neuromagnetic sensors comprise magnetoencephalography (MEG) sensors.

3. The method of claim 2, wherein the MEG sensors comprise optically-pump magnetometers (OPM).

4. The method of claim 1, wherein the spatiotemporal map comprises a plurality of Mel-spectrograms.

5. The method of claim 1, wherein the spontaneous neural speech comprises uncued speech.

6. The method of claim 1, wherein a temporal window of the plurality of neural signals is 10 seconds or less (e.g., 7 seconds or less, 5 seconds or less, or 3 seconds or less).

7. The method of claim 1, wherein the trained Al model is configured to determine the predictive speech classification from a closed-set of speech units.

8. The method of claim 1, wherein the trained Al model is configured to determine an open-set of speech units from the spatiotemporal map.Attorney Docket No. 10046-617W018428 WAN 9. The method of claim 1, wherein the trained Al model comprises a vocoder configured to generate an audio signal from the spatiotemporal map.

10. A method for biometric authentication comprising:recording a neural signal of a subject from a neuroimaging sensor, wherein the neural signal includes features corresponding to neural activity of the subject performing a cognitive task; anddetermining an authentication classification using a trained artificial intelligence (Al) model, wherein the Al model is configured to receive the neural signal and predict an identity of the subject.

11. The method of claim 10, wherein the neuroimaging sensor comprises a non-invasive sensor.

12. The method of claim 10, wherein the neuroiniaging sensor comprises a neuromagnetic sensor.

13. The method of claim 10, wherein the neuroimaging sensor comprises a magnetoencephalography (MEG) sensor.

14. The method of claim 13, wherein the MEG sensor comprises an optically -pump magnetometer (0PM).

15. The method of claim 10, wherein the neuroimaging sensor comprises a plurality of sensors.

16. The method of claim 10, wherein the cognitive task comprises overt speech.

17. The method of claim 10, wherein the cognitive task comprises imagined speech.

18. The method of claim 10, further comprising prompting the subject to perform the cognitive task prior to recording the neural signal.Attorney Docket No. 10046-617W018428 WAN 19. The method of claim 10, further comprising unlocking a device based on the authentication classification.

20. A method for detecting a neurological disorder in a subject, the method comprising:recording a plurality of neural signals of a subject from an array of neuromagnetic sensors, wherein the neural signals includes features corresponding to neural activity of the subject performing a cognitive task; anddetermining an disease state of the subject using attained artificial intelligence (Al) model, wherein the Al model is configured to receive the plurality of neural signals and predict an disease state of the subject based a cortical feature for the neurological disorder.