System for planning and / or providing neuromodulation and / or neurostimulation and method thereof
The system addresses the challenges of decoding brain signals by using self-supervised learning to transform raw brain signals into robust latent features, improving the accuracy and generalizability of neuromodulation and neurostimulation.
Patent Information
- Application Number
- PCT/EP2024/087977
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-20
- Publication Date
- 2025-06-26
AI Technical Summary
Current neuromodulation and neurostimulation systems face challenges in effectively decoding movement intentions from brain signals, particularly due to distribution shifts in neural dynamics across time, tasks, subjects, and therapeutic modalities, and the difficulty in annotating brain activity.
A system that utilizes self-supervised learning (SSL) to extract meaningful information from unlabeled cortical brain data, transforming raw brain signals into latent space features that are robust and transferable across different contexts, thereby improving the accuracy and generalizability of neuromodulation and neurostimulation.
The system achieves improved performance in decoding motor intentions and other behavioral states with reduced reliance on labeled data, enhanced robustness against temporal shifts in neural dynamics, and increased efficiency in calibrating decoders, allowing for more effective and targeted neuromodulation and neurostimulation.
Smart Images

Figure EP2024087977_26062025_PF_FP_ABST
Abstract
Description
[0001] System for planning and / or providing neuromodulation and / or neurostimulation and method thereof
[0002] The present invention belongs to the technical field of neuromodulation and / or neurostimulation, e.g. for restoring and / or controlling a motor function in a subject in need thereof.
[0003] In particular, said subject can be a human patient affected by spinal cord injury (SCI) and / or other neurological disorders such disorders such as stroke, Parkinson’s disease, and / or other neurodegenerative disorders.
[0004] In particular, the present invention provides an improved system for planning and / or providing a neuromodulation and / or neurostimulation.
[0005] The invention further provides a method for planning neuromodulation and / or neurostimulation through said system.
[0006] Still further, the invention provides a computer-readable non-volatile storage medium comprising computer-readable instructions that, when executed by a processing means, causes said processing means to implement said method.
[0007] Neuromodulation and / or neurostimulation can be provided for restoring and / or controlling motor functions of a patient, e.g. movements of the upper and / or lower limbs.
[0008] Here, it is important that delivered stimulation is well-synchronized with a movement intention of the patient, thereby providing a feeling that is close to natural as much as possible.
[0009] The use of brain-computer interfaces (BCI) is known to help for this purpose.
[0010] In particular, BCI rely on extraction and interpretation of brain activity to control a wide variety of effectors.
[0011] However, despite the growing amount of collected neurophysiological data collected globally, classical machine learning methods, relying on expert supervision, fall short of delivering robust and transferable models for clinically relevant applications. Decoding movement intentions involves extracting useful information from raw sequences of neural activity and linking it to specific behaviors, thus establishing a conversion route from thoughts to actions.
[0012] Decoders are typically trained using supervised machine learning algorithms.
[0013] However, numerous challenges arise in this process, notably characterized by distribution shifts in neural dynamics across times, tasks, subjects, and therapeutic modalities.
[0014] The fundamental problem lies in the difficulty of annotating brain activity compared to other domains of artificial intelligence.
[0015] Typically, data labeling is performed through calibration sessions involving repetition of numerous instructed intentions within a controlled clinical environment.
[0016] Apart from being expensive to obtain, labels are often uncertain and compressed descriptions of neural activity.
[0017] Consequently, the trained decoders rely on task-specific characteristics, such that even minor deviations from the task conditions may significantly impair the performance, necessitating constant recalibration.
[0018] Conversely, unlabeled data can be readily obtained for extended periods in everyday life when the participant is not following explicit instructions.
[0019] This renders this option potentially interesting to explore when restoring and / or controlling of motor functions of a patient are concerned.
[0020] Starting from this notion, the inventors investigated the possibility of effectively extracting meaningful information from the abundant unlabeled cortical data, and acquiring generic brain features with transferable properties that require minimal calibration for extension to differing contexts.
[0021] Self-supervised learning (SSL) has recently emerged as a powerful approach for learning intelligible representations of data without the need for explicit labels.
[0022] Unlike supervised learning, which largely depends on carefully annotated data and is limited by its availability, self-supervised approaches have the advantage of being able to leverage vast quantities of unlabeled data to learn generalist models that can be adapted to various tasks.
[0023] SSL has played a pivotal role in significant achievements of artificial intelligence in Natural Language Processing (NLP), ranging from machine translation to the development of large language models trained on vast corpora [1]-[4J.
[0024] SSL has also pushed new boundaries in computer vision, where SSL-pre-trained models were able to outperform models trained on labels [5],
[0025] In particular, the main concept behind SSL is to learn informative representations by defining a pretext task that derives the supervision signal from the input itself.
[0026] In the field of NLP, a frequently used SSL task involves a reconstructive objective, e.g. hiding words within a sentence and training a neural network to predict the missing words based on the surrounding context.
[0027] Through such unsupervised interactions with the data, the model acquires a general- istic understanding of the fundamental structure and intricacies of language.
[0028] These learned representations can subsequently be transferred to a diverse range of downstream tasks, e.g. including text summarization, translation, and even human interaction, e.g. ChatGPT.
[0029] This masking objective was also used in the field of computer vision, termed “masked image modeling”, by similarly masking portions of an image, and training a model to recover the original image [6]-[7] .
[0030] Also, this approach has been recently extended to learning self-supervised representations of audio spectrograms [8],
[0031] These masked autoencoders demonstrated unprecedented capacity to simultaneously learn local and global structures in the data and to achieve competitive performances on a variety of vision and speech tasks.
[0032] The opportunity of transferring these principles to BCI applications appears particularly beneficial, in particular because acquiring comprehensive labels is often challenging and / or specific tasks may not be known in advance. Recent efforts have evaluated SSL strategies in neuro-physiological data such as EEG [9]-
[0012] , neural population spikes
[0013] -
[0014] , or even across modalities
[0015]
[0033] For example, Schneider et al.
[0015] describes the use of a contrastive SSL objective to infer consistent neural embeddings by implicitly labeling over time.
[0034] Contrastive learning consists of creating a contrast in latent space by enforcing variance between incompatible or “negative” samples (e.g., samples distant in time) and promoting invariance between similar or “positive” samples (e.g., samples close in time).
[0035] In particular, these studies provide a preliminary evidence that an SSL-trained encoder could handle some neural variabilities and generate representations that outperform supervised learning in downstream tasks, in particular when there is a limited amount of labeled data available.
[0036] However, these known approaches still fail to demonstrate sufficient evidence of their ability to effectively generalize to novel tasks and / or conditions that were not included in the training set.
[0037] Additionally, these known approaches have proven uncapable of leveraging the potential of SSL to overcome distribution shifts in clinically relevant settings.
[0038] A partial solution to the aforementioned drawbacks is provided in Wang et al.
[0016] , teaching the use of a masked reconstructive objective on time-frequency representations (spectrograms) of intracranial electrodes.
[0039] This solution is adapted to learn unsupervised brain representations that improve decoding accuracy and sample efficiency across tasks involving patients listening to audio stimuli, and can generalize to different subjects.
[0040] Nonetheless, there is still room for further improvement.
[0041] In the light of the above, it is an object for the present invention to provide a system for planning and / or providing neuromodulation and / or neurostimulation, capable of harvesting unlabeled brain data to learn robust features from electrical activity across time, tasks, patients and / or recording modalities, thus allowing for providing stimulation in a more effective and targeted way. The above-specified object is achieved by the provision of a system as defined in claim 1 .
[0042] According to the invention, a system for planning and / or providing control to neuromodulation and / or neurostimulation and / or an actuator such as a brain computer interface (BCI) and / or providing a neural interface system, especially a brain-spinal-cord- interface system, is provided, said system comprising: at least one input module for brain signals, especially raw brain signals, at least one pre-processing module for converting the raw brain signals into input tensors, wherein the input tensors include at least one of temporal and / or spatial and / or spectral information, and at least one conversion module comprising an encoder and decoder network architecture, wherein the encoder and decoder network architecture is configured to be trained by using partially masked samples of the input tensors to reconstruct electrical activity of a given representation of the input tensors, wherein a masking ratio is applied with positional embedding, which provides simultaneous learning of local and global features present in the brain data, thereby transforming the input tensors into a latent space.
[0043] Also provided is a method for planning neuromodulation and / or neurostimulation through the above-described system.
[0044] The present invention provides a system for planning and / or providing neuromodulation and / or neurostimulation for a subject in need thereof.
[0045] Preferably, said subject is a human patient, e.g. affected by spinal cord injury (SCI) and / or other neurological disorders such as stroke, Parkinson’s disease, and / or other neurodegenerative disorders.
[0046] The system comprises at least one input module for brain signals.
[0047] For instance, the system may comprise a single input module. Alternatively, the system may comprise a plurality of input modules, according to the needs.
[0048] Especially, said brain signals are raw brain signals.
[0049] Preferably said brain signals are electrophysiological raw brain signals.
[0050] The at least one input module can be configured and adapted for recording brain signals through invasive techniques.
[0051] Alternatively, the at least one input module can be configured and adapted for recording brain signals through non-invasive techniques.
[0052] The system further comprises at least one pre-processing module. The at least one pre-processing module is adapted for converting the raw brain signals into input tensors.
[0053] For instance, the system may comprise a single pre-processing module.
[0054] Alternatively, the system may comprise a plurality of pre-processing modules, according to the needs.
[0055] In particular, the input tensors may include temporal information.
[0056] Additionally or alternatively, the input tensors may include spatial information.
[0057] Additionally or alternatively, the input tensors may include spectral information.
[0058] Preferably, a 3D masking strategy is implemented, randomly masking patches of spatial, spectral and temporal information.
[0059] This allows addressing the drawbacks of known approaches relying on 2D masking, which are restricted to entire frequency or time bands.
[0060] Additionally, 3D positional embeddings of the input can be incorporated, making the neural network aware of the spatial positions of electrodes relative to the brain atlas, for each patient.
[0061] This approach is beneficial in particular for the purpose of transferring knowledge across patients. The system further includes at least one conversion module. The at least one conversion module comprises an encoder and decoder network architecture.
[0062] For instance, the system may comprise a single conversion module.
[0063] Alternatively, the system may comprise a plurality of conversion modules, according to the needs.
[0064] In particular, the encoder and decoder network architecture is configured to be trained by using partially masked samples of the input tensors to reconstruct a given representation of the input tensors.
[0065] In particular, a masking ratio is applied with positional embedding, which provides simultaneous learning of local and global features present in the brain data, thereby transforming the input tensors into a latent space.
[0066] The invention is based on the basic idea that, by transforming the input tensors into a latent space, it is possible to provide the most relevant and robust features to predict behaviors including motor intentions, sensory or physiological state. In particular, the system of the invention is able to learn this transformation to produce generic features that can handle clinically relevant decoding challenges by accumulating background knowledge about the brain in an unsupervised manner.
[0067] Accordingly, decoding predictions can be improved with smaller number of labelled data and provide reliable inputs to a neuroprosthetic system to restore one or more functions.
[0068] The invention provides a new approach inspired from natural language processing that leverages unlabeled brain activity to extract efficient and generic neural features transferable across tasks and time.
[0069] This new approach allows significantly improving performance in BCI applications, while reducing the amount of required labeled data.
[0070] Advantageously, latent features that are present in brain data can be used in downstream tasks such as motor intention classification, compensation of temporal drift or patient-to-patient generalization. Brain signals can be captured either through experimental sessions or in everyday life conditions.
[0071] The system according to the invention is able to efficiently model the dynamics of change of brain signals over time.
[0072] This knowledge can be used to enhance robustness of decoding against long-term shifts in neural dynamics.
[0073] Advantageously, the encoder and decoder network architecture may be configured to perform self-supervised learning (SSL).
[0074] The use of SSL concepts in this context is beneficial since acquiring comprehensive labels is often challenging, and / or one or more specific tasks may not be known in advance.
[0075] Advantageously, the encoder and decoder network architecture is configured to learn patterns of unlabeled neural activity.
[0076] In particular, by learning patterns in unlabeled neural activity, it is possible to at least partially uncover the underlying dynamical structure of the brain, and establish a generic latent space that correlates with behavior in varying contexts.
[0077] Otherwise stated, by learning the dynamical structure of the brain, it is possible to reliably predict behavior.
[0078] Also, this approach allows overcoming difficulties and drawbacks underlying known solutions broadly relying on labeled data in annotating activities of the brain.
[0079] As mentioned, labels are often uncertain and compressed descriptions of neural activity, so that trained decoders are only enabled to rely on task-specific characteristics.
[0080] Accordingly, even minor deviations from the specific task conditions may significantly impair the performance, requiring constant recalibration.
[0081] To the contrary, unlabeled data can be readily obtained for extended periods, even in situations of everyday life while the patient is not performing a specific task and / or following explicit instructions. Advantageously, said brain signals may at least partially comprise cortical brain signals.
[0082] Preferably, said signals at least partially comprise unlabeled cortical brain signals.
[0083] Advantageously, the input tensors are 3D input tensors.
[0084] Preferably, said 3D input tensors have a spatial and / or temporal and / or spectral dimension, thereby preserving the temporal, spatial and spectral coherence of information.
[0085] Incorporating information in these 3 dimensions of logic guides the neural network to learn patterns specific to each logic, thereby increasing its efficiency and interpretability.
[0086] Advantageously, the pre-processing module is configured to use a wavelet transformation to convert the brain signals into input tensors.
[0087] The use of a wavelet transformation is particularly advantageous for the implementation of real time computation and preservation of temporal information.
[0088] Advantageously, the encoder and decoder network architecture may comprise an encoder module.
[0089] In particular, said encoder module is configured to transform the brain signals into a latent space.
[0090] There may be also a decoder module provided in the encoder and decoder network architecture.
[0091] In particular, said decoder module is configured to map the latent space to the original feature space.
[0092] The decoder module is also configured to compute the reconstruction loss.
[0093] Thus, after SSL pretraining, the learned latents can be used to visualize brain trajectories, estimate domain shifts and provide robust input features to predict behavior.
[0094] In particular, self-supervised pretraining utilizes long-term data including resting states and a range of behavioral tasks, thereby overcoming the limitations of known approaches, only relying on single-task data. By avoiding relying on task-specific data, it is possible to enhance robustness and generalizability of a model to a variety of new tasks that have not been observed during training.
[0095] Advantageously, said masking ratio is a high masking ratio with positional embedding.
[0096] A high masking ratio with positional embedding enables simultaneous learning of local and global features that are present in the brain data.
[0097] In particular, the provision of a high masking ratio in the self-supervised pretraining task largely reduces redundancy, encouraging the neural network to learn informative solutions that effectively capture high-level brain features.
[0098] Conversely, a low masking ratio - given the highly correlated nature of multi-channeled brain signals - only allows for learning of trivial solutions, that interpolate the masked patches from neighboring contexts in the input tensors, without learning interesting representations.
[0099] Further, positional embeddings enable the neural network to use certain position information (such as spatial, spectral and / or temporal positions) of the input tensor as a guiding signal for the simultaneous learning of local and global features.
[0100] The system according to the invention can be used in online, real-time clinical settings.
[0101] The present invention further provides a method for planning neuromodulation and / or neurostimulation through the system described above.
[0102] In particular, the method comprises:
[0103] - providing at least one input module for brain signals, especially raw brain signals, preferably wherein said brain signals at least partially include cortical brain signals, more preferably unlabeled cortical brain signals,
[0104] - converting the raw brain signals into input tensors, preferably 3D input tensors, through a pre-processing module, said input tensors including at least one of temporal and / or spatial and / or spectral information, - providing at least one conversion module, said conversion module comprising an encoder and decoder network architecture,
[0105] - training the encoder and decoder network architecture by using partially masked samples of the input tensors to reconstruct a given representation of the input tensors.
[0106] A masking ratio is applied with positional embedding, which provides simultaneous learning of local and global features present in the brain data, thereby transforming the input tensors into a latent space.
[0107] This enables visualizing brain trajectories and estimating domain shifts to provide the most relevant and robust input features to predict behavior.
[0108] Advantageously, the method may include implementing a 3D masking strategy, randomly masking patches of spatial, spectral and temporal information.
[0109] Additionally, the method may include incorporating 3D positional embeddings of the input, making the neural network aware of the spatial positions of the electrodes relative to the brain atlas, for each patient.
[0110] Preferably, said brain signals at least partially include cortical brain signals.
[0111] More preferably, said brain signals at least partially include unlabeled cortical brain signals.
[0112] Preferably, said input tensors are 3D input tensors.
[0113] Preferably, the dimensions of the 3D input tensors include at least one of spatial, temporal and / or spectral information.
[0114] Advantageously, the method may comprise performing a self-supervised learning (SSL).
[0115] This step is implemented through the encoder and decoder network architecture described above.
[0116] Advantageously, the method may comprise:
[0117] - learning patterns of unlabeled neural activity, thereby at least partially uncovering the underlying dynamical structure of the brain, and - establishing a generic latent space that correlates with behavior in varying contexts.
[0118] These steps are also implemented through the encoder and decoder network architecture described above.
[0119] Advantageously, the step of converting the brain signals into input tensors may be implemented using a wavelet transformation.
[0120] Advantageously, the method may comprise transforming the brain signals into a latent space.
[0121] This step is implemented through the encoder module described above.
[0122] Here, the method may also comprise:
[0123] - mapping the latent space to the original feature space, and
[0124] - computing the reconstruction loss.
[0125] These steps are implemented through the decoder module described above.
[0126] Finally, the invention provides a computer-readable non-volatile storage medium comprising computer-readable instructions that, when executed by a processing means, causes said processing means to implement the method steps described above.
[0127] Further details and advantages of the present invention shall now be disclosed in connection with the drawings, where:
[0128] Fig. 1 is a block diagram schematically showing a system for planning and / or providing neuromodulation and / or neurostimulation to a subject, e.g. a human patient, according to an embodiment of the invention.
[0129] Fig. 2 is a diagram showing how an encoder and decoder network architecture of the system of Fig. 1 implements learning of self-supervised brain representations. Here, brain cortical recordings capture the electrical activity of the neurons. This activity reflects the behavior and intentions of the patient. It can be decomposed in spatial, spectral and temporal features. The system learns the internal structure of this activity by trying to reconstruct the full data from a fraction of it. The later embedding can be viewed as a compressed version of the input that is more informative and provides higher performance in downstream tasks such as movement classification and regression.
[0130] Fig. 3 is a diagram showing how informative natural embeddings are uncovered in the system of the invention. In particular: a. masked auto-encoder framework to learn self-supervised brain representations. Transformer encoder-decoder neural networks learn to restore the spatio-spectro-temporal features from their sparse sampling; b. t-distributed stochastic neighbor embedding t-SNE projection of the learned latents with pre-training epochs across time, with each data point colored according to the day after implantation; c. t-SNE projection of the learned latents across behavioral tasks, unobserved during training, showing a modeling of lower limb, upper limb, and speech brain concepts; d. performance in decoding hip movement intentions (33.3% chance) on a moment-by-moment basis with data points of 200 ms in length. A Random Forest (RF) decoder trained on the latent features (Latent RF), gains 12 points compared to when trained on the input wavelet features (InputRF), with only 4% difference to a transformer decoder (NeuralT), suggesting that meaningful feature extraction is key; e. performance in 6-class arm direction decoding (16.7% chance) with limited labels (10 per class), of a randomly initialized transformer (RandomT) and an SSL-pre-trained transformer (NeuralT).
[0131] Fig. 4 is a diagram schematically illustrating the network’s reconstruction of three brain activity samples, including spatio-spectro-temporal information, from their randomly masked sampling, through the system of the invention.
[0132] Fig. 5 is a diagram schematically illustrating the projection of the latent and the input features in 3D using t-SNE, which is a non-linear dimensionality reduction method. The latent features reveal meaningful structures (here, different electrode configurations), that are less visible in the input features.
[0133] Fig. 6 is a diagram showing how the system of the invention allows achieving highly-decodable and efficient feature extraction. Here, the example refers to a task where a patient was attempting movements of the lower limbs, in particular walking; In particular: a. t-SNE projection of the original input features and the latent features showing increased segregation in the latent embeddings of the left and right hip movement intentions; b. On the left, performance in hip movement intention decoding (33.3% chance) with increasing number of training cues. A classical decoder (Random Forest), trained on the latent features, shows increased performance compared to when trained on the input features, especially with low number of training labels. On the right, decoding probability peaks after movement onsets, showing that the decoder, trained on the latent features, provides more confident predictions compared to the input model, even when the number of given cues is low.
[0134] Fig. 7 is a diagram also showing how the system of the invention allows achieving highly-decodable and efficient feature extraction. Here, the example refers to an upper-limb motor decoding task with a spinal-cord injured patient. In particular: a. Experimental setup of the task involving a robotic arm assisting the patient’s movements; b. Performance of a Random Forest (RF) decoder trained on the latent features or the input features as a function of the percentage of the available training cues; c. t-SNE projection of the original input features and the latent features showing increased segregation in the latent embeddings corresponding to different upper limb joints involved in the movement; d. Decoding probabilities across time in three different states (shoulder, elbow, hand); e. Confusion matrices in 8 states upper limb classification of a RF trained on either the latent or the input features, showing a gain of 13 points in accuracy when using the latent features.
[0135] Fig. 8 is a diagram showing how the latent features effectively model distribution shifts across time. In particular: a. t-SNE projection of the learned latents with pre-training epochs across time, with each data point colored according to the day after implantation; b. Distance in latent space, of each day’s centroid, relative to the centroid of day 167 after implantation; c. PCA projection of the learned latents before and after unsupervised alignment of second order statistics, with each data point colored according to the day after implantation.
[0136] Fig. 9 is a diagram showing the t-SNE projections of the learned latents before and after supervised alignment using a contrastive objective based on the be- havioral labels of left hip, right hip, and resting state. The left portion of the figure illustrates the non-aligned latent features that show a change in the neural dynamics across days for the different behavioral conditions. The right portion of the figure illustrates the latent space obtained after contrastive alignment trained on day 1 to day 10 of rehabilitation (projected from day 1 to day 32 of rehabilitation), suggesting the possibility of obtaining a latent space that is robust and invariant to shifts in neural distributions from session to session.
[0137] Fig. 10 Latent vs. Input online replay. The diagrams in the figure show experimental evidence on the ability of the system of the invention to learn a complex 6-states upper limb task based on latent features in online, real-time clinical settings. In particular: a. Upper-limb decoding probabilities for a decoder trained online using latent features compared to a decoder computed offline in pseudoonline settings using input features; b. t-SNE projections of input features (left) and latent features (right) for five different movement attempts; c. Confusion matrices showing accuracies for the pseudo-online decoder using input features (accuracy = 51 .5%) and the online decoder using latent features (accuracy = 80.7%).
[0138] Fig. 11 Generative Pre-training Transformer (GPT) framework for the system. In particular: a. Overview of the self-supervised GPT framework as used in the experimental activities described below, carried out by the Inventors using the system according to the invention. The raw neuronal activity is converted into a three-dimensional sample that includes spatial, spectral and temporal information. This sample is partially masked and used as input to an autoencoder artificial neural network that is trained to reconstruct a normalized version of the masked parts in the sample; b. Examples showing higher quality reconstruction obtained through the novel approach according to the invention compared to nearest-neighbor interpolation; c. Left, t-SNE projection of the latent representation of synthetic samples, showing an emergent structure with physiolog ically- meaningful spatial and spectral organization. Right, t-SNE projection of the latent representation of both synthetic samples and real samples recorded during a lower-limb task, showing a contralateral organization of movement states across hemispheres at high-frequency regions, with passive resting states emerging in low-frequency regions.
[0139] Fig. 12 System training - Days D1 to D98 post-implantation. In particular: a. GPT training loss (light gray) (MSE at convergence = 0.156) and evaluation loss (dark gray) (MSE at convergence = 0.155) on a separate test set. The dashed line represents the reconstruction loss using nearest neighbor interpolation (MSE = 0.481 ); b. Rank of the latent representations estimated using RankME, as a function of the masking ratio applied on the input samples during training of the system. The reported rank per masking ratio is the maximum value attained during self-supervised training. Peak of rank is reached at 75% masking ratio (RankMe = 119); c. t-SNE projections of representations of pure spatial and spectral synthetic samples, across self-supervised training cycles.
[0140] Fig. 13 System extracts relevant features for decoding motor intentions across tasks. In particular: a. Experimental setup for the lower-limb walking task. The participant performs instructed attempts of left and right-hip flexions with transitory resting periods; b. t-SNE projection of input features versus latent features for the different movement trials; c. 3-state decoding accuracy (left-hip vs. righthip vs. resting) as a function of the number of training cues, comparing decoders trained on input features versus latent features (mean, n = 10 shuffled-cues runs); d. Decoding probabilities for each movement state with 8 and 32 training cues, comparing input and latent features; e. Confusion matrices showing decoding accuracy in pseudo-online conditions for 8 training cues (input decoder accuracy: 33.5%; latent decoder accuracy: 76.4%), and 32 training cues (input decoder accuracy: 89.2%; latent decoder accuracy: 91.4%); f. Experimental setup for the joint-centered upper-limb task. The participant performs instructed attempts with exoskeleton assistance, of 8 different upper-limb joint movements. g. t-SNE projection of input features versus latent features for the different movement trials; h. 8-state decoding accuracy as a function of the number of training cues, comparing decoders trained on input features versus latent features (mean, n = 10 shuffled-cues runs); i. Decoding probabilities for each movement state with all cues available for training (128 cues), comparing input and latent features; j. Confusion matrices showing decoding accuracy in pseudo-online conditions using all available training cues. Left, decoder using input features (accuracy: 55.6%). Right, decoder using latent features (accuracy: 71.9%).
[0141] Fig. 14 Extraction of relevant features for decoding motor intentions across time. In particular: a. t-SNE projections of the latent features of resting state samples across 167 days, along system’s training cycles; b. Feature alignment framework. A projector neural network is trained using contrastive learning to align the latent features to a time-invariant representation space. The contrastive learning is performed by attracting similar movements across time, while repelling dissimilar ones; c. Left, t-SNE projections of the non-aligned latent features during lower-limb movement attempts (left-hip, right-hip, resting) across 24 days. Right, t-SNE projections of the aligned features during lower- limb movement attempts (left-hip, right-hip, resting) across 24 days. The training of the projector was performed using data from day 1 till day 10, then the projector is used to infer the aligned features of future days; d. Comparison of average decoding accuracy using k-NN decoder (k=8) trained on day 1 to 10 and tested on future days, using either the non-aligned (67.9% +- 1.3% s.e.m.) or aligned features (73.7% +- 1.2% s.e.m.) (mean, n = 10 runs). Feature alignment significantly improves the decoding stability over time (two-sided paired t- test, p < 0.001 ).
[0142] Fig. 15 Time shift. In particular: a. t-SNE embeddings of the non-aligned versus aligned features during lower-limb movement trials (left-hip, right-hip, resting) across 24 days, comparing the input features (top) and the latent features (bottom); b. t-SNE projections of the aligned features during lower-limb movement trials (left-hip, right-hip, resting) across 24 days, separately plotting the samples used for training the projector (day 1 to day 10) and the test samples (day > 10), for both the aligned input and latent features; c. Alignment (InfoNCE) loss across training iterations of the projectors, comparing the input and latent features; Dashed lines represent the losses evaluated on the test set (InfoNCE = 6.09 for latent, InfoNCE = 6.14 for input) after the end of training; d. Decoding accuracies using the input features (pre-alignment acc = 64.1 % +- 2.0% s.e.m., post-alignment acc = 68.7% +- 2.7% s.e.m.), or the latent features (prealignment acc = 67.9% +- 1.3% s.e.m., post-alignment acc = 73.7% +- 1.2% s.e.m.). The mean accuracy is reported with n=10 runs with different random seeds. The use of the latent features allows for a significant improvement in decoding stability compared to the use of the input features (two-sided paired t- test, p<0.001 ); e. Decoding accuracy using the aligned input or latent features, evaluated at different complexities of the projector neural network.
[0143] Fig. 16 Real time implementation in BCI paradigm. In particular: a. Online experimental setup. Raw neural activity is processed by the system, which transforms it into informative latent features. These features are then used to train an iterative multilinear model to decode upper-limb motor intentions in real-time clinical settings. The participant attempts five different upper-limb movements; b. Real-time decoding probabilities for the different movements of hand opening, hand closing, pronation, elbow extension and shoulder abduction; c. t-SNE projection of latent features during the different movement attempts; d. Confusion matrix reporting the six-state decoding performance achieved online (accuracy = 80.7%).
[0144] Fig. 17 Implementation of the system as a tool for neural analysis across modalities. In particular: a. data from the rat hippocampus was obtained from Grosmark and Buzsaki 2016
[0019] , Electrophysiology data was gathered during the rat's movement along a linear track measuring 1 .6 meters in either the left or right directions. Left, t-SNE projection of the neural spiking rates. Center, three-dimensional latent features. The mid-splines separate the left side from the right side of the track. Right, linear reconstruction of latent vs. t-SNE embeddings show a 22% score increase with latents; b. Local Field Potential (LFP) data from a participant with Parkinson’s Disease was gathered during walking or standing motor tasks, before and 45 minutes after medication. Left, t- SNE projection of wavelet coefficients, Center, three-dimensional latent features. The mid-splines separate the therapeutic conditions (before and after medication). Right, linear reconstruction of latent vs. t-SNE embeddings also demonstrates a 22% increase in score with the latent features.
[0145] Fig. 1 provides a schematic overview of a system 100 for planning and / or providing control to neuromodulation and / or neurostimulation and / or an actuator such as a brain computer interface (BCI) and / or providing a neural interface system, especially a brain-spinal-cord-interface system according to an embodiment of the present invention.
[0146] Preferably, the system 100 may be used for restoring and / or controlling a motor function in a subject in need thereof.
[0147] Preferably, said subject is a human patient P affected by spinal cord injury (SCI) and / or other neurological disorders such as stroke, Parkinson’s disease, and / or other neurodegenerative disorders.
[0148] The system 100 comprises at least one input module 10 for brain signals S, especially raw brain signals S, preferably be electrophysiological raw brain signals.
[0149] In the shown embodiment, the system 100 comprises a single input module 10 (Fig. 1).
[0150] Not shown is that the system 100 may comprise a plurality of input modules 10, according to the needs.
[0151] The at least one input module 10 can be configured and adapted for recording brain signals through invasive techniques.
[0152] Alternatively, the at least one input modulelO can be configured and adapted for recording brain signals through non-invasive techniques.
[0153] The system further comprises a pre-processing module 20.
[0154] The system 100 further comprises at least one pre-processing module 20.
[0155] In the shown embodiment, the system 100 comprises a single pre-processing module 20 (Fig. 1).
[0156] Not shown is that the system 100 may comprise a plurality of pre-processing module 20s, according to the needs. The pre-processing module 20 is configured to convert the raw brain signals S into input tensors.
[0157] In the present embodiment, said input tensors are 3D input tensors.
[0158] In particular, said input tensors include at least one of temporal and / or spatial and / or spectral information.
[0159] The system further comprises a conversion module 12.
[0160] In the shown embodiment, the system 100 comprises a single conversion module 12 (Fig. 1)
[0161] Not shown is that the system 100 may comprise a plurality of conversion modules 12, according to the needs.
[0162] The at least one conversion module 12 comprises an encoder and decoder network architecture 14 (Fig. 1).
[0163] In particular, the encoder and decoder network architecture 14 is configured to be trained by using partially masked samples of the input tensors, representing the brain’s electrical activity, to reconstruct a given representation of the input tensors.
[0164] A masking ratio is applied with positional embedding, which provides simultaneous learning of local and global features present in the brain data, thereby transforming the input tensors into a latent space.
[0165] Preferably, a 3D masking strategy is implemented, randomly masking patches of spatial, spectral and temporal information, preferably at a high masking ratio.
[0166] A high masking ratio with positional embedding enables simultaneous learning of local and global information.
[0167] In particular, by masking tensors at a high ratio, it is ensured that high-level and interesting representations are learned, rather than trivial interpolation.
[0168] In the present embodiment, the encoder and decoder network architecture 14 is configured to perform self-supervised learning (SSL).
[0169] Also, in the present embodiment, the encoder and decoder network architecture 14 is configured to learn patterns of unlabeled neural activity. This allows at least partially uncovering the underlying dynamical structure of the brain B, and establishing a generic latent space that correlates with behavior in varying contexts.
[0170] In the present embodiment, the brain signals S comprise at least partially cortical brain signals S, especially unlabeled cortical brain signals.
[0171] In the present embodiment, the pre-processing module 20 is advantageously configured to use a wavelet transformation to convert the brain signals S into input tensors.
[0172] In the present embodiment, the encoder and decoder network architecture 14 comprises an encoder module 16 (Fig. 1).
[0173] In particular, the encoder module 16 is configured to compress brain signals S into a latent space.
[0174] There is also a decoder module 18 provided in the encoder and decoder network architecture 14 (Fig. 1).
[0175] In particular, the decoder module 18 is configured to map the latent space to the original feature space and to compute the reconstruction loss.
[0176] By transforming the input tensors into a latent space, it is possible to visualize brain trajectories and estimate domain shifts to provide the most relevant and robust input features to predict behavior.
[0177] Fig. 2 schematically illustrates how the encoder and decoder network architecture 14 implements learning of self-supervised brain representations.
[0178] Here, brain recording devices D, e.g. recording electrodes, preferably epidural recording electrodes, are used to capture electrical activity of the neurons.
[0179] This activity reflects the behavior and intentions of the patient P. It can be decomposed in spatial, spectral and / or temporal features.
[0180] The system 100 learns the internal structure of this activity by trying to reconstruct the full data from a fraction of it. The latent embedding can be viewed as a compressed version of the input that is more informative and provides higher performance in downstream tasks, such as movement classification and regression.
[0181] Fig. 3 schematically illustrates how uncovering informative neural embeddings is implemented in the system 100 according to the present embodiment.
[0182] In particular, Fig. 3 shows the following: a. masked auto-encoder framework to learn self-supervised brain representations. Transformer encoder-decoder neural networks learn to restore the spatio- spectro-temporal features from their sparse sampling; b. t-SNE projection of the learned latents with pre-training epochs across, with each data point colored according to the day after implantation; c. t-SNE projection of the learned latents across behavioral tasks, unobserved during training, showing a modeling of lower limb, upper limb, and speech brain concepts; d. performance in decoding hip movement intentions (33.3% chance) on a mo- ment-by-moment basis with data points of 200 ms in length. A Random Forest (RF) decoder trained on the latent features (Latent RF) gains 12 points compared to when trained on the input wavelet features (InputRF), with only 4% difference to a transformer decoder (NeuralT), suggesting that meaningful feature extraction is key; e. performance in 6-class arm direction decoding (16.7% chance) with limited labels (10 per class) of a randomly initialized transformer (RandomT) and an SSL-pre-trained transformer (NeuralT).
[0183] Fig. 4 schematically shows learning of the natural language of the brain cortex, in particular the motor cortex, in the system 100 of the invention.
[0184] Here, reconstruction is carried out on three different brain activity Samples (Samples 1-3). In particular, Fig. 4 illustrates network’s reconstruction of the three brain activity Samples, including spatio-spectro-temporal information, from their randomly masked sampling.
[0185] Fig. 5 schematically shows learning of meaningful natural representations in the system 100 of the invention.
[0186] By projecting the latent features in 3D (using t-SNE, which is a dimensionality reduction method), reveals interesting structures (in the example of Fig. 5, different electrodes configurations), that are less visible when input tensor are projected in 3D.
[0187] This suggests that the latent features are potentially more efficient representations compared to the input tensors.
[0188] As mentioned, latent embeddings can be seen as a compressed version of the input that is more informative and provides higher performance in downstream tasks.
[0189] The latent features can be used in the downstream tasks, such as motor intention classification, compensation of temporal drift, or patient-to-patient generalization.
[0190] Fig. 6 shows and example referring to a movement of the lower limbs of a patient P, in particular walking.
[0191] Here, the example demonstrates how the system 100 of the invention allows achieving highly-decodable and efficient feature extraction.
[0192] In particular, Fig. 6 shows the following: a. t-SNE projection of the original input features and the latent features showing increased segregation in the latent embeddings of the left and right hip movement intentions; b. on the left, performance in hip movement intention decoding (33.3% chance) with increasing number of training cues. A classical decoder (Random Forest), trained on the latent features, shows increased performance compared to when trained on the input features, especially with low number of training labels. On the right, decoding probability peaks after movement onsets, showing that the decoder trained on the latent features, provides more confident predictions compared to the input model, even when the number of given cues is low. A further example illustrating how the system 100 allows achieving highly-decodable and efficient feature extraction is shown in Fig. 7.
[0193] Here, the example refers to robot-driven movements of the upper limbs of a patient P.
[0194] In particular, Fig. 7 shows the following: a. experimental setup of the task involving a robotic arm assisting the patient’s movements; b. performance of a Random Forest (RF) decoder trained on the latent features or the input features as a function of the percentage of the available training cues; c. t-SNE projection of the original input features and the latent features showing increased segregation in the latent embeddings corresponding to different upper limb joints involved in the movement; d. decoding probabilities across time in three different states (shoulder, elbow, hand); e. confusion matrices in 8 states upper limb classification of a RF trained on either the latent or the input features, showing a gain of 13 points in accuracy when using the latent features.
[0195] Fig. 8 shows an example of model distribution shifts across time, the context of the present invention.
[0196] Here, projection of pre-training epochs across time is shown, with each data point colored according to the day after implantation.
[0197] In particular, Fig. 8 shows the following: a. t-SNE projection of the learned latents with pre-training epochs across time, with each data point colored according to the day after implantation; b. distance in latent space, of each day’s centroid, relative to the centroid of day 167 after implantation; c. PCA projection of the learned latents before and after unsupervised alignment of second order statistics, with each data point colored according to the day after implantation.
[0198] An example of raw latent embeddings and aligned latent embeddings across time is shown in detail in Fig. 9
[0199] Here, the example refers to a right hip movement, a left hip movement, and a rest condition.
[0200] In particular, Fig, 9 shows the t-SNE projections of the learned latents before and after supervised alignment, using a contrastive objective based on the behavioral labels of left hip, right hip, and resting state.
[0201] The left part of the figure illustrates the non-aligned latent features, showing a change in the neural dynamics across days for the different behavioral conditions.
[0202] The right part of the figure illustrates the latent space obtained after contrastive alignment trained on day 1 to day 10 of rehabilitation (projected from day 1 to day 32 of rehabilitation), suggesting the possibility of obtaining a latent space that is robust and invariant to shifts in neural distributions from session to session,
[0203] As mentioned, the system 100 of the invention can be used in online, real-time clinical settings.
[0204] An online, real-time clinical settings experiment was conducted with a human patient to demonstrate this principle.
[0205] The results of this experiment are shown in Fig. 10.
[0206] In particular, the experiment revealed the ability of the system 100 to learn a complex 6-states upper limb task based on latent features, in online, real-time clinical settings.
[0207] The present invention further provides a method for planning neuromodulation and / or neurostimulation, in particular with the system 100 described above.
[0208] In particular, the method comprises: - providing at least one input module 10 for brain signals S, especially raw brain signals S, preferably wherein said brain signals S at least partially include cortical brain signals S, more preferably unlabeled cortical brain signals S,
[0209] - converting the raw brain signals S into input tensors, preferably 3D input tensors, through a pre-processing module 20, said input tensors including at least one of temporal and / or spatial and / or spectral information,
[0210] - providing at least one conversion module 12, said conversion module 12 comprising an encoder and decoder network architecture 14,
[0211] - training the encoder and decoder network architecture 14 by using partially masked samples of the input tensors to reconstruct a given representation of the input tensors.
[0212] A masking ratio is applied with positional embedding, which provides simultaneous learning of local and global features present in the brain data, thereby transforming the input tensors into a latent space.
[0213] In the present embodiment, said brain signals S at least partially include cortical brain signals S, more preferably unlabeled cortical brain signals S.
[0214] In the present embodiment, said input tensors are 3D input tensors.
[0215] Preferably, the dimensions of the 3D input tensors include at least one of spatial, temporal and / or spectral information.
[0216] In the present embodiment, the method further comprises performing a selfsupervised learning (SSL).
[0217] This step is implemented through the encoder and decoder network architecture 14 described above.
[0218] In the present embodiment, the method further comprises: learning patterns of unlabeled neural activity, thereby at least partially uncovering the underlying dynamical structure of the brain B, and establishing a generic latent space that correlates with behavior in varying contexts. These steps are as well implemented through the encoder and decoder network architecture 14 described above.
[0219] In the present embodiment, the step of converting the brain signals S into input tensors is conveniently implemented by using a wavelet transformation.
[0220] In the present embodiment, the method further comprises transforming the brain signals S into a latent space.
[0221] This step is implemented thorough the encoder module 16 described above.
[0222] The method further comprises:
[0223] - mapping the latent space to the original feature space, and
[0224] - computing the reconstruction loss.
[0225] These steps are implemented thorough the decoder module 18 described above.
[0226] The present invention further provides a computer-readable non-volatile storage medium comprising computer-readable instructions that, when executed by a processing means, causes said processing means to implement the method steps described above.
[0227] Experimental evidence
[0228] Premises
[0229] Experimental activities have been carried out by the Inventors using the system according to the invention, which are described in detail in the following.
[0230] Deciphering the language of the human brain remains a central challenge in neuroscience.
[0231] The human brain manifests this language as firing patterns of individual neurons that collectively encode commands to enact neurological functions, including movement.
[0232] One implication of elucidating the language through which neurons encode movement commands is the possibility to decipher the information encoded by neurons to operate prosthetic systems. Self-supervised learning (SSL) methodologies have been instrumental in modeling complex statistical structures, including human language [1],
[0012] ,
[0017] -
[0018] ,
[0233] SSL-trained natural language processing models typically learn linguistic statistical structures by predicting masked (hidden) elements in sequences of words.
[0234] Similarly, the Inventors reasoned that SSL methodologies could be utilized to learn the intrinsic statistical structure of neuronal activity patterns by predicting omitted segments in continuous neuronal recordings.
[0235] The Inventors hypothesized that this learning objective would support the transformation of neuronal activity into informative representations or latent variables. In turn, these latent variables can be used as inputs for downstream decoders to infer movement commands and operate neuroprostheses.
[0236] To test this hypothesis, the Inventors leveraged SSL to train an artificial neural network to learn the statistical structure underlying local field potential signals recorded from the human cerebral cortex.
[0237] The Inventors evaluated whether the resulting latent variables provided informative features from which movement intentions could be efficiently decoded across tasks and time.
[0238] This work new approach offers insights into the structure of brain activity that can be captured through self-supervised learning and opens the door to large brain language models.
[0239] Generative Pre-training Transformer (GPT) framework
[0240] The Inventors aimed to develop a framework capable of learning the statistical structure underlying neuronal activity in humans.
[0241] Neuronal activity is intrinsically constrained by the connectivity between neighboring neurons and distant brain regions, which defines spatial and temporal relationships across local and global scales. These constraints resemble the relational structure in word sequences that characterize human language.
[0242] Given the success of masked language modeling approaches
[0020] -
[0022] in capturing such relational structures, the Inventors hypothesized that a similar solution could effectively model the structure underlying human neuronal activity.
[0243] Central to the design of masked modeling approaches are: (i) the structure of the input signal and its decomposition into basic elements, comparable to phonemes in sequences of words; (ii) the strategy by which these elements are masked, and (iii) the architecture and learning objective of the statistical model. These requirements were adapted to the properties of neuronal recordings.
[0244] (i) Neuronal signals have been converted into a three-dimensional sample, embedding spectral, temporal, and spatial information. Concretely, the Inventors computed spectrograms from temporal changes in each neural signal, which determined the first 2 dimensions of the sample. Information was explicitly integrated on the spatial locations of the electrodes as the third dimension of the sample (Fig. 11a). This strategy preserved the spatial distribution of neuronal activity while capturing the spectral and temporal dynamics of the signals. To define basic elements, the sample has been decomposed into non-overlapping 3D patches. The configuration of these patches can be adapted to the intrinsic properties of the recording modality, since single-unit, multiunit, local field potential, electrocorticograms, and electroencephalograms present unique features.
[0245] (ii) The Inventors aimed to ensure that the model simultaneously learns relationships across space, time, and frequency. To enforce this multi-dimensional learning, a variable fraction of the patches was masked based on random distributions applied uniformly across the three dimensions.
[0246] (iii) An autoencoder artificial neural network was used including 124 million parameters to compress the input data into a lower dimensional latent representation through an encoder and to reconstruct the original data from this latent representation through a decoder. The encoder employs a transformer architecture designed to transform the sequences of visible patches, referenced by their three-dimensional position in the sample, into a latent representation. The decoder is then trained to reconstruct a normalized version of the masked patches from these latent representations. A lowdimensional subset of the representations generated by the encoder contains high- level features of the neuronal signals, referred to as latent features. Finally, these latent features can be used as inputs to a downstream decoder to infer movement commands that are suitable to operate neuroprostheses.
[0247] Training
[0248] The Inventors aimed to evaluate the system on recordings obtained in humans.
[0249] For this purpose, the system was trained using electrocorticographic (ECoG) signals that were recorded from two 5-cm diameter circular devices integrating a grid of 64 recording electrodes. These devices were surgically implanted over the epidural surface of the left and right sensorimotor cortex in an individual with partial upper and lower limb paralysis due to a cervical (C5) spinal cord injury (Clincialtrial.gov, NTC04632290).
[0250] The Inventors acquired 110 hours of ECoG signals (585Hz) over the course of one year while the participants performed a wide variety of tasks, including natural behaviors and instructed tasks (Fig. 12a). A wavelet decomposition was applied to convert the ECoG signals into a three-dimensional sample that embeds spectral, temporal, and spatial information of the raw neuronal signals. These samples were used as inputs to train the system.
[0251] First, the convergence and performance of the system were tested. To this end, the network has been trained using 50 hours of recordings (Fig. 12). The parameters of the neural network converged within 40 training epochs with less than 1 % decrease per additional epoch (Fig. 12).
[0252] The system generated reconstructions with drastically improved accuracy (MSE=0.156 for training loss and MSE = 0.155 for testing loss) compared to reconstructions based on linear or nearest-neighbor interpolations (MSE=0.48) (Fig. 12). The improvement in performance attained by the system over interpolation methodologies suggested that the network learned to infer representations capturing complex, high-level relational structures within the data, rather than relying solely on local dependencies (Fig. 12).
[0253] The informational richness of the learning may be quantified by the effective rank of the latent representation space (
[0023] ).
[0254] The Inventors found that the rank peaked for a masking ratio of 75%, suggesting that this masking ratio provided an optimal balance between the complexity of the reconstruction and the informational richness of the latent space.
[0255] These results show that system effectively modeled the relational structure embedded in ECoG signals generated by the human cerebral cortex. The outcome of the model is an informative latent representation, including latent features that capture high-level properties of the neuronal activity.
[0256] Interpretation
[0257] Understanding how artificial neural networks process and prioritize features is essential to ensure that the network models physiologically relevant information. Therefore, the Inventors aimed to interpret the features learned by system, specifically to investigate whether and how the new approach implemented by the system models the spatial and spectral dependencies of the recorded neuronal signals.
[0258] To answer this question, samples consisting of pure frequencies ranging from 1 to 200 Hz and distinct spatial locations spanning the entire space of the sample have been synthesized.
[0259] Then these synthetic samples have been transformed through the encoder and the emerging structure of the latent features has been visualized using t-SNE (Fig. 11c).
[0260] This visualization revealed that the system learned the spatial organization of the electrodes over the cerebral cortex, notably segregating samples from electrodes distributed over the left versus right hemispheres. Moreover, the system modeled relational structures across the different frequency bands and spatial locations. Next, the Inventors investigated how real samples recorded during instructed behavioral tasks may fit within this emergent structure. Specifically, the Inventors acquired ECoG samples during rest and while the participant attempted to perform flexions of his left and right hips.
[0261] This visualization revealed that the relevant features for each movement were segregated within the regions representing the hemisphere contralateral to the attempted movement, and were attracted to the high-frequency regions (Fig. 11d). In contrast, transitory resting periods emerged in low-frequency regions equidistant to both hemispheres.
[0262] These emergent properties coincide with the well-established principles through which the human cerebral cortex encodes rest and movement.
[0263] Collectively, this interpretation of the latent features extracted by the neural network indicates that the system learned the physiologically relevant spectral and spatial structures inherently embedded in the natural activity of the human brain.
[0264] Enablement of robust inference of motor behaviors
[0265] Next, the Inventors aimed to quantify the potential benefits of latent features generated by the system to decode movement intentions in the perspective of operating neu- roprosthetic systems to regain movement after paralysis.
[0266] Two sets of experiments have been conducted with the same participant.
[0267] First, the participant was asked to perform alternating flexions of the left and right lower limbs during attempts to walk (Fig. 13a).
[0268] Compared to the wavelet features inputs, the latent features extracted from this sample demonstrated a more structured and discernible distribution of movement states when visualized using t-SNE projection (Fig. 13b).
[0269] To quantify this enhanced discriminative power, the Inventors measured the accuracy of movement state predictions generated by a multilinear decoder. The Inventors found that the accuracy of the decoder was significantly improved when latent features were used, especially when a limited amount of repetitions was used to train the decoder (Figs. 13c-d).
[0270] With as few as eight repetitions, the decoder utilizing the latent features achieved a 75% decoding accuracy, whereas the accuracy ceiled at 20% when the input of the sample were used.
[0271] Second, the superiority of this approach has been tested in a more complex task wherein the participant was instructed to perform single-joint movements assisted by a robotic exoskeleton.
[0272] Concurrently, the participant attempted shoulder rotations, elbow flexions, elbow extensions, hand opening, and hand closing. As observed during attempts to walk, visualization of movement states using t-SNE revealed a greater segregation of the 8 states when latent features were used compared to the wavelet features of the sample.
[0273] This improved segregation translated into a pronounced increase in decoding accuracy with latent features (71.8% accuracy) compared to input samples (55.6% accuracy).
[0274] These two experiments showed that the latent features generated by the system outperformed input samples to decode the intended movements of the lower and upper limbs in an individual with paralysis, suggesting that the system learned informative features from which movement intentions could be efficiently decoded across behavioral tasks.
[0275] Achievement of transferable feature extraction across time
[0276] A recurrent challenge in brain-controlled prostheses is the maintenance of stable decoding accuracy despite the non-stationarity of neuronal signals. These signals evolve over time due to alterations of electrode-tissue interfaces, inherent shifts in brain dynamics, and user-dependent adaptations of control strategies. The Inventors hypothesized that, by observing the temporal evolution of neuronal signals through self-supervised learning, the system could develop an internal model of these dynamics. This internal model could in turn be leveraged to enable more robust and stable decoding of movement intentions over time.
[0277] To test this hypothesis, weekly recordings of the neuronal activity have been acquired during rest (sitting with eyes closed) from the same participant over a period of 167 days.
[0278] Then, the latent representations generated by the system have been projected using t- SNE (Fig. 14a).
[0279] Through the successive training iterations, the system progressively learned a latent space that temporally organized the samples into distinct clusters. The emergence of this structure suggested that the self-supervised exposure to evolving neuronal signals enabled the system to elaborate an internal model of temporal shifts in these signals.
[0280] These results suggested that this internal model could be leveraged to compensate for evolving neuronal signals, and thus achieve stable decoding accuracy of movement intentions over time. Indeed, visualization of the latent features underlying attempts to perform alternating flexions of the left and right lower limbs over a period of 24 days confirmed temporal shifts in motor-related neuronal activity over time (Fig. 14b).
[0281] Since the system learned these temporal shifts, the Inventors asked whether this knowledge could support the alignment of the latent features in a time-invariant space.
[0282] To answer the question, the Inventors trained a multilayer perceptron projector based on contrastive learning
[0015] (Fig. 14b) to enforce the alignment of latent features into a representation space that minimizes time-related variance while preserving movement-related variance. Concretely, the network maximized the cosine similarity between randomly selected samples corresponding to the same movements performed at different times, and minimized this metric between randomly selected samples corresponding to different movements. The projector was trained from data acquired during the first 10 days. Compared to the nonaligned latent features, this alignment led to a significant increase in the spatial consistency in later testing days, as quantified by the average k-nearest neighbor (KNN) decoding accuracy (Fig. 14c). Applying the same pipeline to input features led to a poor alignment of features and reduced decoding performance, even when increasing the number of parameters of the projector network (Fig. 15).
[0283] These results show that by modeling the evolving temporal dynamics of neuronal signals, the system enabled compensation of non-stationarity neuronal signals to achieve stable decoding of movement intentions — addressing one of the key challenges for long-term operability of brain-controlled prosthesis.
[0284] Improvement of real-time decoding of prosthetic actions
[0285] Furthermore, the Inventors aimed to evaluate the superiority of this new approach for real-time control of a neuroprosthetic system in a clinical environment.
[0286] To conduct this evaluation, the system was trained using ECoG signals from a second participant who presented with pronounced impairments of upper-limb movements due to a cervical spinal cord injury (NTC05665998).
[0287] The participant was implanted with the same 5-cm diameter circular device integrating a grid of 64 recording electrodes over the arm and hand region of the sensorimotor cortex, and a neurostimulation platform to deliver epidural electrical stimulation over the cervical spinal cord — the region that produces arm and hand movements. The neuroprosthetic system was designed to translate arm and hand movement intentions decoded from ECoG signals into stimulation commands delivered to the cervical spinal cord to elicit the intended movements (Fig. 16).
[0288] 60 hours of continuous ECoG signals (585Hz) have been acquired over the course of 3 months while the participants performed a wide variety of tasks, including natural behaviors and instructed tasks.
[0289] The inventors implemented a real-time decoding paradigm to control six possible movements. The participant was cued to attempt five elementary movements, including shoulder abduction, elbow extension, pronation, hand opening, and hand closing (Figs. 10a and 16b). Latent features extracted by system were fed in real-time to a Markov-Switching multilinear decoder that iteratively updated the regression parameters every 15 seconds.
[0290] As little as 10 repetitions per movement were sufficient to train a decoder that predicted movement intentions with an accuracy of 80.7% on average (Figs. 10c and 16d). In comparison, the same decoder trained using wavelet features only achieved an accuracy of 51.5% on average (Fig. 10c). This decoding framework was subsequently implemented routinely to improve the recovery of arm and hand movements in this participant (BCI award 2024, 2nd place, www.bci-award.com / 2024).
[0291] The effective deployment of this novel approach in a clinical environment to augment the performance of a brain-controlled neuroprosthesis had many implications: (i) the methodological framework underlying this approach was generalizable to a new individual and different tasks; (ii) the approach enabled high-performance decoding of arm and hand movement intentions; (iii) the approach drastically reduced the time required to calibrate task-specific decoders, thus minimizing the efforts for the user of the neu- roprosthetic system; (iv) The computational demands of this approach were compatible with real-time implementation on conventional hardware; and (v) The neuropros- thetic system augmented with the system of the invention was subsequently used daily to restore volitional control of otherwise paralyzed arm and hand movements.
[0292] Generalization of the system to different recording modalities
[0293] Finally, the Inventors sought to demonstrate whether the system architecture was suited to integrate different recording modalities.
[0294] The system trained over hippocampus data recorded in rodents
[0019] demonstrates that the network could learn to represent the position of the animals along a linear maze (Fig. 17a).
[0295] Moreover, the system trained over Local Field Potential data from deep brain stimulation probes in a participant with Parkinson’s disease could learn to segregate movement versus rest in the different medication conditions (Fig. 17b). These experiments demonstrate that the system architecture is suited to integrate various recording modalities during different tasks.
[0296] Discussion
[0297] The Inventors leveraged self-supervised learning to design an artificial neural network that is capable of learning the intrinsic structure embedded in neuronal signals generated by the human brain.
[0298] The above-described experiments revealed that this statistical model achieved high- performance decoding of movement intentions in real-time across diverse participants, tasks, modalities, locations, and time.
[0299] This approach is now routinely implemented to restore walking and upper limb movements in people with paralysis.
[0300] Here, the implications of this approach for understanding the language of the human brain and for various applications, including brain-computer interfaces, are discussed.
[0301] The transformative impact of self-supervised learning in the fields of computer vision and natural language processing cannot be overstated. Unlike traditional supervised approaches that require extensive labeled datasets, self-supervised learning exploits the inherent structure and patterns embedded within data to generate informative representations. This learning strategy supported the design of statistical models that learned generalist representations embedded in large amounts of unlabeled images [2], [6],
[0024] -
[0025] , videos
[0026] -
[0028] , and text [1]-[2],
[0029] - equaling or even surpassing human performance in many tasks.
[0302] By analogy, the brain generates highly nonlinear yet highly structured patterns of activity that unfold as spatial and temporal sequences of electrical signals resembling words and sentences. Therefore, it has been reasoned that self-supervised learning strategies are poised to model the syntactic and semantic rules governing the generation of the human brain language. These strategies are particularly well-suited to model neuronal signals, since explicitly labeling the continuous flow of neuronal activity generated by the brain is a daunting task. In contrast, current decoders for brain-computer interfaces are trained using supervised learning approaches with explicit labeling of behaviors to control various actuators, such as computer cursors
[0030] -
[0032] , robotic arms
[0033] -
[0034] or the patient’s own limbs through electrical stimulation of the nervous system
[0035] -
[0037] , Decoding behavioral intention from neural signals is a crucial aspect of BCI technologies suffering from several limitations. This labelling process requires labor-intensive and time-consuming repetitions of predefined yet potentially ambiguous behaviors, such as “closing the hand” or “extending the arm”. These labels are sparse, imprecise, task-specific, context-dependent, and provide low-bandwidth descriptions of brain activity.
[0303] Instead, the artificial neural network designed by the Inventors has proven capable to learn the statistical structure embedded within neuronal signals acquired during natural conditions. The model thus extracted the structure and semantics underlying neuronal signals without the need of instructions and explicit labels, preserving information that would otherwise be discarded in instructed supervised paradigms.
[0304] Through this self-supervised learning, physiologically relevant spectral and spatial structures inherent to the natural activity of the human brain are captured. The consequence of this learning was the possibility to calibrate decoders of movement intentions with minimal patient involvement and to attain accuracies that exceeded performance of previous decoding frameworks.
[0305] The system was trained on neuronal signals from each patient individually, and remained limited to a few numbers of tasks and objectives. However, the ability to learn the physiological properties of human brain language has implications that expand beyond brain-computer interface technologies.
[0306] Indeed, the properties of the system may help uncover new principles of natural brain communication, the impact of various neurological conditions on the disruption of this communication, and the definition of new treatments to alleviate neurological conditions. Achieving this broader impact will require scaling up the model by incorporating multiple neuronal signal modalities and ubiquitous recording locations. The model will have to be trained on neuronal signals acquired during a variety of natural behaviors that encompass a comprehensive spectrum of sensory, motor, cognitive, and emo- tional states. To ensure robustness and generalizability, the model will have to be fed with data from large cohorts of people with both neurotypical and neurodiverse profiles as well as varying neurological conditions. This endeavor can be compared to the integration of different accents, pronunciation, languages and contexts in natural language processing. Moreover, the simultaneous recording of behavioral variables such as videos, sounds, and physiological states can provide additional information that capture the external context of the neural recordings and enable fine-tuning of the network.
[0307] Ultimately, this novel approach may help solve tasks that cannot be addressed by current machine learning approaches. For example, the human brain transforms the continuous flow of multifaceted sensory information and cognitive intentions into executive commands to produce a vast library of adaptive movements that cannot be achieved with current robotic control schemes.
[0308] A multimodal-aware brain language model as defined in the present invention could learn the transformation between rich sensory perceptions and precise actuator commands, leading to novel controllers for robots and other types of machines.
[0309] Methods
[0310] The time-frequency representation of the 64-channel raw ECoG signals have been extracted using complex continuous wavelet transform (CCWT). Morlet wavelets were used centered around specific frequencies (2, 5:5:100, 125, 150, 200 Hz). The norm of the transform was downsampled to 10 Hz and decomposed into epochs of 1 second long with 100 ms shifts, forming our input tensors X e IR 24x64x10. The input consists of a 3-dimensional tensor, one dimension representing the time, another the frequencies and the last one the spatial components, namely the electrodes. Inspired by the masked autoencoding method from computer vision, each input tensor is decomposed into non-overlapping patches consisting of 2 frequencies, 4 electrodes, and 5 time points, which are then embedded by a linear projector to form the input tokens. Fixed 3D sin-cos positional embeddings are added to the input tokens, which encode their spatial, spectral, and temporal positions. In the present experiment, the system’s en- coder consists of 12 standard transformer layers. Those layers rely on the attention mechanism described in the work from Vashwani. Such mechanism allows each patch to attend to other patches and to update its representation based on those. In alignment with the methodology proposed by He et al. (2021 ) [6], 75% of the tokens are masked at random, and only the remaining visible ones are processed through the encoder layers. This approach reduces the length of the input by 75% which in turn effectively reduces the computational complexity that is quadratic with respect to the number of tokens. The decoder is also composed of conventional transformer blocks, with trainable mask tokens padded to the output of the encoder with added positional embeddings that inform the network about the position of the masked tokens in the original input tensor.
[0311] In this novel approach, the autoencoder directly learns to reconstruct a normalized version of the input tensors consisting of raw signals, obtained by applying z-scores over each frequency band, per recording session. With this approach, the network particularly learns to reveal higher-frequency oscillations (typically exponentially decayed in power). Furthermore, this guides the learning of specificities in each frequency band across different contexts (tasks or days) as the network has these additional normalization factors to predict.
[0312] References
[0313] 100 System
[0314] 10 Input module
[0315] 12 Conversion module
[0316] 14 Encoder I decoder network architecture
[0317] 16 Encoder module
[0318] 18 Decoder module
[0319] 20 Pre-processing module
[0320] B Brain
[0321] D Brain recording devices (recording electrodes)
[0322] S Brain signals
[0323] P Patient
[0324] Bibliography
[0325] [1] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of DeepBi- directional Transformers for Language Understanding.” arXiv, May 24, 2019. doi: 10.48550 / arXiv.1810.04805.
[0326] [2] T. B. Brown et al., “Language Models are Few-Shot Learners.” arXiv, Jul. 22, 2020. doi: 10.48550 / arXiv.2005.14165.
[0327] [3] OpenAI, “GPT-4 Technical Report.” arXiv, Mar. 27, 2023. doi:
[0328] 10.48550 / arXiv.2303.08774.
[0329] [4] M. Popel et al., “Transforming machine translation: a deep learning system reaches news translation quality comparable to human professionals,” Nat. Commun., vol.
[0330] 11 , no. 1 , Art. no. 1 , Sep. 2020, doi: 10.1038 / s41467-020-18073-9.
[0331] [5] M. Oquab et al., “DINOv2: Learning Robust Visual Features without Supervision.” arXiv, Apr. 14, 2023. doi: 10.48550 / arXiv.2304.07193.
[0332] [6] K. He, X. Chen, S. Xie, Y. Li, P. Dollar, and R. Girshick, “Masked Autoencoders Are Scalable Vision Learners.” arXiv, Dec. 19, 2021. doi: 10.48550 / arXiv.2111 .06377.
[0333] [7] Z. Xie et al., “SimMIM: A Simple Framework for Masked Image Modeling.” arXiv, Apr. 17, 2022. doi: 10.48550 / arXiv.2111 .09886.
[0334] [8] P.-Y. Huang et al., “Masked Autoencoders that Listen.” arXiv, Jan. 12, 2023. doi: 10.48550 / arXiv.2207.06405.
[0335] [9] H. Banville, I. Albuquerque, A. Hyvarinen, G. Moffat, D.-A. Engemann, and A. Gramfort, “Self-Supervised Representation Learning from Electroencephalography Signals,” in 2019 IEEE 29th International Workshop on Machine Learning for Signal Processing (MLSP), Oct. 2019, pp. 1-6. doi: 10.1109 / MLSP.2019.8918693.
[0336]
[0010] D. Kostas, S. Aroca-Ouellette, and F. Rudzicz, “BENDR: using transformers and a contrastive self-supervised learning task to learn from massive amounts of EEG data.” arXiv, Jan. 28, 2021. doi: 10.48550 / arXiv.2101 .12037.
[0011] M. N. Mohsenvand, M. R. Izadi, and P. Maes, “Contrastive Representation Learning for Electroencephalogram Classification,” in Proceedings of the Machine Learning for Health NeurlPS Workshop, PMLR, Nov. 2020, pp. 238-253. Accessed: Jul. 01 , 2023. [Online], Available: https: / / proceedings.mlr.press / v136 / mohsenvand20a.html
[0337]
[0012] H.-Y. S. Chien, H. Goh, C. M. Sandino, and J. Y. Cheng, “MAEEG: Masked Autoencoder for EEG Representation Learning.” arXiv, Oct. 27, 2022. doi: 10.48550 / arXiv.2211 .02625.
[0338]
[0013] C. Pandarinath et al., “Inferring single-trial neural population dynamics using sequential auto-encoders,” Nat. Methods, vol. 15, no. 10, Art. no. 10, Oct. 2018, doi: 10.1038 / S41592-018-0109-9.
[0339]
[0014] M. R. Keshtkaran et al., “A large-scale neural network training framework for generalized estimation of single-trial population dynamics,” Nat. Methods, vol. 19, no. 12, Art. no. 12, Dec. 2022, doi: 10.1038 / s41592-022-01675-0.
[0340]
[0015] S. Schneider, J. H. Lee, and M. W. Mathis, “Learnable latent embeddings for joint behavioural and neural analysis,” Nature, vol. 617, no. 7960, Art. no. 7960, May 2023, doi: 10.1038 / S41586-023-06031 -6.
[0341]
[0016] C. Wang et al., “BrainBERT: Self-supervised representation learning for intracranial recordings.” arXiv, Feb. 28, 2023. doi: 10.48550 / arXiv.2302.14367.
[0342]
[0017] Bao, Hangbo, Li Dong, Songhao Piao, and Furu Wei. 2022. “BEiT: BERT PreTraining of Image Transformers.” arXiv. https: / / doi.org / 10.48550 / arXiv.2106.08254.
[0343]
[0018] Liu, Yinhan, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. “RoBERTa: A Robustly Optimized BERT Pretraining Approach.” arXiv. https: / / doi.Org / 10.48550 / arXiv.1907.11692.
[0344]
[0019] Grosmark, Andres D., and Gydrgy Buzsaki. 2016. “Diversity in Neural Firing Dynamics Supports Both Rigid and Learned Hippocampal Sequences.” Science (New York, N. Y.) 351 (6280): 1440-43. https: / / doi.org / 10.1126 / science.aad1935.
[0345]
[0020] Wang, Thomas, Adam Roberts, Daniel Hesslow, Teven Le Scao, Hyung Won Chung, Iz Beltagy, Julien Launay, and Colin Raffel. 2022. “What Language Model Ar- chitecture and Pretraining Objective Work Best for Zero-Shot Generalization?” arXiv. https: / / doi.org / 10.48550 / arXiv.2204.05832.
[0346]
[0021] Raffel, Colin, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2023. “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.” arXiv. https: / / doi.Org / 10.48550 / arXiv.1910.10683.
[0347]
[0022] Tay, Yi, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, et al. 2023. “UL2: Unifying Language Learning Paradigms.” arXiv. https: / / doi.org / 10.48550 / arXiv.2205.05131.
[0348]
[0023] Garrido, Quentin, Randall Balestriero, Laurent Najman, and Yann Lecun. 2023. “RankMe: Assessing the Downstream Performance of Pretrained Self-Supervised Representations by Their Rank.” arXiv. https: / / doi.org / 10.48550 / arXiv.2210.02885.
[0349]
[0024] Wei, Yixuan, Han Hu, Zhenda Xie, Zheng Zhang, Yue Cao, Jianmin Bao, Dong Chen, and Baining Guo. 2022. “Contrastive Learning Rivals Masked Image Modeling in Fine-Tuning via Feature Distillation.” arXiv. https: / / doi.org / 10.48550 / arXiv.2205.14141.
[0350]
[0025] Dong, Xiaoyi, Jianmin Bao, Ting Zhang, Dongdong Chen, Weiming Zhang, Lu Yuan, Dong Chen, Fang Wen, Nenghai Yu, and Baining Guo. 2022. “PeCo: Perceptual Codebook for BERT Pre-Training of Vision Transformers.” arXiv. https: / / d0i.0rg / l 0.48550 / arXiv.2111 .12710.
[0351]
[0026] Tong, Zhan, Yibing Song, Jue Wang, and Limin Wang. 2022. “VideoMAE: Masked Autoencoders Are Data-Efficient Learners for Self-Supervised Video PreTraining.” arXiv. https: / / doi.org / 10.48550 / arXiv.2203.12602.
[0352]
[0027] Feichtenhofer, Christoph, Haoqi Fan, Bo Xiong, Ross Girshick, and Kaiming He. 2021. “A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning.” arXiv. https: / / doi.org / 10.48550 / arXiv.2104.14558.
[0353]
[0028] Schiappa, Madeline C., Yogesh S. Rawat, and Mubarak Shah. 2023. “Self- Supervised Learning for Videos: A Survey.” ACM Computing Surveys 55 (13s): 1-37. https: / / doi.Org / 10.1145 / 3577925.
[0029] Radford, Alec, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. “Improving Language Understanding by Generative Pre-Training.” https: / / cdn.openai.com / research-covers / language- unsupervisedZlanguage_understanding_paper.pdf.
[0354]
[0030] Willett, Francis R., Darrel R. Deo, Donald T. Avansino, Paymon Rezaii, Leigh R. Hochberg, Jaimie M. Henderson, and Krishna V. Shenoy. 2020. “Hand Knob Area of Premotor Cortex Represents the Whole Body in a Compositional Way.” Cell 181 (2): 396-409.e26. https: / / doi.Org / 10.1016 / j.cell.2020.02.043.
[0355]
[0031] Pandarinath, Chethan, Paul Nuyujukian, Christine H Blabe, Brittany L Sorice, Jad Saab, Francis R Willett, Leigh R Hochberg, Krishna V Shenoy, and Jaimie M Henderson. 2017. “High Performance Communication by People with Paralysis Using an In- tracortical Brain-Computer Interface.” Edited by Sabine Kastner. eLife 6 (February)^ 8554. https: / / doi.Org / 10.7554 / eLife.18554.
[0356]
[0032] Nicolelis, Miguel A. L., and Mikhail A. Lebedev. 2009. “Principles of Neural Ensemble Physiology Underlying the Operation of Brain-Machine Interfaces.” Nature Reviews Neuroscience 10 (7): 530-40. https: / / doi.org / 10.1038 / nrn2653.
[0357]
[0033] Hochberg, Leigh R., Daniel Bacher, Beata Jarosiewicz, Nicolas Y. Masse, John D. Simeral, Joern Vogel, Sami Haddadin, et al. 2012. “Reach and Grasp by People with Tetraplegia Using a Neurally Controlled Robotic Arm.” Nature 485 (May):372.
[0358]
[0034] Benabid, Alim Louis, Thomas Costecalde, Andrey Eliseyev, Guillaume Charvet, Alexandre Verney, Serpil Karakas, Michael Foerster, et al. 2019. “An Exoskeleton Controlled by an Epidural Wireless Brain-Machine Interface in a Tetraplegic Patient: A Proof-of-Concept Demonstration.” The Lancet Neurology. https: / / doi.Org / 10.1016 / S1474-4422(19)30321 -7.
[0359]
[0035] Ajiboye, A Bolu, Francis R Willett, Daniel R Young, William D Memberg, Brian A Murphy, Jonathan P Miller, Benjamin L Walter, et al. 2017. “Restoration of Reaching and Grasping Movements through Brain-Controlled Muscle Stimulation in a Person with Tetraplegia: A Proof-of-Concept Demonstration.” The Lancet 389 (10081 ): 1821— 30. https: / / doi.Org / 10.1016 / S0140-6736(17)30601 -3.
[0036] Selfslagh, Aurelie, Solaiman Shokur, Debora S. F. Campos, Ana R. C. Donati, Sabrina Almeida, Seidi Y. Yamauti, Daniel B. Coelho, Mohamed Bouri, and Miguel A. L. Nicolelis. 2019. “Non-lnvasive, Brain-Controlled Functional Electrical Stimulation for Locomotion Rehabilitation in Individuals with Paraplegia.” Scientific Reports 9 (1 ): 6782. https: / / doi.Org / 10.1038 / s41598-019-43041 -9.
[0360]
[0037] Lorach, Henri, Andrea Galvez, Valeria Spagnolo, Felix Martel, Serpil Karakas,
[0361] Nadine Intering, Molywan Vat, et al. 2023. “Walking Naturally after Spinal Cord Injury Using a Brain-Spine Interface.” Nature 618 (7963): 126-33. https: / / doi.Org / 10.1038 / s41586-023-06094-5.
Claims
Claims1. A system (100) for planning and / or providing control to neuromodulation and / or neurostimulation and / or an actuator such as a brain computer interface (BCI) and / or providing a neural interface system, especially a brain-spinal-cord- interface system, the system (100) comprising:- at least one input module (10) for brain signals (S), especially raw brain signals (S),- at least one pre-processing module (20) for converting the raw brain signals (S) into input tensors, wherein the input tensors include at least one of temporal and / or spatial and / or spectral information, and- at least one conversion module (12) comprising an encoder and decoder network architecture (14),- wherein the encoder and decoder network architecture (14) is configured to be trained by using partially masked samples of the input tensors to reconstruct electrical activity of a given representation of the input tensors, wherein a masking ratio is applied with positional embedding, which provides simultaneous learning of local and global features present in the brain data, thereby transforming the input tensors into a latent space.
2. The system (100) according to claim 1 , characterized in that the encoder and decoder network architecture (14) is configured to perform self-supervised learning (SSL).
3. The system (100) according to claim 1 or claim 2, characterized in that the encoder and decoder network architecture (14) is configured to learn patterns of unlabeled neural activity, thereby at least partially uncovering the underlyingdynamical structure of the brain (B) and to establish a generic latent space that correlates with behavior in varying contexts.
4. The system (100) according to any one of the preceding claims, characterized in that the brain signals (S) comprise at least partially cortical brain signals (S), especially unlabeled cortical brain signals (S).
5. The system (100) according to any one of the preceding claims, characterized in that the input tensors are 3D input tensors, preferably wherein the dimensions include at least one of spatial, temporal and / or spectral information.
6. The system (100) according to any one of the preceding claims, characterized in that the pre-processing module (20) is configured to use a wavelet transformation to convert the brain signals (S) into input tensors.
7. The system (100) according to any one of the preceding claims, characterized in that the encoder and decoder network architecture (14) comprises an encoder module (16), which is configured to compress the brain signals (S) into a latent space.
8. The system (100) according to claim 7, characterized in that the encoder and decoder network architecture (14) comprises a decoder module (18), which is configured to map the latent space to the original feature space and to compute the reconstruction loss.
9. The system (100) according to any one of the preceding claims, characterized in that the masking ratio is a high masking ratio with positional embedding.
10. A method for planning neuromodulation and / or neurostimulation through the system (100) according to any one of the preceding claims, the method comprising:- providing at least one input module (10) for brain signals (S), especially raw brain signals (S), preferably wherein said brain signals (S) at least partially include cortical brain signals (S), more preferably unlabeled cortical brain signals (S),- converting the raw brain signals (S) into input tensors, preferably 3D input tensors, through a pre-processing module (20), said input tensors including at least one of temporal and / or spatial and / or spectral information,- providing at least one conversion module (12), said conversion module (12) comprising an encoder and decoder network architecture (14),- training the encoder and decoder network architecture (14) by using partially masked samples of the input tensors to reconstruct a given representation of the input tensors,- wherein a masking ratio is applied with positional embedding, which provides simultaneous learning of local and global features present in the brain data, thereby transforming the input tensors into a latent space.
11. The method according to claim 10, characterized in that the method further comprises performing self-supervised learning (SSL).
12. The method according to claim 10 or claim 11 , characterized in that the method further comprises:- learning patterns of unlabeled neural activity, thereby at least partially uncovering the underlying dynamical structure of the brain (B), and- establishing a generic latent space that correlates with behavior in varying contexts.
13. The method according to any one of claims 10 to 12, characterized in that the step of converting the brain signals (S) into input tensors is implemented by using a wavelet transformation.
14. The method according to any one of claims 10 to 13, characterized in that the method further comprises:- compressing the brain signals (S) into a latent space, preferably wherein the method further comprises:- mapping the latent space to the original feature space, and- computing the reconstruction loss.
15. A computer-readable non-volatile storage medium comprising computer- readable instructions that, when executed by a processing means, causes said processing means to implement the method steps defined in claims 10 to 14.
Citation Information
Patent Citations
Method, device and equipment for training electroencephalogram signal analysis model and storage medium
CN116955983A