Rehabilitation evaluation method and system based on motor imagery and storage medium
By applying cross-modal alignment and optimal transmission theory, the mapping problem between EEG signals and natural language is solved, enabling precise and personalized rehabilitation assessment and supporting end-to-end rehabilitation training guidance.
Patent Information
- Application Number
- CN202510810973.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-11-14
AI Technical Summary
Existing motor intention analysis techniques lack deep mapping in EEG signal encoding, resulting in a lack of personalization and semantic diversity in the output natural language, as well as insufficient analysis accuracy.
By employing cross-modal alignment and optimal transmission theory, and through pre-trained encoders and decoders, we achieve accurate semantic alignment between EEG signals and natural language, generating structured rehabilitation assessment results.
It achieves deep mapping between EEG signals and natural language, generating accurate and personalized rehabilitation assessment results, and supports end-to-end multimodal assessment and personalized rehabilitation training guidance.
Smart Images

Figure CN120954692A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of characterization engineering technology, and in particular to a rehabilitation assessment method and system based on motor imagery. Background Technology
[0002] Electroencephalography (EEG) records brain electrical activity using electrodes and is an important non-invasive method for detecting neural activity. Motor function rehabilitation in subjects with neurological injuries is a significant clinical challenge, and accurately identifying a subject's motor intentions based on EEG signals is crucial for rehabilitation training.
[0003] In recent years, in response to the needs of multimodal modeling and rehabilitation applications of EEG signals, research has proposed a series of innovative solutions for language modeling, cross-modal fusion, and motor intention recognition. However, existing motor intention analysis technologies still face challenges in encoding EEG signals. For example, patent application WO2024258983A1, entitled "Humanaugmentation platform using context, biosignals, and language models," proposes a human ability enhancement system that integrates contextual information, biosignals, and user input, interacting with a generative AI model (GenAI) through a prompt composer. This system uses fixed encoding rules to transform EEG signals into a portion of a large model's prompt word for subsequent model input. However, this application encodes EEG signals through text token embedding, failing to demonstrate a deep mapping between EEG signals and natural language. How to construct a deep mapping between EEG signals and natural language is a bottleneck that urgently needs to be overcome in the field of neurorehabilitation.
[0004] Furthermore, at the application level, existing motor intention analysis methods, when applying EEG signals to identify motor intentions and assist in rehabilitation training, lack personalized and semantically diverse natural language output, and also suffer from analytical accuracy. How to generate accurate and personalized analysis results remains an unsolved problem. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a rehabilitation assessment method and system based on motor imagery, which can more accurately match EEG encoding with natural language in the feature space.
[0006] One aspect of the present invention provides a rehabilitation assessment method based on motor imagery, the method comprising the following steps: acquiring electroencephalogram (EEG) signal data collected by the subject during the performance of a rehabilitation task based on motor imagery, and extracting the frequency domain features of the EEG signal data; The frequency domain features of the EEG signal data are input into a pre-trained encoder for encoding, and the output is an embedded representation of the EEG signal data. The embedded representation of EEG signal data is input into a pre-trained natural language decoder for decoding, and the output is the rehabilitation assessment result of the subject. The rehabilitation assessment results are represented in natural language, and the encoder is trained in the following way: The labeled text contained in the training dataset is input into a pre-set text encoding module, and the embedded representation of the labeled text is output. The frequency domain features of the corresponding EEG signal data are input into a pre-set signal encoding module to obtain the embedded representation of the corresponding EEG signal data. The training dataset contains EEG signal data and corresponding labeled text. The EEG signal data in the training dataset consists of EEG signal samples collected from subjects during the performance of a rehabilitation task based on motor imagery. The labeled text is the motor imagery classification result of the corresponding EEG signal data, and the labeled text is represented in natural language. The encoding loss is calculated based on the embedded representation of the labeled text, the embedded representation of the corresponding EEG signal data, and a preset encoding loss function. This allows the signal encoding module and the text encoding module to be updated according to the encoding loss, achieving cross-modal alignment between the embedded representation of the labeled text and the embedded representation of the corresponding EEG signal data. In this way, a pre-trained encoder is obtained by iteratively updating the signal encoding module.
[0007] In some embodiments of the present invention, the method further includes: For each rehabilitation task based on motor imagery, the rehabilitation assessment results and acquired EEG signal data are input into the corresponding agent. The agent uses the acquired EEG signal data as an index to query the corresponding label text in the training dataset, and compares the queried label text with the rehabilitation assessment results to output the agent's analysis results. Specifically, if the queried label text and the rehabilitation assessment results are consistent, the agent's analysis results are the rehabilitation assessment results; if the queried label text and the rehabilitation assessment results are inconsistent, the agent's analysis results are obtained by integrating the queried label text and the rehabilitation assessment results.
[0008] In some embodiments of the present invention, the rehabilitation assessment results include information on assessment indicators corresponding to at least one rehabilitation task based on motor imagery; Rehabilitation tasks based on motor imagery include tasks for analyzing motor intention features and generating rehabilitation assessment language; and For the motor intention feature analysis task, the corresponding evaluation indicators include motor intention and cortical activation symmetry index; for the rehabilitation assessment language generation task, the corresponding evaluation indicators include alpha rhythm change trend, beta rhythm change trend, rehabilitation trend, attention state feedback, and training intervention strategies.
[0009] In some embodiments of the present invention, the coding loss function is jointly constructed based on the optimal transmission loss function and the information-noise contrast estimation loss function; The text encoding module includes a word segmentation unit, an embedding encoding unit, and a first perceptron unit, with the tag text used as input to the word segmentation unit and the first perceptron unit used as output to the embedded representation of the tag text. The signal encoding module includes a second perceptron unit and an attention unit built based on Transformer, and uses the output of the second perceptron unit as the input of the attention unit.
[0010] In some embodiments of the present invention, the natural language decoder is obtained by training a hybrid expert model through model distillation; the process of obtaining the natural language decoder by training a hybrid expert model through model distillation is as follows: The EEG signal data and corresponding label text contained in the training dataset are input into the teacher model, and the EEG signal data contained in the training dataset are input into the hybrid expert model to be trained. The hybrid expert model to be trained is updated based on the natural language text output by the teacher model, so that the text output by the hybrid expert model is semantically similar to the natural language text output by the teacher model. The natural language decoder is obtained through iterative updates. The semantic similarity between the text output by the hybrid expert model and the natural language text output by the teacher model is defined as the semantic similarity between the text output by the hybrid expert model and the natural language text output by the teacher model exceeds a set semantic similarity threshold.
[0011] In some embodiments of the present invention, the frequency domain features of the obtained electroencephalogram (EEG) signal data are extracted, including: A time window was used to divide the acquired EEG signal data into multiple signal data segments; The Fast Fourier Transform is used to perform spectral analysis on each signal data segment, and the frequency domain statistical characteristics are calculated based on the frequency domain analysis results and the set frequency domain statistical indicators.
[0012] In some embodiments of the present invention, the embedded representation of EEG signal data includes the number of EEG signal data, the number of signal data segments into which the EEG signal data is divided, and the coding dimension; the embedded representation of tag text includes the number of tag texts, the number of sentences in the tag text, and the coding dimension.
[0013] Another aspect of the present invention provides a rehabilitation assessment system based on motor imagery, including a processor, a memory, and a computer program / instructions stored in the memory. The processor is used to execute the computer program / instructions, and when the computer program / instructions are executed, the system implements the steps of the method described in any of the above embodiments.
[0014] Another aspect of the present invention provides a computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implement the steps of the method described in any of the above embodiments.
[0015] Another aspect of the present invention provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the method described in any of the above embodiments.
[0016] The rehabilitation assessment method and system based on motor imagery proposed in this invention can generate an embedded representation of EEG signal data through a specially designed encoder, and bypass the traditional text embedding stage by inputting the embedded representation of EEG signal data into a natural language decoder to generate structured rehabilitation assessment results expressed in natural language. This application can achieve end-to-end generation from raw neural electrical signals to multimodal assessment information, and demonstrates a deep mapping between EEG signals and natural language in the process.
[0017] Additional advantages, objects, and features of the invention will be set forth in part in the description which follows, and will also become apparent in part to those skilled in the art upon studying the description, or may be learned by practice of the invention. The objects and other advantages of the invention can be realized and obtained by means of the structures specifically pointed out in the description and drawings.
[0018] Those skilled in the art will understand that the objectives and advantages achievable with the present invention are not limited to those specifically described above, and that the above and other objectives achievable with the present invention will become clearer from the following detailed description. Attached Figure Description
[0019] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, are not intended to limit the scope of the invention. In the drawings: Figure 1 This is a flowchart illustrating a rehabilitation assessment method based on motor imagery in one embodiment of the present invention.
[0020] Figure 2 This is a schematic diagram of the process by which the training signal encoding module obtains the encoder in one embodiment of the present invention.
[0021] Figure 3 This is a schematic diagram of the embedded representation of input EEG signal data via soft coding in one embodiment of the present invention.
[0022] Figure 4 This is a flowchart illustrating a rehabilitation assessment method based on motor imagery in another embodiment of the present invention. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and descriptions of this invention are used to explain the invention, but are not intended to limit the invention.
[0024] It should also be noted that, in order to avoid obscuring the invention with unnecessary details, only the structures and / or processing steps closely related to the solution according to the invention are shown in the accompanying drawings, while other details that are not closely related to the invention are omitted.
[0025] It should be emphasized that the term "including / comprises" as used herein refers to the presence of a feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0026] It should also be noted that, unless otherwise specified, the term "connection" in this article can refer not only to a direct connection, but also to an indirect connection involving an intermediary.
[0027] In the following description, embodiments of the invention will be illustrated with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar parts, or the same or similar steps.
[0028] In the field of neurorehabilitation, although some technologies can achieve motor intention recognition and simple assessment, existing EEG signal encoding does not reflect the deep mapping between EEG signals and natural language, and the output rehabilitation assessment results lack personalization and semantic diversity.
[0029] To overcome the aforementioned challenges, this application creatively combines cross-modal alignment and optimal transport (OT) theory to achieve precise semantic alignment between EEG signals and natural language, and performs deep mapping between neurophysiological signals and clinical semantics, generating semantically diverse analysis results end-to-end. Specifically, cross-modal alignment in this application refers to mapping the embedded representations of different modalities (such as EEG signals and text) to the same semantic space, establishing semantic consistency relationships.
[0030] Figure 1 This is a flowchart illustrating a rehabilitation assessment method based on motor imagery in one embodiment of this application. Figure 1 As shown, the method includes steps S110 to S130.
[0031] Step S110: Obtain EEG signal data collected from the subject during the performance of a rehabilitation task based on motor imagery, and extract the frequency domain features of the EEG signal data.
[0032] Motor imagery refers to the active, conscious process in which participants imagine their own limbs or muscles (e.g., left and right hands and feet) moving without actual physical movement. It is crucial for motor function compensation and repair, requiring no physical movement or external stimulation. In the field of neurorehabilitation, specific brain regions are activated when participants perform rehabilitation tasks based on motor imagery, resulting in corresponding electroencephalogram (EEG) signals. Analyzing these EEG data allows for the detection and identification of activation effects in different brain regions to determine user intent.
[0033] In some embodiments of the present invention, electroencephalography (EEG) can be used to record the brain's own weak bioelectrical signals. Analysis of EEG signals typically requires the extraction of EEG features. This application extracts frequency domain features from the EEG signal data by: using a time window to divide the acquired EEG signal data into multiple signal data segments; performing spectral analysis on each signal data segment using Fast Fourier Transform (FFT) (extracting frequency features such as alpha and beta rhythms from the EEG signal data); and calculating frequency domain statistical features based on the frequency domain analysis results and set frequency domain statistical indicators. The time length contained in each signal data segment is the same. This application does not specifically limit the length of the time window; for example, the time window length can be 1 second or 0.5 seconds, and can be set according to requirements.
[0034] As an example, frequency domain statistical features are determined based on multiple statistical indicators of frequency domain features. These statistical indicators may include average spectral value, spectral variance, entropy, skewness, kurtosis, alpha rhythm energy proportion, beta rhythm energy proportion, and spectral width, among others (alpha and beta rhythms are typical frequency bands in EEG, widely used in studies related to motor intention and perception). The statistical indicators of frequency domain features can be selected according to the needs of EEG analysis to achieve the feature extraction process; this application does not impose specific limitations on them. The statistical indicators of each frequency domain feature can be calculated using the following formula: Average spectral amplitude Spectral amplitude variance Spectral entropy in, Spectral Energy Spectral skewness Spectral kurtosis Frequency-weighted standard deviation (spectral width) in, Alpha rhythm (8–13 Hz) power proportion β rhythm (13–30 Hz) power proportion Fourth-order frequency difference moment Where N is the total number of spectral points in the frequency domain analysis result obtained by FFT, k represents the index number of the frequency point (k = 1, 2, ..., N), F k p represents the spectral power value at the k-th frequency point. k f represents the normalized spectral energy at the k-th frequency point. k μ represents the frequency value (Hz) corresponding to the k-th frequency point. f This represents the frequency-weighted average.
[0035] Furthermore, the subject's EEG signal data can be an electroencephalogram (EEG) or related signal data obtained from the subject's EEG, and the EEG can record the subject's brain activity through electrodes.
[0036] Currently, there are various acquisition methods for EEG hardware devices. For example, Chinese invention patent application No. 202380050506.0, entitled "Wearable Device Capable of Long-Term Measurement of EEG and ECG," discloses a wearable device comprising: a sensor unit containing multiple electrodes mounted on the user's head above the neck to detect electrical signals from the head; different electrodes can monitor the user's EEG or ECG; and a control unit electrically connected to the sensor unit, which converts the analog signals of the EEG and ECG detected by the sensor unit into digital signals. However, traditional EEG acquisition devices have insufficient signal-to-noise ratio and spatiotemporal resolution (typically ≤1000Hz), making it difficult to capture fine neural activity related to motor imagery.
[0037] Considering the problems of unstable contact, poor wearing comfort, and significant signal attenuation in currently widely used dry electrodes and silver-silver chloride electrodes, especially in long-term rehabilitation training or home monitoring scenarios where artifacts and interruptions are prone to occur, this invention can choose to use high-performance, high-precision hydrogel electrodes to improve the quality of EEG signal data acquisition in terms of neurophysiological signal acquisition. The improved adhesion, conductivity, and stability of hydrogel electrodes can significantly enhance the anti-interference ability of EEG acquisition in dynamic rehabilitation scenarios, ensuring low-noise acquisition from the data source and guaranteeing a closed-loop advantage from the physical contact layer to downstream application models. For example, this application can use the hydrogel material disclosed in Chinese invention patent application No. 202410250422.8, entitled "A Thermally Reversible Hydrogel for Human Physiological Electrical Signal Acquisition and Its Preparation Method," to prepare the acquisition electrode. Its skin impedance is controlled to be below 10kΩ in the frequency range of 0.1Hz to 100,000Hz, reaching the gold standard level of wet electrodes, and the sampling frequency can reach up to 2000Hz.
[0038] As an example, hydrogel electrode materials can be adapted to various electrode structures, such as patch, coating, filling, and freestanding types. Furthermore, high-density electrode arrays (i.e., hydrogel sensor arrays) can be used to acquire high-quality EEG signals (i.e., electroencephalogram data). Flexible electrode arrays fabricated using hydrogel materials exhibit superior conductivity and wearing stability compared to traditional wet electrodes. Regarding electrode array layout, hydrogel electrodes support various flexible arrangements, including standard array electrodes, flexible mesh electrodes, and three-dimensional porous electrodes. Array electrodes are suitable for multi-point simultaneous monitoring of the cerebral cortex; mesh electrodes improve fit and breathability with their soft mesh structure, reducing discomfort during prolonged wear; three-dimensional porous electrodes utilize three-dimensional arrangement and microporous design to further enhance the contact area with the scalp, improving the comprehensiveness and quality of signal acquisition. In addition, this application can also add an additional control unit to control the electrode sensors, for example, using a Seeed Studio XIAO ESP32S3 chip as the control unit.
[0039] In some embodiments of the present invention, although hydrogel electrodes can solve the problem of impedance changes that traditional electrodes easily generate during long-term monitoring, resulting in high noise and low signal-to-noise ratio in the acquired signals, and achieve high-quality, long-term and high signal-to-noise ratio continuous signal acquisition in brain regions related to rehabilitation training such as motor control and motor activation, as well as dynamic monitoring and feedback during individualized rehabilitation training, significantly improving user comfort, biocompatibility and monitoring duration, the raw signal data packets directly acquired by the electrodes usually contain a lot of noise, which can easily interfere with delicate neural activities, so noise reduction processing is required for the raw acquired signal data.
[0040] As an example, the acquired EEG signal data is obtained by filtering the original acquired signal data. The specific steps of the filtering process are as follows: using a Notch filter to remove 50Hz power frequency interference, and employing a bandpass filter to retain signal data in the 1–40Hz frequency band, thereby achieving noise reduction of the original acquired signal data. The noise-reduced EEG signal data can effectively suppress artifact interference while preserving motor intention-related rhythmic information (such as alpha and beta rhythms), improving the accuracy and stability of subsequent feature extraction steps.
[0041] The above-described method of combining Notch filtering and bandpass filtering is merely an example. Wavelet transform can also be used to denoise and enhance the features of the original EEG signal (the process involves selecting a suitable wavelet basis (such as Daubechies or Symlet, etc.) for multi-scale decomposition and reconstructing the signal components within the required frequency band to extract alpha and beta rhythm features, ultimately obtaining EEG signal data). This invention is not limited to this.
[0042] Furthermore, when subjects engage in motor imagery, specific brain regions are activated. Collecting EEG signals from these regions yields diverse information that can be used to analyze the subject's motor imagery categorization or motor intentions during task execution. For instance, in neurorehabilitation procedures, to accurately reflect a subject's cognitive activity related to rehabilitation training and thus identify their motor intentions, signals reflecting brain regions involved in motor control and activation can be acquired. Specifically, EEG data from specific brain regions can be obtained to target rehabilitation tasks. These regions include the primary motor cortex, premotor area, supplementary motor area, parietal cortex, basal ganglia, and cerebellum.
[0043] Step S120: Input the frequency domain features of the EEG signal data into the pre-trained encoder for encoding, and output the embedded representation of the EEG signal data (EEG_embed).
[0044] In some embodiments of this invention, current research mostly employs contrastive learning or adversarial alignment to match labeled text tokens with EEG signals during encoder training. However, these methods rely on strongly supervised sample pairing, making it difficult to align distribution structures in small sample or unlabeled scenarios, and their convergence is unstable. Therefore, this application employs a cross-modal alignment method to train the encoder, such as... Figure 2 As shown, the specific training process is as follows: The labeled text in the training dataset is input into a pre-set text encoding module, which outputs the embedded representation of the labeled text. The frequency domain features of the corresponding EEG signal data (obtained by feature extraction from the EEG signal data in the training dataset) are input into a pre-set signal encoding module, which then outputs the embedded representation of the corresponding EEG signal data. Based on the embedded representation of the labeled text, the embedded representation of the corresponding EEG signal data, and a pre-set encoding loss function, an encoding loss is calculated. The text encoding module and the signal encoding module are updated according to the calculated encoding loss. This iterative update of the signal encoding module yields the encoder, enabling cross-modal alignment between the embedded representations of the labeled text and the corresponding EEG signal data. For example, this application can use a designed encoding loss function for gradient descent to iteratively update the parameters of the signal encoding module, resulting in a pre-trained encoder. In this application, the text encoding module and the signal encoding module can be considered as two independent encoding modules.
[0045] More specifically, the training dataset contains EEG signal data related to motor imagery and corresponding labeled text; therefore, the training dataset can also be called a motor imagery training dataset. In this application, the training dataset can be an existing dataset (e.g., EEG Motor Movement / Imagery Dataset v1.0.0) or a self-compiled dataset. The EEG signal data in the training dataset can also be EEG signal samples collected from subjects during a rehabilitation task based on motor imagery. The labeled text in the training dataset represents the motor imagery classification results of the corresponding EEG signal data, and the labeled text is expressed in natural language.
[0046] As an example, the embedding representation of labeled text can include the number of labeled texts, the number of sentences in the labeled text, and the encoding dimension, which can be represented as `train-diag_batch_embed.shape = [batch_size, number_of_sequence, dim]`. Here, `batch_size` represents the number of input labeled texts, `number_of_sequence` is the number of content sentences in the labeled text, and `dim` is the encoding dimension. Similarly, the embedding representation of EEG signal data can include the EEG signal data, the number of signal data segments into which the EEG signal data is divided, and the encoding dimension, which can be represented as `train-EEG_batch_embed.shape = [batch_size, number_of_segments, dim]`, where `batch_size` is the number of EEG signal lines, `number_of_segments` is the number of segments, and `dim` is the encoding dimension. Both the embedding representation of labeled text and the embedding representation of corresponding EEG signal data are sequences; therefore, they can also be referred to as the embedding representation sequence of labeled text and the embedding representation sequence of corresponding EEG signal data, respectively. In this application, the corresponding EEG signal data refers to the EEG signal data in the training dataset that corresponds to the labeled text.
[0047] In addition, the process of extracting features from the EEG signal data in the training dataset to obtain the frequency domain features of the EEG signal data in this application can be the same as the feature extraction process in step S110, and the statistical indicators of the frequency domain features can also be the same.
[0048] Furthermore, the text encoding module and the signal encoding module can map the labeled text and the corresponding EEG signal data to the same feature space for alignment, thereby generating the corresponding embedding representation. This mapping process can more fully extract the potential relationship between the labeled text and the corresponding EEG signal data, making the obtained embedding representation richer. The text encoding module includes a word segmentation unit, an embedding encoding unit, and a first perceptron unit in sequence, with the labeled text used as input to the word segmentation unit and the first perceptron unit used as output to the embedded representation of the labeled text; the signal encoding module includes a second perceptron unit and an attention unit built based on Transformer, with the output of the second perceptron unit used as the input of the attention unit.
[0049] The tokenization unit in the text encoding module is used to split the input labeled text to obtain the symbol representation sequence corresponding to the labeled text. For example, the tokenization unit can split "Hello, World" into [Hello, World]; the embedding encoding unit is used to convert the symbol representation sequence into a digital representation form; the first perceptron unit is used to extract the relationships between the individual digital representations in the digital representation form. That is, by inputting the labeled text in the training dataset into the tokenization unit and successively passing through the embedding encoding unit and the first perceptron unit, the embedding representation of the labeled text can be output. The attention unit in the signal encoding module has the ability of context understanding and is used to encode the frequency domain features of the input EEG signal data. For example, the attention unit can be a deep learning model constructed based on the self-attention mechanism; the second perceptron unit is then used to extract the potential relationships between the encoded features and output the embedding representation of the EEG signal data. That is, by inputting the EEG signal data in the training dataset into the attention unit and inputting the output of the attention unit into the second perceptron unit, the embedding representation of the EEG signal data can be output.
[0050] As an example, the first perceptron unit and the second perceptron unit can be constructed based on perceptrons. For example, the first perceptron unit in the text encoding module can be composed of 3 layers of perceptrons, and the second perceptron unit in the signal encoding module can be composed of 6 layers of perceptrons. In this application, the network structures of the first perceptron unit and the second perceptron unit can be set by themselves, and the number of perceptron layers mentioned above is only an example (and there is no proportional relationship between the number of perceptron layers of the first perceptron unit and the second perceptron unit), and the present invention is not limited thereto. In addition, the text encoding module can also directly use the encoding part of an existing large language model (LLM) (such as the tokenizer and parallel embedding layer of DeepSeek-V3) to tokenize and encode the labeled text in the motor imagery training dataset.
[0051] In some embodiments of this invention, during the process of obtaining the encoder from the training signal encoding module, this application introduces an optimal transmission loss function and designs an encoding loss function in combination with a conventional loss function to achieve cross-modal alignment between the embedded representation of the labeled text and the embedded representation of the corresponding EEG signal data in the feature space. Furthermore, when the encoding loss reaches a set threshold or the number of training iterations reaches a set number, the iterative update of the signal encoding module can be stopped, and the encoder is obtained. Taking a conventional loss function as an example of a point-to-point information noise contrast estimation loss function, this application constructs an encoding loss function based on a distribution-level optimal transmission loss function and an information noise contrast estimation loss function. The conventional loss function is used to constrain the accuracy of the embedded representation generated by the signal encoding module, while the OT loss function is used to align and match the distribution geometry of the EEG encoded sequence and the corresponding text embedded sequence, achieving structure-level modal alignment. That is, the optimal transmission loss function is used to align the distribution geometry of the embedded representation of the labeled text and the embedded representation of the corresponding EEG signal data.
[0052] As an example, the encoding loss function is based on the optimal transmission loss function. Estimating the loss function by comparing information noise with information noise This can be expressed by the formula: in, Where λ represents the weight of the optimal transmission loss, and E is the embedding representation e of all EEG signal data in the current batch (referring to all EEG signal data contained in the training dataset). i The set can be represented as E = {e1, e2, ..., e} i , ..., e N}(i∈M), where T is the embedding representation t of all tag texts in the current batch. j The set formed can be represented as T = {t1, t2, ..., t} j , ..., t N}(j∈M), π ij This is the optimal transfer matrix, used to describe the transfer from e i to t j Mass flow, π∈Π(μ) E μ T The embeddings represent sets E and T that follow probability distributions μ, respectively. E and μ T (The probability distribution in this application can be a uniform distribution), ε is the entropy regularization strength, and H(π) is the entropy regularization term of the transfer matrix, which can be expressed by the formula H(π)=-∑ i,j π ij logπ ij get.
[0053] and, Where τ represents the temperature coefficient, M represents the number of samples in each batch (i.e., batch size), and cos(e i , t j ) for e i and t j Normalized cosine similarity, The same applies to the others.
[0054] Apart from Common loss functions include cross-entropy loss and Dice loss. The cross-modal alignment scheme designed in this application based on the optimal transmission mechanism can measure the global distribution difference between the embedding representations of two modalities without requiring strict pairing information. It is particularly suitable for modal fusion in high-dimensional, weakly labeled medical scenarios, and combines theoretical interpretability with engineering feasibility.
[0055] After training the encoder, step S120 can be executed, whereby the frequency domain features of the EEG signal data (if multiple frequency domain feature indicators are selected, the frequency domain features of the EEG signal data can be used as input in sequence form) are input into the pre-trained encoder (i.e., the trained signal encoding module), and the embedded representation of the EEG signal data can be output. The embedded representation of the EEG signal data includes the number of EEG signal data, the number of signal data segments, and the encoding dimension, which can be represented as EEG_batch_embed.shape = [batch_size, number_of_segments, dim], where batch_size is the number of EEG signal data, number_of_segments is the number of segments, and dim is the encoding dimension.
[0056] Step S130: The embedded representation of the EEG signal data is input into a pre-trained natural language decoder for decoding, and the rehabilitation assessment result of the subject is output. The rehabilitation assessment result is represented in natural language.
[0057] More specifically, the Mixture of Experts (MoE) architecture is a deep model structure that can divide different tasks through multiple "expert sub-models," which helps to improve model capacity and generalization performance. To adapt to different user groups, this application iteratively trains the MoE model (the model to be trained based on the MoE mechanism) to obtain a natural language decoder with strong generalization decoding capabilities for motion intent.
[0058] In some embodiments of the present invention, the natural language decoder can be obtained by training a hybrid expert model through model distillation. Specifically, the process of obtaining the natural language decoder by training a hybrid expert model through model distillation is as follows: The EEG signal data and labeled text from the motor imagery training dataset are input into the teacher model, while the EEG signal data from the training dataset are input into the hybrid expert model to be trained. The hybrid expert model is updated based on the natural language text output by the teacher model, ensuring that the text output by the hybrid expert model is semantically similar to the natural language text output by the teacher model. This process is repeated iteratively to obtain the natural language decoder. Specifically, if the semantic similarity between the text output by the hybrid expert model and the natural language text output by the teacher model exceeds a set semantic similarity threshold, the text output by the hybrid expert model is considered semantically similar to the natural language text output by the teacher model.
[0059] The above-mentioned method of using model distillation to train the MoE model to obtain the natural language decoder is only an example. Other training methods can also be used (but you need to build your own dataset to train the MoE model). This application does not specifically limit them.
[0060] As an example, considering that the natural language decoding performance of the teacher model is superior to that of the hybrid expert model, the teacher model can be a large language model with strong contextual understanding and natural language generation capabilities, such as the DeepSeek-R1 model. In addition to using LLM for model distillation to obtain the natural language decoder, the MoE part of an existing LLM can also be directly applied as the natural language decoder. Furthermore, since the embedding representations of labeled text and corresponding EEG signal data in the same feature space are aligned across modalities, the text encoding module and signal encoding module in this application can be trained together; this application iteratively updates the signal encoding module and the MoE model by adjusting their parameters. The encoder and natural language decoder can be trained separately or as a whole.
[0061] Furthermore, existing technologies often employ hard encoding to input the embedded representation of EEG signal data into a natural language decoder. This means the encoded EEG signal data is treated as natural language and directly used by the natural language decoder for word segmentation and embedding. Therefore, hard encoding results in significant EEG information loss. This invention considers a soft embedding injection mechanism (which embeds continuous feature vectors into the language model input layer). By using this embedded representation, the embedded representation of the EEG signal data is input into the natural language decoder. This not only significantly simplifies engineering implementation but also allows for compatibility with different types of language models (such as DeepSeek, ChatGLM, or BERT) to build the natural language decoder, thus possessing broader versatility. Additionally, the rehabilitation assessment results output by the natural language decoder can be presented in text or chart format.
[0062] As an example, considering the good transferability of soft coding mechanisms, they are suitable for semantic understanding tasks in the field of brain signals, such as... Figure 3 As shown, this application uses the MoE part of Deepseek-V3 (e.g., Figure 3 The DeepSeek core computation module in the model is used as the natural language decoder, and the input of the MoE part in the DeepSeek model is adapted and modified. This allows the embedded representation of the EEG signal data to be directly embedded into the input layer of the natural language decoder through soft coding. In other words, this application skips the traditional text token embedding path and maintains the core MoE mechanism of the large language model to execute step S130. Furthermore, if the encoder uses the encoding part of an existing large language model and the natural language decoder uses the MoE part of an existing large language model, then the encoder and natural language decoder can be collectively referred to as the Electroencephalogram Rehabilitation Large Model (EEG-RehabLM).
[0063] In some embodiments of the present invention, the rehabilitation assessment results output by the natural language decoder can be applied to downstream applications to guide clinical monitoring and postoperative rehabilitation tasks for subjects. The rehabilitation assessment results obtained in step S130 can be used to guide motor imagery-based rehabilitation tasks, including motor intention feature analysis tasks and rehabilitation assessment language generation tasks. The motor intention feature analysis task is used to identify and analyze motor intentions, while the rehabilitation assessment language generation task is to generate rehabilitation suggestions and content in natural language form for clinicians and subjects. Furthermore, the rehabilitation assessment results can include information on different types of assessment indicators for different types of motor imagery-based rehabilitation tasks. For example, if the type of motor imagery-based rehabilitation task for which this application is applied is not specified before the rehabilitation assessment (e.g., using EEG signal data from the whole brain region without specifying a specific brain region), the rehabilitation assessment results can output information on assessment indicators corresponding to each motor imagery-based rehabilitation task; if the type of motor imagery-based rehabilitation task for which this application is applied is specified before the rehabilitation assessment, the rehabilitation assessment results can output information on assessment indicators corresponding to that type of motor imagery-based rehabilitation task. That is, the rehabilitation assessment results include information on assessment indicators corresponding to at least one rehabilitation task based on motor imagery.
[0064] As an example, for a motor intention feature analysis task, the corresponding rehabilitation assessment indicators may include motor intention and the cortical activation symmetry index; for a rehabilitation assessment language generation task, the corresponding rehabilitation assessment indicators may include the alpha rhythm (8–13 Hz) change trend, the beta rhythm (13–30 Hz) change trend, the subject's rehabilitation trend, the subject's attentional state feedback, and the subject's rehabilitation training intervention strategies. Among these, the cortical activation symmetry index is used to assess whether the activation of the subject's left and right hemiplegic motor areas is balanced; for example, this index can be used to assess the recovery status of hemiplegia. The change trend of alpha or beta rhythm can indicate the recovery status of neural activation in brain regions. The subject's rehabilitation trend is an indicator used to quantify the effectiveness of rehabilitation training, and can be graded using a scoring system. Attentional state feedback reflects the subject's cooperation and the stability of the training effect. Training intervention strategies may include active training, alternating training, or passive guidance, among other intervention strategy suggestions.
[0065] Furthermore, the rehabilitation assessment results output by the natural language decoder are only preliminary assessments, and there are still significant challenges in generating guidance reports based on them. For example, in rehabilitation assessment language generation tasks, after obtaining the rehabilitation assessment results, expert rules, template completion, or scoring table derivation are often used to construct the rehabilitation assessment results. Although this facilitates structured management, it cannot dynamically generate diverse and personalized language outputs, and it lacks the ability to express non-standard indicators (such as the subject's subjective motor performance, cognitive participation, etc.). Therefore, to overcome the above challenges and achieve a deep mapping between neurophysiological signals and clinical semantics, this invention further employs an agent to perform language mapping on structured rehabilitation information, generating rehabilitation suggestions, trend assessments, and personalized training prompts in natural language format. This process can support multi-turn interactions and interpretable output, and has greater value for human-computer interaction and clinical application.
[0066] In some embodiments of the present invention, such as Figure 4 As shown, the intelligent analysis method proposed in this application further includes step S140: for each rehabilitation task based on motor imagery, the rehabilitation assessment results and the acquired EEG signal data are input into the intelligent agent corresponding to the task. The intelligent agent uses the acquired EEG signal data as an index to query the tag text corresponding to the EEG signal data in the motor imagery training dataset, and compares the queried tag text with the rehabilitation assessment results, thereby outputting the intelligent agent's analysis results. Specifically, this application can set up a corresponding intelligent agent for each type of rehabilitation task based on motor imagery to generate personalized instructions based on the corresponding analysis information, or it can use a common intelligent agent to output corresponding personalized prompts based on the rehabilitation assessment results of each rehabilitation task based on motor imagery.
[0067] More specifically, the agent can treat the motor imagery training dataset as a knowledge base and use EEG signal data as an index to query the tag text corresponding to the EEG signal data in the motor imagery training dataset. By comparing the queried tag text with the rehabilitation assessment results, it can identify the analytical information missing from the rehabilitation assessment results (i.e., identify the rehabilitation assessment information corresponding to the ignored part of the EEG signal data). The agent can then compile the tag text content corresponding to the missing analytical information and the rehabilitation assessment results to output personalized agent analysis results for the subject. For example, if a short-duration spur signal exists after a stable EEG signal, the encoder and natural language decoder might treat this spur signal as noise and ignore it during analysis. After querying the original EEG signal data, the agent can supplement the analytical information corresponding to the spur signal based on the queried tag text (if the agent determines that the rehabilitation assessment results lack analytical information for this spur signal, but the dataset does not contain the tag text for this spur signal, the agent only needs to point out the ignored part of the EEG signal data).
[0068] That is, when the retrieved tag text matches the rehabilitation assessment results, the agent analysis results are obtained by organizing the rehabilitation assessment results; when the retrieved tag text and rehabilitation assessment results do not match, the agent analysis results are obtained by integrating the retrieved tag text and rehabilitation assessment results. For example, the preliminary rehabilitation assessment results output by the natural language decoder may include the following information: basic information of the subject (including name, gender, age, height, and weight, etc.), the type of rehabilitation task based on motor imagery, and the corresponding indicator information of the rehabilitation task based on motor imagery; the personalized rehabilitation assessment results output by the agent may include the following text information: basic information of the subject, the type of rehabilitation task based on motor imagery, the corresponding indicator information of the rehabilitation task based on motor imagery, and the analysis information of the ignored EEG signal data portion. In addition, the preliminary rehabilitation assessment results and personalized rehabilitation assessment results may also include the subject's EEG signal data (such as visualized EEG).
[0069] As an example, the agent can utilize LangChain to query the knowledge base for the tag text corresponding to the retrieved EEG signal data. Furthermore, considering that EEG signal data typically involves long time periods, making it difficult to find completely identical EEG signals, it can also use individual signal data segments as indexes to query the motor imagery training dataset for the tag text corresponding to each segment. Additionally, the agent's output analysis results can be presented in a multilingual and interpretable format, making it suitable for applications such as medical institutions, rehabilitation centers, and home assessments.
[0070] The rehabilitation assessment method based on motor imagery proposed in this application has the following significant advantages: ① Regarding cross-modal alignment, this application utilizes a joint encoding loss function constructed from conventional loss and optimal transmission loss to more accurately match the embedded representation of the labeled text with the corresponding embedded representation of the EEG signal data in the feature space, thereby improving the consistency and accuracy of the encoder in semantic generation. The encoder proposed in this invention possesses significant uniqueness.
[0071] ②This application realizes the end-to-end generation of raw neural electrical signals into multimodal rehabilitation assessment information.
[0072] ③ This invention uses an agent to organize rehabilitation assessment results based on different motor imagery-based rehabilitation tasks, thereby achieving customized and structured output and improving the readability and intuitiveness of the analysis results. Specifically, the agent, based on the preliminary rehabilitation assessment results output by the natural language decoder and combined with EEG signals, generates corresponding analytical information according to the corresponding motor imagery-based rehabilitation task, providing a basis for doctors and subjects to read the information. Simultaneously, if the subject's EEG signals show abnormalities beyond the specified indicators, the agent can also provide explanations and analysis.
[0073] ④ Compared with traditional dry electrodes and conductive gel wet electrodes, this application uses electrodes made of hydrogel material to acquire EEG signal data. The hydrogel electrode used in this application (1) has strong adhesion and flexibility, and has high skin compatibility (continuous wearing time exceeds 14 hours), which is particularly suitable for long-term neural signal acquisition in dynamic motion scenarios; (2) the effective working time is at least 24 hours, which is 6 to 8 times longer than the traditional gold standard wet electrode (conductive gel), and can realize high-frequency EEG monitoring and continuous acquisition; (3) the material is synthesized by green and controllable condensation reaction, which has high stability; (4) the impedance adjustment range is wide, usually covering 100Ω to 100,000Ω (100 to 100,000Ω). 5 (5) It exhibits highly stable frequency coding in EEG evoked experiments and is suitable for motor intention analysis.
[0074] Corresponding to the above method, the present invention also provides a rehabilitation assessment system based on EEG motor imagery. The system includes a computer device, which includes a processor and a memory. The memory stores computer programs / instructions, and the processor is used to execute the computer programs / instructions stored in the memory. When the computer programs / instructions are executed by the processor, the system performs the steps of the method described above.
[0075] This invention also provides a computer-readable storage medium storing a computer program / instructions thereon, which, when executed by a processor, implements the steps of the aforementioned edge computing server deployment method. The computer-readable storage medium can be a tangible storage medium, such as random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, floppy disks, hard disks, removable storage disks, CD-ROMs, or any other form of storage medium known in the art.
[0076] This invention also provides a computer program product storing a computer program / instructions thereon, which, when executed by a processor, implements the steps of the aforementioned edge computing server deployment method. The computer program product can be a tangible product, such as random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, floppy disks, hard disks, removable storage disks, CD-ROMs, or any other form of product known in the art.
[0077] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this invention are programs or code segments used to perform the desired tasks. The programs or code segments can be stored in a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried in a carrier wave.
[0078] It should be clarified that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of the present invention.
[0079] In this invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.
[0080] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations of the embodiments of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A rehabilitation assessment method based on motor imagery, characterized in that, The method includes the following steps: EEG signal data collected from subjects during a rehabilitation task based on motor imagery was obtained, and the frequency domain features of the EEG signal data were extracted. The frequency domain features of the EEG signal data are input into a pre-trained encoder for encoding, and the output is an embedded representation of the EEG signal data. The embedded representation of the EEG signal data is input into a pre-trained natural language decoder for decoding, and the output is the subject's rehabilitation assessment result. The rehabilitation assessment results are represented in natural language, and the encoder is trained as follows: The labeled text in the training dataset is input into a pre-set text encoding module, which outputs the embedded representation of the labeled text. The frequency domain features of the corresponding EEG signal data are input into a pre-set signal encoding module, which outputs the embedded representation of the corresponding EEG signal data. The training dataset contains EEG signal data and corresponding labeled text. The EEG signal data in the training dataset consists of EEG signal samples collected from subjects during a rehabilitation task based on motor imagery. The labeled text is the motor imagery classification result of the corresponding EEG signal data, and the labeled text is represented in natural language. An encoding loss is calculated based on the embedded representation of the labeled text, the embedded representation of the corresponding EEG signal data, and a pre-set encoding loss function. This allows the signal encoding module and the text encoding module to be updated according to the encoding loss, achieving cross-modal alignment between the embedded representation of the labeled text and the embedded representation of the corresponding EEG signal data. The pre-trained encoder is then obtained by iteratively updating the signal encoding module.
2. The method according to claim 1, characterized in that, The method further includes: For each rehabilitation task based on motor imagery, the rehabilitation assessment results and the acquired EEG signal data are input into the agent corresponding to the task. The agent uses the acquired EEG signal data as an index to query the tag text corresponding to the EEG signal data in the training dataset, and compares the queried tag text with the rehabilitation assessment results to output the agent analysis results. Wherein, if the retrieved tag text and the rehabilitation assessment result are consistent, the agent analysis result is the rehabilitation assessment result; if the retrieved tag text and the rehabilitation assessment result are inconsistent, the agent analysis result is obtained by integrating the retrieved tag text and the rehabilitation assessment result.
3. The method according to claim 2, characterized in that, The rehabilitation assessment results include information on assessment indicators corresponding to at least one rehabilitation task based on motor imagery; The rehabilitation task based on motor imagery includes a motor intention feature analysis task and a rehabilitation assessment language generation task. as well as For the motor intention feature analysis task, the corresponding evaluation indicators include motor intention and cortical activation symmetry index; for the rehabilitation assessment language generation task, the corresponding evaluation indicators include alpha rhythm change trend, beta rhythm change trend, rehabilitation trend, attention state feedback, and training intervention strategies.
4. The method according to claim 1, characterized in that, The encoding loss function is jointly constructed based on the optimal transmission loss function and the information-noise contrast estimation loss function; The text encoding module includes a word segmentation unit, an embedding encoding unit, and a first perceptron unit in sequence, and the tag text is used as input to the word segmentation unit, and the first perceptron unit is used as output to the embedded representation of the tag text; The signal encoding module includes a second perceptron unit and an attention unit built based on Transformer, and uses the output of the second perceptron unit as the input of the attention unit.
5. The method according to claim 1, characterized in that, The natural language decoder was obtained by training a hybrid expert model through model distillation. The process of obtaining the natural language decoder by training the hybrid expert model through model distillation is as follows: The EEG signal data and corresponding label text contained in the training dataset are input into the teacher model, and the EEG signal data contained in the training dataset are input into the hybrid expert model to be trained. The hybrid expert model to be trained is updated based on the natural language text output by the teacher model, so that the text output by the hybrid expert model is semantically similar to the natural language text output by the teacher model, thereby obtaining the natural language decoder through iterative update; Specifically, the semantic similarity between the text output by the hybrid expert model and the natural language text output by the teacher model is defined as the semantic similarity between the text output by the hybrid expert model and the natural language text output by the teacher model exceeding a set semantic similarity threshold.
6. The method according to claim 1, characterized in that, The extraction of frequency domain features from the EEG signal data includes: A time window was used to divide the acquired EEG signal data into multiple signal data segments; The Fast Fourier Transform is used to perform spectral analysis on each signal data segment, and the frequency domain statistical characteristics are calculated based on the frequency domain analysis results and the set frequency domain statistical indicators.
7. The method according to claim 6, characterized in that, The embedding representation of the EEG signal data includes the number of EEG signal data, the number of signal data segments into which the EEG signal data is divided, and the coding dimension; the embedding representation of the tag text includes the number of tag texts, the number of sentences in the tag text, and the coding dimension.
8. A rehabilitation assessment system based on motor imagery, comprising a processor, a memory, and a computer program / instructions stored in the memory, characterized in that, The processor is configured to execute the computer program / instructions, and when the computer program / instructions are executed, the system implements the steps of the method as described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1 to 7.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Thermally reversible hydrogel for collecting human physiological electrical signals and preparation method of thermally reversible hydrogel
CN118307802A
Wearable device capable of measuring EEG and ECG for long time
CN119451626A
Human augmentation platform using context, biosignals, and language models
WO2024258983A1
Rehabilitation training system and method based on motor imagery-brain-computer interface and virtual reality
CN113398422A
Electroencephalogram signal recognition method based on language imagination and motor imagination time sequence coding
CN115590532A
Cited By
Rehabilitation evaluation system and implantable neural rehabilitation equipment
CN121641463A
EEG signal feature extraction method, system and equipment based on frequency band-to-attention mapping mechanism
CN121667720A