Decoding method before upper limb movement
By extracting data augmentation and electrophysiological source imaging preprocessing in the pre-motion period, combined with recursive graphs and attention residual networks, the classification performance problem in MRCP signal multi-classification decoding is solved, and five high-precision upper limb motion types are realized.
Patent Information
- Application Number
- CN202510283747.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art has the problem of insufficient classification performance in the multi-classification decoding of motor-related cortical potential (MRCP) signals, especially in the case of low frequency characteristics and low signal-to-noise ratio, it is difficult to achieve high-precision decoding of five upper limb motion types.
A pre-motion decoding method of upper limbs is adopted, including extracting the pre-motion period of interest, performing data augmentation and electrophysiological source imaging (ESI) pre-processing, selecting source channels and performing dimensionality reduction, calculating recursive maps (RPs), and using Attention-ResNet for classification.
It realizes the decoding of five upper limb movement types with high accuracy using only pre-motion data, with a classification accuracy of more than 90%, which significantly improves the separability and classification performance of signal characteristics.
Smart Images

Figure CN120387098A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of movement decoding, and particularly to a method for pre-decoding upper limb movements. Background Art
[0002] Movement-related cortical potential (MRCP) is a negative brain potential generated and observed in electroencephalogram (EEG) recordings, and is usually related to voluntary movements. As a standard paradigm in EEG-based brain-computer interface (BCI) systems, MRCP is often used as a control signal to operate the subordinate units of BCI systems. When the start of movement is set at time 0 seconds, MRCP consists of three components: the readiness potential (-2 seconds to -0.1 seconds), the movement potential (-0.1 seconds to 0 seconds), and the movement monitoring potential (0 seconds to 1 second), which correspond to the preparation and planning of movement, the execution of movement, and the evaluation and monitoring of the executed movement, respectively. Given the strong correlation between the generation of MRCP and the activation of movement intentions in various movement-related functional areas of the cerebral cortex, MRCP is not only observed during movement execution (ME), but also during movement imagination (MI) and movement attempt (MA), especially in patients with movement disorders such as stroke, cerebral palsy, locked-in syndrome, and amyotrophic lateral sclerosis. These characteristics make MRCP a favorable control signal for BCI systems because they can activate movement-related brain regions and have significant potential in establishing precise temporal correlations between motor cortex activation and related somatosensory feedback afferents. The simultaneous realization of these two processes is crucial for neuroplastic rehabilitation based on Hebbian theory.
[0003] MRCP is widely used in brain-computer interface (BCI) to decode movement intentions. However, most studies have focused on classifying two types of movements or studying multi-classification problems by constructing multiple binary classifiers. Although these studies have verified the feasibility of decoding movement intentions through MRCP and made progress in performance, from the perspective of rehabilitation needs and BCI system applications, the binary classification method obviously has limitations. When the goal is to achieve more natural control, multi-classification using MRCP remains a considerable challenge. Many studies have investigated multi-classification decoding by designing paradigms such as movement speed, force magnitude, limb direction, and sequential instructions. However, overly detailed paradigm design may limit the generalization of applications and result in an unnatural user experience. Therefore, algorithm-based methods may represent a more promising solution.
[0004] Research on multi-class decoding algorithms for movement-related cortical potential (MRCP) signals is limited. Mohseni et al. used wavelet common spatial patterns to decode five complex upper limb movements from premotor electroencephalogram (EEG) signals, achieving an accuracy of 0.70 on a subset of the most efficient channels. Tao et al. applied multivariate empirical mode decomposition and a convolutional neural network (CNN) to decode EEG signals for four hand movements, achieving an average classification accuracy of 0.8124 across 13 subjects using data from the first 2 seconds of movement onset. Jia et al. proposed a TTSNet that combined filter bank task-related component analysis with a CNN, achieving an accuracy of 0.4588±0.0724 on a multi-classification task. In summary, the classification performance of current methods still needs to be improved. This may be due to the simple temporal features resulting from the low-frequency nature of MRCP signals and the reduced signal-to-noise ratio (SNR) and spatial resolution caused by volume conduction effects. Consequently, the discriminability of the extracted features remains insufficient. Summary of the invention
[0005] In order to solve the problems existing in the prior art, the purpose of the present invention is to provide an upper limb pre-movement decoding method, which can achieve high-precision decoding of five upper limb movement types using only pre-movement data.
[0006] To achieve the above object, the present invention adopts a technical solution: a method for decoding upper limb movement before movement, comprising the following steps:
[0007] Step 1: Extract the pre-exercise period of interest and perform data augmentation to construct a new dataset;
[0008] Step 2: Preprocessing: Electrophysiological source imaging (ESI) is directly applied to the data, followed by source channel selection and downsampling to obtain the source signal.
[0009] Step 3, feature extraction: calculate the recurrence graph RP of each channel of the source signal and obtain the average recurrence graph;
[0010] Step 4: Take the average recurrence graph as input, train the attention residual network Attention-ResNet and perform classification.
[0011] As a further improvement of the present invention, in step 2, applying electrophysiological source imaging (ESI) to the data includes forward modeling and inverse problem solving. The forward modeling projects the neuronal source activity in the brain onto the scalp electrodes, while taking into account the geometric shapes and conductive properties of various tissues during signal propagation. The projection process is as shown in formula (1):
[0012] x(t)=Gq r (t) (1)
[0013] Where r = 1,...,Ns , is the vector of the three-dimensional directional primary current related to position 3×N s , with a magnitude of q r (t), where N s represents the number of possible source positions in the cortex, and G is the lead field matrix, i.e., the head model, which maps the primary current q r (t) to the scalp signal x(t), with a magnitude of N c ×3N s , where N c is the number of scalp electrodes;
[0014] The inverse problem is solved by backtracking and estimating the source q r (t) corresponding to the acquired scalp signal x(t) and the lead field matrix G, using a linearly constrained minimum variance beamformer, and estimating the source activity at each source position through spatial filtering. According to equation (1), the output of the beamformer is expressed as follows:
[0015]
[0016] Given the source signal position r, the spatial filter W T satisfies the following constraint conditions:
[0017] W T (r i )*G(r i ) = 1 (3)
[0018] W T (r i )*G(r j ) = 0 (4)
[0019] For each case where i is not equal to j, 1 and 0 are 3x3 matrices, representing the identity matrix and the zero matrix respectively; assume i is the position of the current source; j is regarded as an interfering source, and its influence is minimized; is a good approximation of q r when the filter output is minimized, when is minimized:
[0020]
[0021] The good approximation q obtained through LCMV beamforming estimation r (t) is used as the source signal in the cerebral cortex.
[0022] As a further improvement of the present invention, in step 2, the source channel selection and downsampling are specifically as follows:
[0023] Source channel selection includes region of interest (ROI) selection and dimensionality reduction, and determines the positions of all sources and source signals through a head model and an inverse solution method; subsequently, ROIs are extracted according to the cell structure or tissue structure of the cerebral cortex and the regions defined by the cellular tissue;
[0024] Dimensionality reduction of the source signals of the ROI is performed through multi-channel signal correlation analysis. During the dimensionality reduction process, first, the correlation coefficient r between channel signals is calculated pairwise, and a correlation coefficient matrix is generated; then, the following operations are applied to all values in the matrix: (i) taking the absolute value; (ii) setting the values of |r| greater than or equal to a preset threshold to 1, and setting the remaining values to 0; (iii) summing the values of each row or column, arranging them in descending order, and selecting the top n channels as the target channels for signal extraction; finally, the signal is downsampled.
[0025] As a further improvement of the present invention, in step 3, RP is used for feature extraction, the time series is embedded into the phase space, and the distance between trajectory states is calculated; when the distance exceeds the threshold, a black dot is plotted in RP; otherwise, a white dot is plotted; for a given time series length x = x i i = 1,..., N, whose phase space vector is represented as And the distance between the phase space trajectory states is given by the following equation:
[0026]
[0027] where R is the recurrence matrix, ||·|| is the norm, ε is the threshold distance, and Θ(·) is the Heaviside function; the values of R are 0 and 1, corresponding to the white and black dots in RP, and the distribution and geometry of the dots reflect the recurrence characteristics of the time series.
[0028] As a further improvement of the present invention, in step 4, the data first sequentially performs two-dimensional convolution, regularization, activation function, and max pooling calculations; subsequently, the data passes through four large residual modules, each module contains two small residual blocks, and each block performs two convolutions, each convolution using a 3×3 kernel and a padding of 1; after average pooling, the final feature vector is passed to the fully connected FC layer to obtain the feature vector, and the feature vector is then input to the Softmax layer for final calculation, and its output represents the probability distribution of each category in the classification task;
[0029] As a further improvement of the present invention, the attention residual network Attention-ResNet includes a channel attention module, a spatial attention module, and a convolutional block attention module inserted after the final feature vector, between the small residual blocks, and inside the small residual blocks.
[0030] The beneficial effects of the present invention are:
[0031] The present invention proposes a multi-class decoding framework that combines ESI, RP, and an attention residual neural network for classifying five motion types including four upper limb movements and rest, while considering the characteristics of MRCP signals; specifically reflected in:
[0032] (1) A new decoding framework for the MRCP signal paradigm is proposed, which can achieve high-precision decoding of five upper limb movement types using only pre-movement data;
[0033] (2) To address challenges such as low signal-to-noise ratio, low spatial resolution, and strong temporal correlation and low-frequency characteristics in MRCP, the ESI and RP technologies are combined, significantly improving the separability of signal features;
[0034] (3) ResNet is optimized by integrating an attention mechanism to adapt to the obtained two-dimensional features, and the influence of the attention mechanism and the optimization mechanism on the classification performance is further verified, ultimately achieving a five-class classification result with an accuracy exceeding 90%. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a flowchart of the multi-class pre-movement decoding framework in an embodiment of the present invention;
[0036] Figure 2 It is a schematic diagram of the source channel selection process in an embodiment of the present invention;
[0037] Figure 3 It is a schematic diagram of the feature extraction process in an embodiment of the present invention;
[0038] Figure 4 It is a schematic diagram of the ResNet model and the optimization mechanism in an embodiment of the present invention;
[0039] Figure 5 It is a schematic diagram of the confusion matrix results of 10 models in an embodiment of the present invention;
[0040] Figure 6 It is the overall average Bereitschaftspotential map in an embodiment of the present invention;
[0041] Figure 7 It is a schematic diagram of the visualization of high-dimensional features based on UMAP in an embodiment of the present invention;
[0042] Figure 8 It is the five-class class activation map (CA-map) based on ResNet-c-CAM in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0044] Embodiment
[0045] As Figure 1 shown, for a pre - movement decoding method of the upper limb, first, extract the pre - movement period of interest and perform data augmentation to construct a new dataset. In the subsequent pre - processing process, directly apply ESI, region of interest (ROI) selection, dimensionality reduction, and down - sampling to the data. In the feature extraction process, calculate the recurrence plot (RP) for each channel of the source signal and obtain the average recurrence plot. Finally, use the average recurrence plot as the input to train an attention residual network (Attention - ResNet) for classification; again, as Figure 1 shown, the first module shows the process of extracting the target data segment for five different types of movements; after data augmentation, the second module shows the ESI process, where the head model projects the scalp signal to obtain the cortical source signal; subsequently, select the source signal position (including ROI selection, dimensionality reduction, and down - sampling), and calculate the recurrence plot (RP) for each channel; finally, input the average RP into the attention residual network (Attention - ResNet) for classification.
[0046] The following further explains this embodiment:
[0047] A. Pre - processing:
[0048] Classify the operations before calculating the recurrence plot (RP) as pre - processing steps. In the current analysis, pre - processing does not involve any conventional operations on electroencephalogram (EEG) data, but focuses on ESI, followed by source channel selection (including region of interest (ROI) selection and dimensionality reduction) and down - sampling.
[0049] 1) ESI:
[0050] The core of the pre - processing process is ESI, which is a neuroimaging technique that determines the source of the recorded scalp potential by integrating the temporal and spatial components of the electroencephalogram (EEG), thus facilitating the study of signals at the source level. Currently, it is widely regarded as an effective solution to the low spatial resolution of EEG. The implementation of ESI involves two basic steps: forward modeling and inverse problem solving. Forward modeling describes how the neuronal source activity in the brain projects onto the scalp electrodes, taking into account the geometry and electrical conductivity of various tissues (such as gray matter, white matter, cerebrospinal fluid, skull, and skin) during the signal propagation process. This projection process can be expressed by equation (1):
[0051] x(t) = Gq r (t) (1)
[0052] where, (r(r = 1,...,N s )) is the vector of the three - dimensional directional primary current related to the position (3×N s ) with a magnitude of (qr (t)), where Ns represents the number of possible source locations in the cortex. (G) is called the "lead field" matrix (head model), which maps the primary current (q r (t)) to the scalp potential (x(t)), and its size is (N c ×3N s ), where (N c ) is the number of scalp electrodes. The head model used in this embodiment is a high-precision cortical surface grid finite element model based on the anatomy of Haufe et al. according to ICBM152s, also known as the New York Head (NYH) model. This model is very detailed, including six tissue types: scalp, skull, cerebrospinal fluid, gray matter, white matter, and cavities, with a total of 2004 source locations.
[0053] The inverse problem is solved by backtracking and estimating the source (q r (t)) corresponding to the acquired scalp signal (x(t)) and the lead field matrix (G).
[0054] However, solving the inverse problem faces significant challenges because the number of dipoles to be estimated far exceeds the number of available sensors, and there are infinitely many solutions. To address this issue, the most widely adopted methods include minimum norm estimation (MNE), beamformers, and dipole modeling. This framework uses the linearly constrained minimum variance (LCMV) beamformer, which provides a compromise between dipole modeling and MNE, allowing the estimation of source activity at each source location through spatial filtering and having the advantage of not requiring a predefined number of dipoles. Grosse Wentrup et al. have demonstrated the advantages of beamformers in electroencephalogram (EEG)-based brain-computer interface (BCI) applications.
[0055] According to (1), the output of the beamformer can be expressed as follows:
[0056]
[0057] Given the source signal location (r), the spatial filter (W T ) should satisfy the following constraint conditions:
[0058] W T (r i )*G(r i ) = 1 (3)
[0059] W T (r i )*G(r j ) = 0 (4)
[0060] For each case where (i) is not equal to (j), 1 and 0 are 3x3 matrices representing the identity matrix and the zero matrix respectively. Assume that (i) is the position of the current source. (j) is regarded as an interfering source, and its influence is minimized. is a good approximation of when the filter output is minimized (q r ), when is minimized:
[0061]
[0062] In this case, this embodiment follows the good approximation of (q r (t)) obtained by LCMV beamforming estimation as the source signal in the cerebral cortex.
[0063] 2) Source channel selection and downsampling:
[0064] The positions of all sources and source signals are determined by the head model and the inverse solution method. Subsequently, regions of interest (ROIs) are extracted according to the cytoarchitecture or organizational structure of the cerebral cortex and the regions defined by cytoarchitecture (called Brodmann areas). The Brodmann areas divide each hemisphere into 52 regions. Each region corresponds to a different brain area, and each brain area is associated with a specific function. Among them, area 4 is located in the precentral gyrus, medial to the central sulcus, corresponding to the primary motor cortex responsible for controlling voluntary movement. Area 6 is located in the precentral gyrus of the frontal lobe in the anterior frontal lobe, corresponding to the premotor cortex that guides sensorimotor coordination and the supplementary motor area that controls the proximal and trunk muscles of the body. Therefore, in this process, as Figure 2 shown, Brodmann areas 4 and 6 located in the frontal lobe together constitute the ROI. Among the 2004 source positions in the head model, by filtering the source positions corresponding to the ROI ( Figure 2 the yellow dots in), a total of 205 source positions belonging to Brodmann areas 4 and 6 are identified; Figure 2 In, (a) is the sagittal view of the head model, and the yellow dots represent the source positions after ROI selection; (b) is the horizontal view of the source positions after dimensionality reduction, where the yellow and red dots together define the ROI, and the red dots represent the source positions selected after dimensionality reduction.
[0065] To further reduce the computational cost and time delay, in this embodiment, the source signals of the ROI are dimensionally reduced through multi-channel signal correlation analysis. During the dimensional reduction process, first, the correlation coefficient "r" between 205 channel signals is calculated pairwise, and a correlation coefficient matrix (205×205) is generated. Then, the following operations are applied to all the values in the matrix: (i) taking the absolute value; (ii) setting the values where (|r|≥0.8) to 1 and the remaining values to 0; (iii) summing the values in each row (or column), arranging them in descending order, and selecting the top 64 channels as the target channels for signal extraction. To minimize the error caused by subject differences, a sample is randomly selected from the data of each subject, and the correlation matrix is averaged over 15 samples. Then, the dimensional reduction procedure is applied to the averaged data. Thus, the 205-channel source signals are reduced to 64-channel source signals (marked as red dots in (b) of Figure 2 ).
[0066] Finally, the signals are downsampled to several frequencies including 256Hz - 8Hz, and the results are verified through classification performance. The classification results show that the performance of the 16Hz signal is the best, and then the signals at this frequency are used for experiments.
[0067] Figure 3 . Schematic diagram of the feature extraction process. (a) Preprocessing multi-channel electroencephalogram; (b) Calculating multi-channel RP for each channel, represented in a 3D structure; (c) Average RP obtained by averaging along the channel axis, represented in a 2D structure; (d) RQA derived from the average RP, represented in a 1D structure.
[0068] B. Feature extraction:
[0069] As Figure 3 shown, Figure 3 in, (a) is the preprocessed multi-channel electroencephalogram; (b) is the multi-channel RP calculated for each channel, represented in a 3D structure; (c) is the average RP obtained by averaging along the channel axis, represented in a 2D structure; (d) is the RQA derived from the average RP, represented in a 1D structure. RP is used for feature extraction, which is a non-linear analysis method that can reveal the hidden dynamic features in EEG signals. The principle behind the formation of RP involves embedding the time series into the phase space and calculating the distance between the trajectory states. When the distance exceeds the threshold, black dots are plotted in the RP; otherwise, white dots are plotted. For a given time series length (x = x i i = 1,..., N), its phase space vector is represented as and the distance between the phase space trajectory states (vectors) is given by the following equation:
[0070]
[0071] Where (R) is a recurrence matrix, (||·||) is a norm (Euclidean norm is used in this embodiment), (ε) is a threshold distance, and (Θ(·)) is the Heaviside function. The values of (R) are 0 and 1, corresponding to the white and black dots in the figure, where the distribution and geometry of the dots reflect the recurrence characteristics of the time series. In addition, since different systems correspond to different thresholds, a value range (from 0.1 to 0.9, with an interval of 0.1) is set in this embodiment, and the classification results are used for comparison. The results show that (ε) = 0.1 yields the best classification performance, so all subsequent experiments are carried out under the condition of (ε) = 0.1.
[0072] For the experimental details, during the feature extraction process, three different types of recurrence features are derived: multi-channel RP, average RP, and recurrence quantification analysis (RQA). The procedure is as Figure 3 shown, where (b), (c), and (d) show the RP features of three stages. Initially, the RP of each individual channel is calculated to generate a three-dimensional matrix, which represents the multi-channel RP feature, and then it is stacked according to the number of channels; subsequently, the multi-channel RP is averaged along the channel dimension to generate a two-dimensional image of the average RP; finally, the quantification eigenvalue of each average RP is calculated by RQA, and the obtained features are represented as a one-dimensional vector.
[0073] C. Attention Residual Network Attention-ResNet:
[0074] ResNet is an advancement in convolutional neural networks. It solves the problem of network degradation by constructing a constant mapping in the residual structure, shortening the path between the network input and output without increasing network complexity or adding extra parameters. This structure also has strong anti-overfitting ability and enhances robustness, making it a very promising deep learning model. In this embodiment, ResNet is used as the base model for classifying the average RP features.
[0075] In this embodiment, ResNet-18 is improved and an original model named ResNet is constructed. Further, it is optimized by integrating three attention modules and three optimization mechanisms, and the impact of these improvements on the model performance is explored. The structure of ResNet consists of five main modules, as Figure 4As shown in the left box, ResNet is shown on the left, which is an improved version of ResNet-18; the three optimization mechanisms adopted in this embodiment are illustrated on the right, where the attention module is inserted after the final feature vector, between ResBlocks, and inside ResBlocks. The data is input in the format of [batch size, number of channels, number of rows, number of columns]. In the first module, two-dimensional convolution, regularization, activation function, and max pooling calculations are sequentially performed; subsequently, the data passes through four large residual modules (ResLayer), each module containing two small residual blocks (ResBlock), and each block performs convolution twice, with a 3×3 kernel and a padding of 1 for each convolution; after average pooling, the final feature vector is passed to the fully connected (FC) layer to obtain a feature vector, which is then input to the Softmax layer for the final calculation, and its output represents the probability distribution of each category in the classification task. The three optimization mechanisms are as Figure 4 shown in the right box, and the three attention modules proposed by Woo et al. include the channel attention module (SAM), the spatial attention module (CAM), and the convolutional block attention module (CBAM), which is a hybrid attention module. In the experiment, the Adam optimizer is used to optimize the cross-entropy loss, the initial learning rate is set to 0.001, each experiment is conducted for 100 training epochs, the training batch size is 64, the ReLU activation function is used, and a dropout rate of 0.3 is used.
[0076] Experiments and Results:
[0077] The proposed decoding framework is applied to decode five categories of electroencephalogram (EEG) data during the period from -1 second to 0 second before the start of movement. The experimental results will be presented in the following sections: First, the dataset and how this embodiment obtains the new dataset required for this embodiment are described. Then the experimental settings are outlined, including the training methods and evaluation metrics used. Before presenting the decoding performance of the framework, the optimal framework is verified. These sections include the verification of the preprocessing steps, the feature extraction process, and the network structure used in the classification process, with the first two parts combined and presented. Then the results of the optimal framework are compared with the current state-of-the-art research results, and finally, the conclusion is drawn with the visualization and interpretability analysis of the feature maps and the model.
[0078] Dataset Description:
[0079] This experiment used an open-source dataset provided by Ofner et al., which is currently the only publicly available dataset that contains both high-density electroencephalogram (61 electrodes) and motion data (glove sensors and exoskeleton data). The dataset contains electroencephalogram data of 15 healthy subjects. During the experiment, the subjects sat on a chair, and their right arms were fully supported by an antigravity exoskeleton and a glove to prevent muscle fatigue. Each trial lasted for 5 seconds. At 0 seconds, the trial started with a beep, and a "+" was displayed on the computer screen. Two seconds later, the subject was prompted to perform the required action. Each subject conducted the experiment once on two different days with an interval of no more than one week between the two experiments. The first stage was actual movement execution, and the second stage was motor imagery. For more detailed information about the experiment, please refer to the original text.
[0080] This dataset contains motor imagery (MI) and motor execution (ME) data for six upper limb movements and the rest (RE) state, including elbow flexion (EF), elbow extension (EE), supination (SU), pronation (PR), hand opening (HO), and hand clenching (HC), all of which are right upper limb movements. In each stage, 10 experiments were conducted on each subject, and each experiment contained 42 instructions, ensuring that each instruction was repeated 6 times in a random order. For each subject, each instruction type was executed 60 times. Importantly, only the motor execution (ME) data was used in this embodiment because it can not only accurately identify the start time of the movement, thus facilitating pre-movement decoding, but also is more natural and verifiable compared to motor imagery (MI). Research has shown that motor execution or attempts can produce better effects in motor rehabilitation.
[0081] This embodiment adopted the method of Ofner et al. to determine the start time of the movement, and then used the start time of the movement as a reference point (time 0), and defined the time period from -1 second to 0 second as a time segment for pre-movement decoding. To accurately identify the start time of the movement, all motion data was visually inspected to eliminate errors caused by data corruption. In addition, since the motion data of hand opening (HO) and hand clenching (HC) was of poor quality, these two groups of data were discarded, and only the data of four movements and the rest state were retained. Finally, the time-frequency data augmentation method proposed by Xie et al. was adopted to address the overfitting problem that may be caused by the limited open-source data. Data augmentation was performed on the same action type group of each subject, so as to maintain the sample ratio across subjects while preserving individual differences. For each group of 60 samples, 120 augmentation processes were performed, and finally the total number of augmented samples in the dataset reached 13,500.
[0082] Cross-validation and evaluation metrics:
[0083] Due to the small dataset, a 5-fold cross-validation method was adopted in this embodiment for training. The individuals were randomly shuffled and divided into five equal parts, four of which were used as the training set, and the remaining one was used as the test set to evaluate the classification accuracy of the model. Finally, the five accuracies obtained were averaged as the overall classification result of the model.
[0084] The evaluation metrics used in this experiment included accuracy (ACC) and confusion matrix. Accuracy represents the proportion of correctly classified samples among all samples and is a basic metric for evaluating the performance of the model. Its calculation formula is as follows:
[0085]
[0086] where TP represents true positive; TN represents true negative; P represents all positive cases; N represents all negative cases.
[0087] The confusion matrix shows the specific number of instances for each classification result. It includes four metrics: true positive (TP), false positive (FP), true negative (TN), and false negative (FN). These metrics together provide a clearer understanding of the performance of the classification model. The diagonal of the confusion matrix represents the prediction accuracy for each class, and the higher the value, the better the classification performance. The values outside the diagonal represent the probability of being misclassified into other classes.
[0088] Verification of the preprocessing and feature extraction process:
[0089] To evaluate the effectiveness of the preprocessing steps and investigate the impact of three different RP feature structures on classification performance, this embodiment retained the signals from different preprocessing stages, including the original EEG signal (61 channels), ESI-EEG signal (205 channels), and ESI-DR-EEG signal (64 channels). For each of these signals, three RP features were subsequently extracted for classification, including multi-channel RP, average RP, and RQA. Accuracy was used as the performance metric to compare the classification results.
[0090] Given that these three different feature structures are represented in 3D, 2D, and 1D formats respectively, three typical classification models were selected to adapt to these structures for parallel comparison during the classification process. For multi-channel RP and average RP, 3D-CNN and 2D-CNN models were used, while for the classification of RQA, a random forest (RF) model was used. For detailed information on the model structures and parameters, please refer to the supplementary information. The 5-fold cross-validation was adopted in the model training process, and the validation results are shown in Table I.
[0091] Table I Verification results of preprocessing steps and feature structures
[0092]
[0093] Note: Due to the huge amount of data of the ESI multi-channel RP feature, a large amount of memory and computing power are required, resulting in significant latency. Therefore, this feature has obvious limitations in the practical application of motion intention detection, and thus this feature is not classified.
[0094] As shown in Table I, for the signals processed by different preprocessing methods, the classification accuracies of all feature structures show the same trend: ESI-DR-EEG > ESI-EEG > raw EEG. For the signals with different feature structures, compared with the image features, the separability of the RQA feature is the lowest. Specifically, the classification accuracy of the RQA feature in the raw EEG is lower than the random probability (20%). In addition, among the image features, the average RP is superior to the multi-channel RP in terms of classification accuracy. The comprehensive comparison shows that the average RP feature after ESI downsampling has the highest classification accuracy, reaching 82.6%.
[0095] Verification of multiple classifier network structures:
[0096] Based on three attention modules and three optimization mechanisms, nine optimized models are developed in this embodiment. These ten models are compared according to the classification accuracy, and the results are shown in Table II. Although ResNet-a and CAM show the most stable optimization results in terms of the optimization mechanism and the attention mechanism respectively, the best classification performance is achieved by ResNet-c-CAM. Compared with ResNet-c-CAM, ResNet-a-SAM also performs well, slightly inferior to ResNet-c-CAM, while ResNet-b-SAM, ResNet-b-CBAM, and ResNet-c-SAM perform poorly compared with ResNet. The comparison of the optimization mechanisms shows an obvious trend, in which the classification performance of ResNet-b is the worst. However, the comparison of the attention mechanisms needs to be analyzed one by one because CAM and SAM show opposite trends in ResNet-a and ResNet-b. The CAM module is more suitable for ResNet-c, while the SAM module is more effective for ResNet-a. In addition, when the CBAM module combines the two attention mechanisms, it is more affected by CAM.
[0097] Table II Accuracy results of 10 models
[0098]
[0099] Figure 5The confusion matrix results of 10 models are shown, and the results of the confusion matrix are basically consistent with the trends observed in the classification accuracy. However, the confusion matrix provides more details. For example, in addition to ResNet-b-SAM, ResNet-b-CBAM, and ResNet-c-SAM (whose accuracies are lower than that of ResNet), there are also significant misclassifications between classes in ResNet-a-CAM and ResNet-b-CAM. In contrast, ResNet-a-SAM, ResNet-a-CBAM, ResNet-c-CAM, and ResNet-c-CBAM maintain the balanced performance of ResNet in the classification of five classes. Among them, ResNet-a-SAM and ResNet-c-CAM with the highest accuracies further improve the recognition rate of each class compared with ResNet, especially overcoming the problem of low recognition rates of ResNet in the EF and SU classes. The interaction between the optimization mechanism and the attention mechanism is more obvious in the confusion matrix. Obviously, among different optimization mechanisms, CAM and SAM show opposite behaviors, while CBAM shows intermediate performance and is more biased towards CAM. For example, the same misclassification pattern appears in ResNet-a-CBAM, while ResNet-c-CBAM shows more balanced recognition performance, similar to ResNet-c-CAM, although its average value is lower.
[0100] Comparison with the state-of-the-art methods:
[0101] In this embodiment, the optimal framework is determined through experimental verification. Table III shows the comparison results between the method proposed in this embodiment and the existing state-of-the-art methods. The results show that the classification framework proposed in this embodiment achieves the highest performance, and the accuracy of five-class classification based on the signals recorded before the motor seizure exceeds 90%. This result significantly improves the multi-class classification performance of the movement-related cortical potential (MRCP).
[0102] Table III Comparison with the state-of-the-art methods
[0103]
[0104] Visualization:
[0105] Feature visualization:
[0106] To further evaluate the effectiveness and necessity of the feature extraction process and visually examine and analyze the grand average readiness potential (RP) feature maps, all the average RPs of the same subject and movement type were averaged, taking subject S1 as an example. In addition, the pre-movement readiness potential signal is a segment of the complete readiness potential signal, which contains additional movement information. To study this, the data was extended by 0.5 seconds in both directions (-1.5 seconds to 0.5 seconds), and a grand average map of the complete readiness potential signal from -1.5 seconds to 0.5 seconds was generated for comparison, as shown in Figure 6 shown in (a) of
[0107] As Figure 6 shown in (b) of
[0108] Network model visualization:
[0109] To more detailedly examine the results of Attention-ResNet and improve the interpretability of the deep learning (DL) model, in this embodiment, the UMAP method was first used to perform dimensionality reduction visualization on the best-performing ResNet-c-CAM model, the worst-performing ResNet-b-SAM model, and the high-dimensional features of ResNet. As Figure 7 shown, there is almost no overlap between categories in the best-performing ResNet-c-CAM model for classification performance, and all categories are clearly separable; in contrast, the ResNet model has a closer spacing between the EF and PR clusters, with a higher degree of confusion between features compared to the ResNet-c-CAM model; for the worst-performing ResNet-b-SAM model, the EE, SU, and PR clusters are well-distinguished, but some overlap and connection are observed between the EF and RE clusters. These results explain the differences in model performance.
[0110] For the best-performing ResNet-c-CAM model, the contributions of different regions of the feature map to the results were subsequently explored through saliency map techniques. As Figure 8 shown, the saliency maps of the four movements are skew-symmetrically distributed, similar to the feature map, and the focus positions of the heatmaps are all near the lower right corner. This region is close to the readiness potential (RP) obtained by calculating the electroencephalogram (EEG) signals near the onset of movement (0 s). When the movement is about to start, the oscillation of the EEG signal increases, which will show stronger features in the phase space and appear on the RP. The classification process of the network captures this important feature, resulting in a greater response in this region; the saliency map of the rest state (RE) pays more attention to the region in the lower left corner of the RP. According to the calculation process of the RP, this region consists of the state distance of the phase space trajectory between the signals close to -1 s and the signals close to 0 s. Compared with the state difference of the movement signal between -1 s and 0 s, the state difference shown by the RE is significantly smaller, which is reflected in the RP and may be an important feature of the RE.
[0111] This embodiment proposes a new pre-movement decoding framework for the movement-related cortical potential (MRCP) and evaluates its performance through experiments. This framework solves two main problems: the first is pre-movement decoding. The goal of this embodiment is to make full use of the advantages of the MRCP signals that appear before the onset of movement to offset the delay in the recognition process. To achieve this goal, this embodiment accurately identifies the moment of movement onset and strictly uses the data before the start of movement. The second problem is the multi-class classification problem. This embodiment proposes a framework that can be applied to multi-class MRCP in various paradigms or scenarios, significantly improving the classification performance.
[0112] By comparing the classification results of the signals with different preprocessing procedures and feature types (Table I), it can be observed that compared with the original electroencephalogram (EEG), the signals processed by equivalent source imaging (ESI) show better classification performance. This indicates that ESI enhances the spatial resolution of the signals, thereby extracting deeper information and improving the separability of the signals. The improvement of the classification results after dimensionality reduction indicates that the dimensionality reduction process may extract more consistent signals, further enhancing the signal features. The source position map in the channel selection process ( Figure 2As shown in (b) of , due to the differences between the Montreal Neurological Institute (MNI) and Talairach coordinates and the errors in the coordinate conversion process, the source positions at the edges are not completely symmetric, and the filtered source positions still contain some errors. However, the dimensionality reduction process removes the interference of the signals at the edges of the region of interest (ROI) by selecting highly correlated features, eliminating the possible interference caused by coordinate conversion. In addition, this result shows that the channels near the center of the ROI exhibit stronger functional connectivity, which is consistent with the physiological structure of the brain, further verifying the effectiveness and necessity of the preprocessing steps. The comparison of the results of different feature types shows that the recurrence plot (RP) images may contain richer information, while the recurrence quantification analysis (RQA) metrics do not fully capture the complete or key information in the images. The difference between the average RP and multi-channel RP results may be attributed to the fact that averaging the RP across multiple channels reduces the amount of inter-channel noise, thereby enhancing the features and improving the signal-to-noise ratio (SNR) and separability. To improve the framework, two parameters were experimentally verified. According to the results, the downsampling frequency was set to 16 Hz, which may be because reducing the signal frequency suppresses high-frequency interference, thereby improving the separability of the features. In addition, 16 Hz is also the signal frequency used in many MRCP-based classification studies, which is somewhat consistent with the empirical values. The setting of the RP threshold has been discussed in many studies, and the threshold is mainly related to the distance between the trajectories of brain activity signals in the phase space. Therefore, it must be discussed according to different signals or research objectives. The optimal threshold obtained here is only a conclusion based on the conditions of this framework.
[0113] A series of comparative experiments were conducted on the multi-class classifier Attention-ResNet. Among the three fusion mechanisms, ResNet-b showed significant performance limitations (Table I). This may be because an attention module was inserted between the two ResBlocks of ResNet-b, partially destroying the residual structure. When the shortcut branch is enabled, a multi-layer attention mechanism structure may be formed, which instead inhibits the advantages of ResNet. The comparison of the attention mechanisms shows that among all three optimization mechanisms, the class activation mapping (CAM) exhibited stable performance, while the results of the spatial attention mechanism (SAM) and the convolutional block attention module (CBAM) showed greater variability under different optimization mechanisms (Table Ι). This shows that, in the context of this embodiment, CAM provides better robustness for the model, while SAM is more affected by the model architecture and optimization mechanisms and requires more precise conditions. In addition, when combining the accuracy (Table I) and the confusion matrix results ( Figure 5) When [condition], it was found that the models combining CAM and SAM in ResNet-a and ResNet-c showed opposite trends. This indicates that there is an interaction between the optimization mechanism and the attention mechanism. In addition, the results of CBAM show that its trend is similar to that of CAM, which means that in the hybrid attention mechanism, the influence of the channel attention mechanism may be more significant.
[0114] In the research related to deep learning-based electroencephalogram (EEG) decoding, time series signals are usually used as input features to make full use of the end-to-end classification model while minimizing data preprocessing. Although the movement-related cortical potential (MRCP) features are mainly observed in the time domain, the amplitude change over time is relatively simple. This leads to lower feature separability during multi-class decoding, thus limiting the decoding performance. Recurrence plot (RP) is a method for visualizing the recurrence characteristics of a dynamic system and can reveal the non-linear dynamics represented by the time series. It transforms the time series into an N×N two-dimensional matrix and usually performs better than other complex time series network methods. In addition, the brain is a complex non-linear dynamic system, and EEG signals have non-linear, dynamic, and complex characteristics. Therefore, RP features can reveal the hidden dynamic characteristics in EEG signals, thereby improving feature separability. The result comparison between Table I and Table III shows that after feature extraction by the method proposed in this embodiment, the classification results of the average RP features using the standard convolutional neural network (CNN) exceed the results of several advanced studies, which strongly proves the effectiveness and accuracy of the feature extraction process.
[0115] Finally, the average large-scale recurrence plot (RP) of the subjects and the high-order features of the model were visualized. Importantly, the result of including the complete movement-related cortical potential (MRCP) signal in the average large-scale RP is based on a reasonable assumption that the pre-movement signal and the complete MRCP signal show similar trends and characteristics in the feature map. Therefore, more information can be obtained from the feature map of the complete MRCP signal. From the result image ( Figure 6 ) it can be observed that the average large-scale RP of the complete MRCP signal shows good traceability and interpretability in terms of movement categories. Although its feature distinctiveness and pairing are poor due to the limitations of signal length and the region of interest, this trend can also be extended to the pre-movement MRCP signal. This is still one of the challenges in multi-class classification. In terms of model visualization, the high-dimensional feature visualization of models with different performance levels partially explains the result differences and enhances the interpretability of the model. In addition, the combination of the average large-scale RP ( Figure 6 ) and the saliency map analysis of ResNet-c-CAM ( Figure 8 ) shows that the model correctly focuses on important features, making it an effective network architecture.
[0116] Although some studies have explored pre-movement decoding, there is still much room for improvement in the classification performance of multiple classes, far from achieving the level of natural movement intention detection. This embodiment proposes a multi-class movement intention decoding framework that uses only the movement-related cortical potentials (MRCPs) before the start of movement to classify five categories: upper limb movement or rest. The optimal structure and parameters of the framework are determined through experimental verification. The results show that the proposed framework is both reasonable and effective, achieving a classification accuracy of 0.915 on a public dataset.
[0117] This embodiment aims to contribute to the development of a more natural brain-computer interface (BCI) system. To this end, movement-related cortical potential (MRCP) signals are selected because of their high potential in practical applications. The method of this embodiment not only improves the classification performance but also shows considerable promise for real-time applications in online BCI systems, which is worthy of further exploration in future research.
[0118] Movement-related cortical potential signals (MRCPs) have great potential in realizing a rehabilitation brain-computer interface (BCI) system based on Hebbian theory, but the accuracy of multi-class decoding limits the development of a more natural control process. This embodiment proposes a novel multi-class decoding framework that can decode five different upper limb movements using only pre-movement signals. In the preprocessing and feature extraction stages, electrophysiological source imaging (ESI) is used to improve the spatial resolution and signal-to-noise ratio (SNR) of the signals. Subsequently, recurrence plot (RP) technology is adopted to capture the dynamic behavior of the system, thereby enhancing the separability of the extracted features. Finally, these features are classified by an enhanced attention residual network (Attention-ResNet). In 5-fold cross-validation, this method achieved an average classification accuracy of 91.5%, showing a significant improvement compared to the current state-of-the-art methods. The results show that the proposed pre-movement decoding method for upper limb movements significantly improves the multi-class decoding performance based only on pre-movement signals, providing a valuable reference for the development of a more natural BCI system.
[0119] The above-described embodiments merely represent the specific implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention.
Claims
1. An upper limb pre-movement decoding method, characterized in that, It includes the following steps: Step 1: Extract the pre-movement period of interest and perform data augmentation to construct a new dataset; Step 2: Preprocessing: Directly apply electrophysiological source imaging (ESI) to the data, then perform source channel selection and downsampling to obtain source signals; Step 3: Feature extraction: Calculate the recurrence plot (RP) for each channel of the source signal and obtain the average recurrence plot; Step 4: Use the average recurrence plot as the input, train the attention residual network Attention-ResNet and perform classification.
2. The upper limb pre-movement decoding method according to claim 1, wherein In Step 2, applying electrophysiological source imaging (ESI) to the data includes forward modeling and inverse problem solving; Forward modeling projects the neuronal source activities in the brain onto the scalp electrodes, taking into account the geometry and conductivity characteristics of various tissues during the signal propagation process. The projection process is as shown in Equation (1): x(t) = Gq r (t) (1) where r = 1, ..., N s , is the vector of the three-dimensional directional primary current related to position 3×N s , with a magnitude of q r (t), where N s represents the number of possible source positions in the cortex, G is the lead field matrix, i.e., the head model, which maps the primary current q r (t) to the scalp signal x(t), with a size of N c ×3N s , where N c is the number of scalp electrodes; The inverse problem is solved by backtracking and estimating the source q r (t) corresponding to the acquired scalp signal x(t) and lead field matrix G. Using a linearly constrained minimum variance beamformer, the source activity at each source location is estimated through spatial filtering. According to Equation (1), the output of the beamformer is expressed as follows: Given the source signal position r, the spatial filter W T satisfies the following constraint conditions: W T (r i )*G(r i )=1 (3) W T (r i )*G(r j ) = 0 (4) For each case where i is not equal to j, 1 and 0 are 3x3 matrices representing the identity matrix and the zero matrix respectively; assume that i is the position of the current source; j is regarded as the interfering source whose influence is minimized; is a good approximation of q when the filter output is minimized r when is minimized: The good approximation q obtained by LCMV beamforming estimation r (t) as the source signal in the cerebral cortex.
3. The upper limb movement pre-decoding method according to claim 2, wherein In Step 2, the source channel selection and downsampling are specifically as follows: The source channel selection includes region of interest (ROI) selection and dimensionality reduction. Determine the positions of all sources and source signals through the head model and inverse solution method; Subsequently, extract the ROI according to the regions defined by the cytoarchitecture or tissue structure and cytoarchitecture of the cerebral cortex; Perform dimensionality reduction on the source signals of the ROI through multi-channel signal correlation analysis. During the dimensionality reduction process, first calculate the correlation coefficient r between channel signals in pairs and generate a correlation coefficient matrix; Then apply the following operations to all values in the matrix: (i) Take the absolute value; (ii) Set the values where |r| is greater than or equal to the preset threshold to 1, and set the remaining values to 0; (iii) Add the values of each row or column, sort them in descending order, and select the first n channels as the target channels for signal extraction; Finally, downsample the signal.
4. The upper limb pre-movement decoding method according to claim 1 or 3, characterized in that, In Step 3, use RP for feature extraction, embed the time series into the phase space and calculate the distance between trajectory states; When the distance exceeds the threshold, draw black dots in the RP; Otherwise, draw a white dot; for a given time series length \(x = x\) i \(i = 1,\ldots,N\), whose phase space vector representation is And the distance between phase space trajectory states is given by the following equation: Where R is the recurrence matrix, ||·|| is the norm, ε is the threshold distance, and Θ(·) is the Heaviside function; The values of R are 0 and 1, corresponding to the white dots and black dots in the RP, where the distribution and geometry of the dots reflect the recurrence characteristics of the time series.
5. The upper limb movement pre-decoding method according to claim 1, characterized in that In Step 4, the data first sequentially performs two-dimensional convolution, regularization, activation function, and max pooling calculations; Subsequently, the data passes through four large residual modules, each module contains two small residual blocks, and each block performs two convolutions, each convolution using a 3×3 kernel and a padding of 1; After average pooling, the final feature vector is passed to the fully connected (FC) layer to obtain the feature vector, and the feature vector is then input to the Softmax layer for final calculation, and its output represents the probability distribution of each category in the classification task; 6. The upper limb pre-movement decoding method according to claim 5, wherein The attention residual network Attention-ResNet includes channel attention modules, spatial attention modules, and convolutional block attention modules inserted after the final feature vector, between small residual blocks, and inside small residual blocks.