Method and system for dyskinesia rehabilitation training based on electroencephalographic signal recognition
By using self-attention convolution neural network and multi-time attention mechanism in EEG signal recognition, and combining feature distillation and logical distillation optimization models, the complex problems of noise interference and feature extraction are solved, and the effect of efficient identification and rapid recovery is achieved, and the real-time needs of motor dysfunction rehabilitation training is met.
Patent Information
- Application Number
- PCT/CN2023/134054
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art faces the problems of noise interference, complex feature extraction and high real-time requirements in the recognition of EEG signals, especially in the rehabilitation training of motor disorders, which are difficult to achieve efficient identification and rapid recovery.
The self-attention convolution neural network is used to combine a multi-time attention mechanism to extract the spatiotemporal feature of EEG signals, and optimize the model through feature distillation and logical distillation to reduce the complexity of the network structure to improve the recognition rate.
It realizes efficient filtering and feature extraction of EEG signals, improves recognition accuracy and real-time performance, and meets the needs of rapid recovery of motor dysfunction rehabilitation training.
Smart Images

Figure CN2023134054_30052025_PF_FP_ABST
Abstract
Description
A movement disorder rehabilitation training method and system based on EEG signal recognition Technical Field
[0001] The present invention relates to the technical field of electroencephalogram (EEG) signal processing, and in particular to a movement disorder rehabilitation training method and system based on EEG signal recognition. Background Art
[0002] Motor imagery EEG signals contain information about movement-related neural activity. Collecting, classifying, and identifying these signals through brain-computer interface technology is a promising solution for helping patients with movement disorders effectively recover their motor abilities. However, there are currently two major challenges: 1. The weak, non-stationary, and nonlinear nature of EEG signals complicates feature extraction; 2. High real-time performance requirements are required, including high recognition speed and accuracy.
[0003] To address the above issues, traditional EEG signal decoding is based on machine learning and manual feature extraction of signals. The recognition accuracy depends on the effectiveness of feature selection, and the data features are single and subject to a lot of noise interference. Although the application of deep learning technology can extract effective identification features from complex EEG signals, the resulting network model structure is complex and cannot meet the real-time requirements of rehabilitation training. It also requires a certain amount of data set training, and long-term training will make patients feel tired, resulting in poor data results.
[0004] Therefore, how to design an EEG signal recognition model so that it can efficiently filter EEG signals to remove noise interference, use limited data to extract deep spatial and temporal features from multiple dimensions and scales, and reduce the network structure as much as possible to speed up the recognition rate while improving the classification and recognition accuracy is a key issue that must be faced and solved in the current research and application of multi-category EEG signal recognition in limb movement disorder rehabilitation training.
[0005] Summary of the Invention
[0006] To address the problems of the above-mentioned prior art, the present invention provides a method and system for movement disorder rehabilitation training based on EEG signal recognition. The purpose is to extract features that better characterize motor imagery to improve signal recognition accuracy, while also accelerating the formation speed and ensuring the rate of motor imagery EEG signal recognition, thereby achieving high-speed and efficient movement disorder rehabilitation training. To achieve the above-mentioned objectives, the present invention provides the following technical solutions:
[0007] In a first aspect, a method for rehabilitation training of movement disorders based on EEG signal recognition is provided, comprising:
[0008] S1, performing offline training on the patient to collect training phase data, wherein the training phase data is multimodal EEG signals of the movement disorder patient when imagining different types of movements;
[0009] S2, preprocessing the multimodal EEG signal, including establishing a frequency domain filter based on bandpass filtering, normalization processing, and improving the common spatial pattern algorithm to establish a spatial domain filter;
[0010] S3, using the preprocessed EEG signal to establish a self-attention convolutional neural network to extract the spatiotemporal features of the EEG signal, and determining a first loss function based on the predicted classification and the true classification of the self-attention convolutional neural network, and determining a second loss function based on the distance between the spatiotemporal features and the center of the corresponding spatiotemporal features;
[0011] S4, using the deep network layer of the self-attention convolutional neural network as a teacher model and the shallow network layer as a student model, performing feature distillation and logical distillation on the self-attention convolutional neural network to optimize the model, and obtaining a third loss function; the third loss function includes a feature similarity loss function in the feature distillation process and a classification loss function in the logical distillation process;
[0012] S5, constructing a model total loss function based on a linear combination of the first loss function, the second loss function, and the third loss function, and iteratively training the self-attention convolutional neural network to obtain an EEG signal recognition model;
[0013] S6, collecting the patient's EEG signals from online training as rehabilitation stage data, and inputting them into the EEG signal recognition model established during the patient's most recent offline training to obtain classification and recognition results;
[0014] S7, converting the classification and recognition results into control commands recognizable by an external device, driving corresponding limb movements based on the control commands, and completing movement disorder rehabilitation training for the patient.
[0015] In a second aspect, a movement disorder rehabilitation training system based on EEG signal recognition is provided, comprising:
[0016] The EEG signal acquisition module is used to collect EEG signals when patients imagine different types of movements;
[0017] A signal processing module is used to perform preprocessing operations on the collected EEG signals, including a frequency domain filtering unit and a normalization processing unit; the frequency domain filtering unit is used to establish a frequency domain filter to filter out irrelevant noise in the EEG signals in the frequency domain, the normalization processing unit is used to reduce the volatility and non-stationarity of the data, and the spatial domain filtering unit is used to establish a spatial domain filter to maximize the difference in variance values of multiple types of EEG signals;
[0018] The recognition network building module is used to build a self-attention convolutional neural network to extract the spatiotemporal features of the preprocessed EEG signals;
[0019] A distillation module, configured to optimize the self-attention convolutional neural network model from two aspects: feature distillation and logic distillation.
[0020] a loss function calculation module, configured to determine a first loss function based on the predicted classification and the true classification of the self-attention convolutional neural network, determine a second loss function based on the distance between the spatiotemporal feature and the center of the corresponding spatiotemporal feature, determine a third loss function based on the distillation module, and construct a total model loss function based on a linear combination of the first loss function, the second loss function, and the third loss function;
[0021] A model training module, configured to iteratively train the self-attention convolutional neural network based on the model total loss function to obtain an EEG signal recognition model;
[0022] A classification and recognition module, configured to classify and recognize EEG signals based on the EEG signal recognition model and convert them into control commands recognizable to external devices;
[0023] The motion execution module is used to execute the control command and drive the corresponding exoskeleton device to perform rehabilitation training of the corresponding limb.
[0024] The present invention provides a method and system for rehabilitation training of movement disorders based on EEG signal recognition, which has the following beneficial effects:
[0025] 1. The present invention filters and standardizes EEG signals from the frequency domain and spatial domain respectively, effectively reducing the influence of irrelevant noise. Among them, the improved common space pattern algorithm extends the two-classification algorithm to the multi-classification range to maximize the difference in variance values of multiple types of signals, which can greatly reduce the complexity and time overhead of the overall algorithm. Considering the problem that the eigenvalues of different categories in the common space pattern obtained by multiple classification tasks may be the same, by selecting as few eigenvectors as possible based on whether the eigenvalues are the same under the premise of maximum importance, it can ensure the projection space requirement of maximizing the energy difference of multiple types of EEG signals while effectively reducing the dimension of the eigenvector. Compared with the prior art that directly uses the size of the eigenvalue to select, it can accurately filter and classify EEG signals of different movement types.
[0026] 2. The present invention adopts a multi-head self-attention mechanism with multiple time attention levels to establish a global self-attention module group, which expresses the input data from different dimensions while considering the relationship between all data in the spatiotemporal feature sequence of the EEG signal, thereby improving the expression ability of global relevant information. The present invention compresses the attention weight matrix from a global level, pays more attention to the effective features of some key time frames, reduces the number of dot product operations, and speeds up the feature extraction. The mean feature and the fusion of multiple self-attention units with different time attention levels make up for the fact that the aforementioned compression process only focuses on important information. Compared with the self-attention mechanism of the prior art, which directly ignores minor information during weighted average calculation, the integrity of the feature extraction of EEG signals can be improved.
[0027] 3. The present invention takes into account the complex characteristics of the EEG signal itself and the multi-head self-attention mechanism with multiple temporal attention in the feature extraction, which makes the feature extraction process more complicated. Therefore, the EEG signal recognition model is optimized from the perspectives of feature distillation and logical distillation through self-distillation training, and the network with more parameters and more complex structure is refined into a target network with smaller scale and smaller time delay, which can reduce the complexity of the network structure to ensure that the complexity and real-time requirements of the EEG signal are met. In feature distillation, a student model with self-matching and self-selecting teachers is adopted, so that shallow features select deep features suitable for their own learning and move closer to them, and perform knowledge supervision training on feature similarity; in logical distillation, the difference between the classification results of two different deep and shallow layers is obtained through classification loss, which can better guide the shallow network to learn spatial features.
[0028] 4. The present invention provides a method and system for rehabilitation training for movement disorders based on EEG signal recognition. This method performs multi-faceted signal preprocessing before feature extraction, effectively filtering out noise and better separating different signals. During feature extraction, a low-complexity model is established, extracting features from both temporal and spatial dimensions and incorporating a multi-temporal self-attention mechanism to capture the correlation between features' global temporal positions. Self-distillation learning is also incorporated to accelerate network fitting and improve the system's real-time performance. The introduction of multiple loss function optimization constraints enhances classification and recognition capabilities, enabling efficient and accurate rehabilitation training. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] FIG1 is a flow chart of a movement disorder rehabilitation training method based on EEG signal recognition according to the present invention;
[0030] FIG2 is a schematic diagram of a flow chart of spatial domain filtering using the improved common spatial pattern algorithm of the present invention;
[0031] FIG3 is a schematic diagram of the structure of the self-attention convolutional neural network established by the present invention;
[0032] FIG4 is a block diagram of the decomposition structure of the global self-attention module group using multiple temporal attention levels of the present invention;
[0033] FIG5 is a structural block diagram of a movement disorder rehabilitation training system based on EEG signal recognition according to the present invention. DETAILED DESCRIPTION
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0035] First, an embodiment of the present invention provides a movement disorder rehabilitation training method based on EEG signal recognition, as shown in FIG1 , comprising:
[0036] S1, performing offline training on the patient to collect training phase data, wherein the training phase data is multimodal EEG signals of the movement disorder patient when imagining different types of movements;
[0037] Patient rehabilitation training consists of two parts: the training phase and the rehabilitation phase. During the training phase, it is necessary to collect the patient's EEG data to establish a labeled motor imagery EEG signal dataset for subsequent model training;
[0038] S2, preprocessing the multimodal EEG signal, including establishing a frequency domain filter based on bandpass filtering, normalization processing, and improving the common spatial pattern algorithm to establish a spatial domain filter;
[0039] S3, using the preprocessed EEG signal to establish a self-attention convolutional neural network to extract the spatiotemporal features of the EEG signal, and determining a first loss function based on the predicted classification and the true classification of the self-attention convolutional neural network, and determining a second loss function based on the distance between the spatiotemporal features and the center of the corresponding spatiotemporal features;
[0040] S4, using the deep network layer of the self-attention convolutional neural network as a teacher model and the shallow network layer as a student model, performing feature distillation and logical distillation to obtain a third loss function; the third loss function includes the feature similarity loss function in the feature distillation process and the classification loss function in the logical distillation process;
[0041] S5, constructing a total loss function of the model based on the linear combination of the first loss function, the second loss function, and the third loss function, and iteratively training the self-attention convolutional neural network to obtain an EEG signal recognition model;
[0042] S6, collecting the patient's EEG signals during online training as rehabilitation stage data, and inputting the patient's rehabilitation stage data into the EEG signal recognition model established during the patient's most recent offline training to obtain classification and recognition results;
[0043] During the rehabilitation phase, patients undergo rehabilitation testing using the EEG signal recognition-based movement disorder rehabilitation training system of the present invention. The acquisition module transmits rehabilitation data to the classification and recognition module, which then uses the trained EEG recognition model to predict the results and displays the predicted results on the screen. To account for the fact that EEG signals vary over time, the recognition model established during the patient's most recent offline training session must be loaded.
[0044] S7, converting the classification and recognition results into control commands recognizable by an external device, driving corresponding limb movements based on the control commands, and completing movement disorder rehabilitation training for the patient.
[0045] The pre-processing operation in step S2 includes:
[0046] S21, bandpass filtering is performed using a type II Chebyshev filter to filter the data at 4-40 Hz to remove frequency components not related to motion. The transfer function is: ω0 is the effective passband cutoff frequency, ε is a parameter related to the passband ripple, 0<ε<1, T n The present invention uses a 6th-order Chebyshev filter to maintain the rhythm related to the motor imagery task.
[0047] S22, use Z-score normalization to normalize the non-standardized EEG signal to balance and normalize it to eliminate the fluctuation interference during the acquisition process: where x i and x0 represent the bandpass filtered data and the normalized output, μ and σ respectively. 2 represent the mean and variance of the training data respectively.
[0048] Preprocessing multimodal EEG signals can reduce the impact of noise and other non-research factors on EEG feature extraction and decoding, thereby extracting useful signal components. Bandpass filtering can filter out irrelevant high- and low-frequency noise, and normalization can reduce data volatility and non-stationarity, facilitating model training.
[0049] S23, using the improved common space pattern algorithm to design a spatial domain filter. This embodiment takes the three-classification of EEG signals as an example, see Figure 2, and the specific steps are as follows:
[0050] S231, grouping the multimodal EEG signals according to motion types to form n types of EEG signals, where n is the total number of EEG signal motion types;
[0051] S232, calculate the normalized covariance matrix of each type of EEG signal i characterizes different motion types; and obtains a mixed space covariance matrix R of the multimodal EEG signal;
[0052] Specifically, calculate the covariance matrix Ri of various EEG signals: where X i are EEG signals of different movement types, T represents matrix transpose, and trace() represents the sum of the elements on the diagonal of the matrix;
[0053] The covariance matrix Ri of each type of EEG signal is averaged to obtain the normalized covariance matrix
[0054] The mixed space covariance matrix R is:
[0055] S233, perform principal component decomposition on the mixed space covariance matrix R of the EEG signal: R = UVU T ,
[0056] Where V is the eigenvalue diagonal matrix, and U is the eigenvector matrix corresponding to the eigenvalues in V;
[0057] S234, based on the normalized covariance matrix of various EEG signals And principal component decomposition to obtain the common eigenvector matrix S i , and S i The principal component decomposition is performed in step S233 to obtain the eigenvalue diagonal matrix V of each type of EEG signal. i Sum and eigenvalue diagonal matrix V i The corresponding eigenvector matrix U i ;
[0058] Specifically, find the common eigenvector matrix S i include:
[0059] Calculate the whitening matrix P, Then the common eigenvector matrix S i for
[0060] For the common eigenvector matrix S i The principal component decomposition is performed in step S233 to obtain the eigenvalue diagonal matrix V of each type of EEG signal. i and the corresponding eigenvector matrix U i :
[0061] Where U1=U2=…=U i , V1+V2+…+Vi =I, I is the unit matrix.
[0062] Using steps S231-S234, we can obtain U1=U2=U3, V1+V2+V3=I.
[0063] Since the classical co-spatial pattern algorithm is a spatial domain filtering method for two-classification tasks, and the present invention includes multiple types of motion imagery, it needs to be expanded.
[0064] S235, for all the common eigenvector matrices S i Perform approximate joint diagonalization to obtain the relevant diagonal matrices corresponding to various types of EEG signals, and calculate the eigenvalues λ in each of the relevant diagonal matrices j The importance of the characteristic value λ j and the eigenvalue λ j The larger value of the inverse proportional function of
[0065] It should be understood that the existing technology is used to perform approximate joint diagonalization on the matrix, which will not be described in detail here.
[0066] Specifically, the importance is the eigenvalue λ j and The larger value among them; where τ is the eigenvalue λ j The inverse proportional function of α is determined according to the number of classification tasks.
[0067] S236, in each related diagonal matrix, the eigenvalue λ is sorted according to the importance. j Sort in descending order and record the eigenvalue λ corresponding to the maximum importance value in each related diagonal matrix j The same number of values;
[0068] The existing common spatial pattern algorithm sorts the eigenvalues and selects the maximum eigenvalue to construct the filter. However, considering that the eigenvalues of different classifications in the common spatial pattern spatial domain filtering obtained from multiple classification tasks may be the same, the present invention uses the importance of the eigenvalue to select the appropriate eigenvalue.
[0069] S237, if the eigenvalue λ corresponding to the maximum importance value in each related diagonal matrix j If the values of are the same, the eigenvalues λ of the first n eigenvalues in importance order are used. j The corresponding eigenvector is subjected to spatial filtering;
[0070] In this embodiment, n=3, and if the eigenvalues with the greatest importance are all the same, the first three columns of the relevant eigenvector matrix corresponding to the relevant diagonal matrix are selected.
[0071] If the eigenvalue λ corresponding to the maximum importance value in each related diagonal matrix jIf the number of eigenvalues with the same value is m, and m < n, then the top m + 1 eigenvalues λ sorted by importance are used. j The corresponding eigenvectors are used to construct a spatial filter.
[0072] The present invention sorts according to the importance of eigenvalues, which can avoid the situation that different motor imagery classifications may correspond to the same eigenvector, and solves the problem that the common spatial pattern is only applicable to the binary classification problem of finding the maximum and minimum eigenvalues.
[0073] Denote the matrix composed of the eigenvectors corresponding to the eigenvalues as the relevant eigenvector matrix. In this embodiment, n = 3, that is, m can be 0 or 1 (when m = 2, they are all the same);
[0074] If m = 0, that is, the eigenvalues λ corresponding to the maximum importance values in each relevant diagonal matrix j are all different. At this time, the eigenvectors corresponding to the eigenvalues with the maximum importance in each relevant diagonal matrix are used for spatial filtering, that is, the first column of the relevant eigenvector matrix is selected.
[0075] If m = 1, that is, the eigenvalues λ corresponding to the maximum importance values in each of the said relevant diagonal matrices j have two identical values. At this time, the eigenvectors corresponding to the eigenvalues of the first m + 1 = 2 columns sorted by importance in each category are used for spatial filtering, that is, the first two columns of the relevant eigenvector matrix are selected.
[0076] The spatial filtering step after eigenvalue selection in the present invention is the same as the common spatial pattern algorithm in the prior art, and will not be elaborated here.
[0077] By using the common spatial pattern algorithm to maximize the variance value difference of multi-class signals, eigenvectors with high discrimination are obtained. Considering the limitation of the number of classifications and the problem that the eigenvalues of different classifications in the common spatial pattern spatial filtering obtained from multiple classification tasks may be the same, the present invention improves the common spatial pattern algorithm by using the importance of eigenvalues and whether the selected eigenvalues are the same, and extends the binary-class common spatial pattern algorithm to the multi-class range. Compared with the one-versus-one and one-versus-many strategies in the prior art, it can greatly reduce the complexity and time overhead of the overall algorithm. At the same time, on the premise of the maximum importance, by selecting as few eigenvectors as possible according to whether the eigenvalues are the same, it can ensure the requirement of maximizing the energy difference projection space of multi-class EEG signals while effectively reducing the dimension of the eigenvectors. Compared with directly using the magnitude of eigenvalues to select in the prior art, it can more accurately filter and classify EEG signals of different motion types.
[0078] As shown in Figure 3, the self-attention convolutional neural network structure in step S3 is as follows:
[0079] Feature extraction layer: Perform time and space convolution along the time dimension and channel dimension to extract spatiotemporal features, and use the spatiotemporal features through the self-attention module to extract the global temporal correlation information of the signal. It specifically includes four layers:
[0080] The first layer uses k convolution kernels of size (1,25) and stride (1,1) to perform convolution operations to extract the time domain features of the EEG signal;
[0081] The second layer uses a kernel of size (N, 1) and a stride of (1, 1) to extract spatial features;
[0082] The third layer is the pooling layer of the time dimension, with a kernel size of (1,75) and a step size of (1,15); it is used to smooth the time features, which not only avoids overfitting but also reduces the computational complexity.
[0083] The fourth layer uses a self-attention module to obtain global temporal position correlation. This effectively compensates for the convolution operation's focus on the local receptive field, and the resulting features are more discriminative.
[0084] Center loss layer: The initial center point of each EEG signal category is set to a zero vector or a random vector, the spatiotemporal feature output of the fourth layer of the feature extraction layer of the convolutional neural network is used as the sample feature vector, and the Euclidean distance between each sample feature vector and the center point of its corresponding category is calculated as the second loss function.
[0085] The center loss layer is used to shrink the distance between EEG signals of the same category and increase the distance between EEG signals of different categories, so as to improve the performance of the EEG signal recognition model.
[0086] Classification layer: Use the fully connected layer classifier to predict and classify the spatiotemporal features extracted by the feature extraction layer, and calculate the cross entropy loss function between the predicted classification result and the real EEG signal classification label as the first loss function;
[0087] The self-attention module is a global self-attention module group, in which the self-attention units use a multi-head scaled dot-product attention mechanism. Each self-attention unit at each temporal attention level includes a multi-head self-attention mechanism and a fully connected network, each of which is connected to a residual connection and a normalization module. Multiple self-attention units at different temporal attention levels are connected in parallel via a splicing normalization layer. The specific structure of the global self-attention module group of the present invention is shown in Figure 4.
[0088] The output of the global self-attention module group of the present invention is
[0089] Among them, Norm is the layer normalization operation, MulH T Represents the output of a self-attention unit with temporal attention of T.
[0090] The multi-head self-attention mechanism considers the relationships between all data in the spatiotemporal feature sequence while expressing the input data from different dimensions, capturing different attention information and improving the ability to express globally relevant information. Residual connections provide cross-layer connectivity for feature information, thereby reducing the difficulty of training deep neural networks. Layer normalization aligns the output of each self-attention module with a uniform distribution, allowing the global self-attention module group to adapt to the changes in the compression of the attention matrix due to different temporal attention levels, thereby accelerating model convergence.
[0091] The present invention establishes a global self-attention module group as follows:
[0092] S51, determining the temporal attention of multiple self-attention units based on the time series length of the spatiotemporal features and the number of self-attention units, and establishing a global self-attention module group with multiple self-attention units of different temporal attentions connected in parallel;
[0093] The global self-attention module group contains multiple self-attention units. The internal composition structure of the multiple self-attention units is the same, but the temporal attention is different, that is, the temporal receptive field of each self-attention unit is different.
[0094] Specifically, this embodiment adopts three self-attention units with different temporal attention levels, namely T, T / 2, and T / 4, where T is the time series length of the spatiotemporal feature sequence, and the temporal receptive fields are the entire / one-half / one-quarter of the spatiotemporal feature sequence, respectively. That is, when processing the sequence, each element is associated with other elements in the entire / one-half / one-quarter of the sequence, and the relative importance between the elements is calculated to adaptively capture the long-range dependencies between the elements.
[0095] S52, linearly mapping the spatiotemporal features based on the self-attention units of the different temporal attention levels to obtain attention matrices under corresponding temporal attention levels: Q matrix, K matrix, and V matrix; the attention matrix has the same dimension as the matrix of the spatiotemporal features;
[0096] The linear mapping formula is: Q = XW q , K=XW k , V=XW v . Where X is the spatiotemporal feature sequence, W q , W k , W v are the learnable weight matrices required to linearly map X to Q, K, and V, respectively.
[0097] Considering that EEG signals often show a periodic change in the time domain, and the changes between adjacent time frames are relatively slow, and the feature changes are not obvious, it is necessary to compress the length of the time series and extract important time frames.
[0098] S53, calculating the cumulative eigenvalues of each eigenvector in the same time frame in the K matrix in the corresponding self-attention unit to extract the key representation vector from the K matrix;
[0099] S54, calculating a first attention weight matrix based on the key representation vector and the Q matrix, and performing a compression operation on the Q matrix according to the weights in the first attention weight matrix to obtain a compressed Q matrix;
[0100] S55, using a zero vector to complete the dimension of the compressed Q matrix to the same dimension as the K matrix, to obtain a key Q matrix;
[0101] Since the subsequent attention weights need to perform dot product operations on the K matrix and the Q matrix, the dimensions of the two need to be the same. Therefore, the time frames corresponding to the Q values removed in the compression are filled with zero vectors.
[0102] S56, calculating a second attention weight matrix based on the K matrix and the key Q matrix, and performing weighted summation on the V matrix to obtain the output of the corresponding self-attention unit;
[0103] in is the compressed Q matrix, d is the dimension of K matrix, and T represents the transpose of the matrix. Then the output of the self-attention unit is: MulH=Concat(head1,head2,…,head h )W
[0104] Among them, Concat represents the concatenation operation, A 2,h represents the second weight matrix of the h-th self-attention head, head h represents the hth self-attention head, and W is a learnable parameter matrix used to fuse multi-head self-attention information.
[0105] S57, performing mean filling on the output of the self-attention unit using the V matrix of the spatiotemporal features;
[0106] S571, calculating the mean value of the V matrix of the spatiotemporal features within the corresponding time attention range;
[0107] When the time attention is T, the mean calculation formula is:
[0108] T is time attention, V i,Tis the i-th V value of the spatiotemporal feature in the T-th time frame of the V matrix.
[0109] When the time focus is T / 2, T / 4, just replace T in the above formula with T / 2, T / 4.
[0110] S572, using the mean to replace the data 0 in the output of the self-attention unit.
[0111] The original Q matrix is compressed and the weights of uncritical time frames in the attention weight matrix are directly ignored. However, considering the weakness and complexity of EEG signals, direct ignoring may lose some signal features and lead to result deviation. Therefore, the mean V matrix is used to replace the 0 vector before output splicing, that is, the relatively unimportant time frame signal is replaced by the global average feature, which can improve the integrity of the feature extraction of EEG signals.
[0112] S58, executing S52-S57 for the self-attention units of the global self-attention module group with different time attention levels;
[0113] S59, after splicing and normalizing the outputs of the self-attention units of different time attention levels, the output of the global self-attention module group is obtained, which is the global temporal correlation information of the EEG signal.
[0114] By compressing the attention weight matrix globally by taking the mean of the key representation vectors and spatiotemporal features within the temporal attention range, the loss of spatiotemporal features can be minimized while reducing the number and complexity of self-attention calculations in existing technologies, thereby focusing more on the more effective features of key time frames. At the same time, through multiple self-attention units with different temporal attention levels, the data at each time point can focus on other data within windows of different time ranges. By fusing the results of multiple self-attention units, the correlation between data within global and local time ranges can be obtained simultaneously, compensating for the fact that the aforementioned compression process only focuses on the "most important information." Compared to the existing self-attention mechanism that directly ignores "secondary information" during weighted average calculations, the integrity of EEG signal feature extraction can be improved.
[0115] The step S53 specifically includes:
[0116] S531, obtain the K matrix of each head of the multi-head self-attention mechanism, and calculate the mean and variance of the eigenvalues of each eigenvector in the same time frame in the K matrix of all heads;
[0117] Each head of the multi-head self-attention mechanism mines the representation information of different time frames in different subspaces. Each element can be associated with other elements in the sequence. In order to compress the attention matrix, the present invention needs to find the key time, that is, only consider the K value with relatively large key information in each self-attention head.
[0118] S532 , sorting the eigenvalues of each eigenvector in the same time frame in the K matrix of all heads in descending order according to the size of the average value, and selecting the first eigenvalue whose variance meets the variance threshold from top to bottom in each time frame to obtain a key characterization vector.
[0119] Specifically, for each feature vector in the same time frame, a K value is calculated under the self-attention mechanism. The average and variance of the K values calculated by all heads for the feature vector are calculated. In each head, only the K value with a larger average and smaller variance is selected as the key representation K value. The key representation K values of multiple heads in each time frame are combined to obtain the key representation vector:
[0120] in represents the key representation vector, K i,T The K value representing the K matrix of the i-th head of the spatiotemporal feature under time frame T has a larger average value and a variance that meets the variance threshold;
[0121] The present invention takes into account the relatively slow feature changes between adjacent time frames of EEG signals, the lack of clear boundaries between adjacent frames, and the difficulty in using attention to calculate the similarity between each pair of time frames. By selecting eigenvalues with relatively large key information in each self-attention head in the original K matrix, the key information is extracted, thereby making the self-attention mechanism pay more attention to the key time frame information.
[0122] Step S54 performs a compression operation on the Q matrix to obtain a compressed Q matrix including:
[0123] S541, based on the key characterization vector and Q matrix according to Calculate the first attention weight matrix;
[0124] S542, sort the first attention weight matrix in descending order according to the size of the weight, select the time frames corresponding to the first p weights in the sort as the key time frames, extract the Q values under the key time frames from the Q matrix to form a compressed Q matrix, where p is a trainable parameter.
[0125] The key representation vector is composed of K values with high weights in multiple self-attention heads. Therefore, time frames with small weights in the first attention weight matrix do not need to be paid much attention, so the original Q matrix is compressed by the weight size. The number of time frames selected can be adjusted during training based on the results of the self-attention module.
[0126] The original Q matrix is compressed through the first attention weight matrix, and only the time frames with large weights are selected as key time frames, while other time frames with low weights are ignored, and the self-attention weights of these time frames are assigned to 0, thereby reducing the number of dot product operations and speeding up the model formation.
[0127] Due to the complex characteristics of the EEG signal itself and the addition of a multi-head self-attention mechanism with multiple temporal attention levels, the feature extraction process is more complicated and the recognition model network has more layers. However, considering the real-time nature of EEG recognition, the present invention uses a self-distillation method to optimize the EEG signal recognition model.
[0128] Through self-distillation, a network with more parameters and more complex structure (teacher model) is refined into a target network with smaller scale and lower latency (student model). This can reduce the complexity of the network structure and achieve the characteristics of lightweight training and high feature transfer efficiency to ensure that the complexity and real-time requirements of EEG signals are met.
[0129] The steps of feature distillation in step S4 are as follows:
[0130] S81, taking each feature extraction layer and the intermediate layer of the self-attention convolutional neural network as a candidate distillation layer; the intermediate layer is a respective attention unit;
[0131] S82, adding a suitable classification structure to each candidate distillation layer, wherein the classification structure is used to output a weak classification result for each candidate distillation layer;
[0132] It should be noted that the classification structure added to each candidate distillation layer can be deleted after the network training is completed, and the final EEG signal recognition model does not increase the response time.
[0133] S83, obtaining the mean precision of the classification structure of each candidate distillation layer, and calculating a distillation correlation value between any candidate distillation layers based on the mean precision; the distillation correlation value is the product of the mean precisions between two candidate distillation layers and the quotient of the square of the number of layers between them;
[0134] By distilling the correlation value, teachers are independently selected for the middle layer of the original network. Compared with the existing technology of manually selecting the last layer as the teacher layer, the knowledge transfer effect is effectively improved, ensuring that the deep layer contains as much knowledge as possible that the shallow layer does not have, ensuring that the two are not too close or too far, and ensuring that the deep knowledge is as suitable as possible for shallow learning.
[0135] S84, assigning a single other candidate distillation layer to each candidate distillation layer based on the distillation correlation value and the preset interval number of teacher and student layers to form multiple pairs of teacher and student groups;
[0136] Among all candidate distillation layers whose interval number of layers is not greater than the preset interval number, another candidate distillation layer whose distillation correlation value is greater than the correlation threshold and is closest to the current candidate distillation layer is assigned to the current candidate distillation layer, where the shallow candidate distillation layer is used as the student layer and the deep candidate distillation layer is used as the teacher layer.
[0137] S85, calculating the feature similarity of two feature vectors of the same EEG signal in each pair of teacher and student groups; the feature similarity uses Euclidean distance to measure the degree of difference between candidate distillation layers of different depths;
[0138] In a neural network, the deeper the network layer, the closer it is to the real features. As the number of layers continues to increase, the features also gradually approach the deepest features. By calculating the similarity of the features extracted by the student layer and the teacher layer for the same signal, the feature extraction of the n+1 layer is made close to that of the n layer.
[0139] S86, calculating the feature similarity for all EEG signals to obtain a similarity matrix for each pair of teacher and student groups, and the feature similarity loss function is to minimize the similarity matrix;
[0140] Use deeper branches to extract deep knowledge, introduce global feature information in deep features into shallow features to balance the feature differences between deep and shallow layers, use deep information to feed back and optimize the shallow layer, refine the similarity of the shallow layer, and learn more accurate information from the original deep layer in the shallow layer. In this way, the shallow convolution can better match the deep layer results when predicting spatial classification results.
[0141] The steps of logic distillation in step S4 are as follows:
[0142] S91, using the classification layer of the self-attention convolutional neural network as a teacher layer, and adding a shallow fully connected classifier after the second layer of the self-attention convolutional neural network feature extraction layer as a student layer;
[0143] It should be understood that, similar to the aforementioned classification structure, the shallow fully connected classifier added after the second layer of the feature extraction layer can be deleted after the network training is completed, which will not increase the complexity of the formed model.
[0144] S92, calculating the outputs of the student layer and the teacher layer through KL divergence to obtain the classification loss function.
[0145] During logistic distillation, the deepest layer classifier serves as the teacher classifier, guiding the learning of the second layer classifier, which is the student classifier for feature extraction. The classification loss can be used to determine the difference between the classification results of the two different layers, which can better guide the shallow network to learn spatial features.
[0146] [Corrected 05.12.2023 according to Rule 91] The present invention also provides a movement disorder rehabilitation training system based on EEG signal recognition, as shown in FIG5 , comprising:
[0147] The EEG signal acquisition module is used to collect EEG signals when patients imagine different types of movements;
[0148] A signal processing module is used to perform preprocessing operations on the collected EEG signals, including a frequency domain filtering unit and a normalization processing unit; the frequency domain filtering unit is used to establish a frequency domain filter to filter out irrelevant noise in the EEG signals in the frequency domain, the normalization processing unit is used to reduce the volatility and non-stationarity of the data, and the spatial domain filtering unit is used to establish a spatial domain filter to maximize the difference in variance values of multiple types of EEG signals;
[0149] Identification network building module, used to build a self-attention convolutional neural network to extract spatiotemporal features;
[0150] A distillation module, configured to optimize the self-attention convolutional neural network model using a self-distillation method;
[0151] a loss function calculation module, configured to determine a first loss function based on the predicted classification and the true classification of the self-attention convolutional neural network, determine a second loss function based on the distance between the spatiotemporal feature and the center of the corresponding spatiotemporal feature, determine a third loss function based on the distillation module, and construct a total model loss function based on a linear combination of the first loss function, the second loss function, and the third loss function;
[0152] A model training module, configured to iteratively train the self-attention convolutional neural network based on the model total loss function to obtain an EEG signal recognition model;
[0153] A classification and recognition module, configured to classify and recognize EEG signals based on the EEG signal recognition model and convert them into control commands recognizable to external devices;
[0154] The motion execution module is used to drive the corresponding exoskeleton equipment to perform rehabilitation training of the corresponding limbs.
[0155] The present invention is not limited to the above-mentioned specific implementation methods. Various changes made by ordinary technicians in this field based on the above-mentioned concept without creative work are all within the scope of protection of the present invention.
Claims
1. A rehabilitation training method for movement disorders based on electroencephalogram (EEG) signal recognition, characterized in that, it includes the following steps: S1. Collect training phase data through offline training for patients. The training phase data is multi-modal EEG signals when movement disorder patients imagine different movement types; S2. Preprocess the multi-modal EEG signals, which sequentially includes establishing a frequency-domain filter based on band-pass filtering, normalization processing, and improving the common spatial pattern (CSP) algorithm to establish a spatial-domain filter; S3. Use the preprocessed EEG signals to establish a self-attention convolutional neural network to extract the spatio-temporal features of the EEG signals, determine the first loss function based on the predicted classification and the true classification of the self-attention convolutional neural network, and determine the second loss function based on the distance between the spatio-temporal features and the corresponding spatio-temporal feature centers; S4. Optimize the self-attention convolutional neural network in a self-distillation manner to obtain the third loss function; S5. Construct a total model loss function based on the linear combination of the first loss function, the second loss function, and the third loss function, and perform iterative training on the self-attention convolutional neural network to obtain an EEG signal recognition model; S6. Collect the EEG signals of patients during online training as rehabilitation phase data, and input the rehabilitation phase data of patients into the EEG signal recognition model established during the patient's most recent offline training to obtain a classification and recognition result; S7. Convert the classification and recognition result into a control command recognizable by an external device, and drive the corresponding limb movement based on the control command to complete the rehabilitation training for the patient's movement disorder.
2. The rehabilitation training method for movement disorders based on EEG signal recognition according to claim 1, characterized in that, in step S2, the band-pass filtering uses a type-II Chebyshev filter to filter the data to remove irrelevant high-frequency and low-frequency noises; the normalization processing uses Z-score for normalization to reduce the volatility and non-stationarity of the data.
3. The rehabilitation training method for movement disorders based on EEG signal recognition according to claim 1, characterized in that, the improvement of the CSP algorithm in step S2 to establish a spatial-domain filter includes: S231. Group the multi-modal EEG signals according to the movement type to form n types of EEG signals, where n is the total number of movement types of the EEG signals; S232, calculate the normalized covariance matrix of each type of EEG signal respectively i represents different motion types; and based on the normalized covariance matrix obtain the mixed spatial covariance matrix R of the multi-modal EEG signals; S233. Perform principal component decomposition on the mixed spatial covariance matrix R of the multi-modal EEG signals; S234, based on the normalized covariance matrices of the various types of EEG signals and obtain a common eigenvector matrix S through the principal component decomposition i , and perform principal component decomposition on the common eigenvector matrix S i in the manner of step S233 to obtain an eigenvalue diagonal matrix V of various electroencephalogram signals i and an eigenvector matrix U i corresponding to the eigenvalue diagonal matrix V i ; S235, for all the common eigenvector matrices S i perform approximate joint diagonalization to obtain the relevant diagonal matrices corresponding to various EEG signals, and calculate the importance of each eigenvalue λ j in the relevant diagonal matrices; the importance is the larger value between the inverse proportional function of the eigenvalue λ j and the eigenvalue λ j ; S236, in each of the relevant diagonal matrices, perform a descending order sorting on the eigenvalue λ according to the importance degree j and record the eigenvalue λ corresponding to the maximum value of the importance degree in each of the relevant diagonal matrices j and the same number of the values thereof; S237, if the values of the eigenvalues λ corresponding to the maximum importance values in each of the relevant diagonal matrices are the same, then the eigenvectors corresponding to the first n eigenvalues λ sorted by importance in each of the relevant diagonal matrices are used for spatial domain filtering; j If the values of the eigenvalues λ corresponding to the maximum importance values in each of the relevant diagonal matrices are the same, then the eigenvectors corresponding to the first n eigenvalues λ sorted by importance in each of the relevant diagonal matrices are used for spatial domain filtering; j If the values of the eigenvalues λ corresponding to the maximum importance values in each of the relevant diagonal matrices are the same, then the eigenvectors corresponding to the first n eigenvalues λ sorted by importance in each of the relevant diagonal matrices are used for spatial domain filtering; If the number of the same values of the eigenvalue λ corresponding to the maximum importance value in each of the relevant diagonal matrices is m, and m < n, then the eigenvectors corresponding to the first m + 1 eigenvalues λ sorted by importance in each of the relevant diagonal matrices are used to establish a spatial domain filter. j j 4. The rehabilitation training method for movement disorders based on EEG signal recognition according to claim 1, characterized in that, the self-attention convolutional neural network in step S3 includes: Feature extraction layer: Perform time-domain and spatial-domain convolutions along the time dimension and the lead channel dimension to extract spatio-temporal features, and extract the global correlation information at the time positions of the EEG signals through the self-attention module; Center loss layer: Define the center loss function based on the distance between the spatio-temporal features and the corresponding spatio-temporal feature centers as the second loss function, which is used to minimize the Euclidean distance between the centers of the class features and the sample features; Classification layer: Use a fully connected layer classifier to predict and classify the spatio-temporal features of the EEG signals extracted by the feature extraction layer, and calculate the cross-entropy loss function between the predicted classification result and the true EEG signal classification label as the first loss function.
5. The method for motor disorder rehabilitation training based on EEG signal recognition according to claim 4, wherein, the feature extraction layer divides the two-dimensional convolution operator into two one-dimensional convolutions to extract time domain features and spatial domain features respectively, specifically including four layers of structures: The first layer uses k convolutional kernels of size (1, 25) with a stride of (1, 1) for convolution operations to extract the time domain features of the EEG signals; The second layer uses a kernel of size (N, 1) with a stride of (1, 1) to learn the interaction between different lead channels and extract spatial domain features; The third layer is a pooling layer in the time dimension with a kernel size of (1, 75) and a stride of (1, 15); The fourth layer uses a self-attention module to obtain the correlation of global time positions.
6. The method for motor disorder rehabilitation training based on EEG signal recognition according to claim 5, wherein, the self-attention module is a global self-attention module group, including: S61, determine the time attention degrees of multiple self-attention units based on the time series length of the spatio-temporal features and the number of self-attention units, and establish multiple self-attention units with different time attention degrees connected in parallel as the global self-attention module group; S62, perform a linear mapping on the spatio-temporal features based on the self-attention units with different time attention degrees to obtain attention matrices corresponding to the time attention degrees: Q matrix, K matrix, V matrix; the matrix dimensions of the attention matrices are the same as those of the spatio-temporal feature matrix; S63, calculate the cumulative eigenvalue of each feature vector in the same time frame in the K matrix in the corresponding self-attention unit to extract the key characterization vector for the K matrix; S64, calculate the first attention weight matrix based on the key characterization vector and the Q matrix, and perform a compression operation on the Q matrix according to the weight magnitudes in the first attention weight matrix to obtain a compressed Q matrix; S65, use zero vectors to complete the dimension of the compressed Q matrix to be the same as that of the K matrix to obtain a key Q matrix; S66, calculate the second attention weight matrix based on the K matrix and the key Q matrix, and weighted sum the V matrix using the second attention weight matrix to obtain the output of the corresponding self-attention unit; S67, use the V matrix of the spatio-temporal features to perform mean filling on the output of the self-attention unit; S68, execute the above steps S62 - S67 for each self-attention unit with different time attention degrees in the global self-attention module group; S69, splice and normalize the outputs of each self-attention unit with different time attention degrees to obtain the output of the global self-attention module group, and the output of the global self-attention module group is the global correlation information in the time position of the EEG signals.
7. The method for motor disorder rehabilitation training based on EEG signal recognition according to claim 5, wherein, The self-attention unit adopts a multi-head scaled dot-product attention mechanism. Each self-attention unit for time attention includes a multi-head self-attention mechanism and a fully-connected network. After both the multi-head self-attention mechanism and the fully-connected network, there are residual connections and normalization modules connected; multiple self-attention units with different time attentions are connected in parallel through a splicing normalization layer.
8. The method for motor disorder rehabilitation training based on electroencephalogram signal recognition according to claim 7, wherein, in step S63, calculating the cumulative eigenvalue of each eigenvector in the same time frame in the K matrix in the corresponding self-attention unit to extract a key representation vector from the K matrix: S631, obtaining the K matrix of each head of the multi-head self-attention mechanism, and calculating the average value and variance of the eigenvalues of each eigenvector in the same time frame in the K matrices of all heads; S632, sorting the eigenvalues of each eigenvector in the same time frame in the K matrices of all heads in descending order according to the average value, and selecting the first eigenvalue whose variance satisfies the variance threshold from top to bottom in each time frame to obtain the key representation vector.
9. The method for motor disorder rehabilitation training based on electroencephalogram signal recognition according to claim 7, wherein, step S64 includes calculating a first attention weight matrix based on the key representation vector and the Q matrix, and performing a compression operation on the Q matrix according to the weight magnitudes in the first attention weight matrix to obtain a compressed Q matrix: S641, calculating a first attention weight matrix based on the key representation vector and the Q matrix; S642, sorting the first attention weight matrix in descending order according to the weight magnitudes, selecting the time frames corresponding to the top p weights as key time frames, and extracting the Q values at the key time frames from the Q matrix to form the compressed Q matrix, where p is a trainable parameter.
10. The method for motor disorder rehabilitation training based on electroencephalogram signal recognition according to claim 7, wherein, step S67 of using the V matrix of the spatio-temporal features to perform mean filling on the output of the self-attention unit includes: S671, calculating the mean value of the V matrix of the spatio-temporal features within the corresponding time attention range; S672, using the mean value to replace the data 0 in the output of the self-attention unit.
11. The method for motor disorder rehabilitation training based on electroencephalogram signal recognition according to claim 1, wherein, step S4 adopts a self-distillation method to optimize the model of the convolutional neural network to obtain a third loss function, including: using the deep network layer of the self-attention convolutional neural network as the teacher model and the shallow network layer as the student model to perform feature distillation and logic distillation on the neural network; the third loss function includes a feature similarity loss function in the feature distillation process and a classification loss function in the logic distillation process.
12. The method for motor disorder rehabilitation training based on electroencephalogram signal recognition according to claim 11, wherein, the feature distillation includes: S1201. Take each layer of the feature extraction layer and the intermediate layer of the self-attention convolutional neural network as candidate distillation layers; the intermediate layer is a self-attention unit. S1202. Add a suitable classification structure to each of the candidate distillation layers, and the classification structure is used to output a weak classification result for each candidate distillation layer. S1203. Obtain the mean accuracy of the classification structure of each candidate distillation layer, and calculate the distillation correlation value between any candidate distillation layers based on the mean accuracy; the distillation correlation value is the quotient of the product of the mean accuracies between two candidate distillation layers and the square of the number of layers between them. S1204. Based on the distillation correlation value and the preset number of layers between the teacher and student layers, assign a single other candidate distillation layer to each of the candidate distillation layers to form multiple pairs of teacher-student groups. S1205. For the same segment of EEG signal, calculate the feature similarity between the two feature vectors of this segment of EEG signal in each pair of teacher-student groups; the feature similarity uses the Euclidean distance and is used to measure the difference degree between candidate distillation layers of different depths. S1206. Calculate the feature similarity for all EEG signals to obtain the similarity matrix for each pair of teacher-student groups, and then the feature similarity loss function is to solve the minimization of the similarity matrix.
13. The method for motion disorder rehabilitation training based on EEG signal recognition according to claim 11, wherein, the logical distillation includes: S1301. Take the classification layer of the self-attention convolutional neural network as the teacher layer, and add a shallow fully-connected classifier after the second layer of the feature extraction layer of the self-attention convolutional neural network as the student layer. S1302. Calculate the output of the student layer and the teacher layer through KL divergence to obtain the classification loss function.
14. The method for motion disorder rehabilitation training based on EEG signal recognition according to claim 4, wherein, the center loss layer defines the center loss function as the second loss function based on the distance between the spatio-temporal feature and the corresponding spatio-temporal feature center, including: Set the initial center point of each EEG signal category as a zero vector or a random vector, and use the spatio-temporal feature output of the fourth layer of the feature extraction layer of the convolutional neural network as the sample feature vector, and calculate the Euclidean distance between each sample feature vector and its corresponding category center point as the second loss function.
Citation Information
Patent Citations
Electroencephalogram emotion classification system based on frame-level feature distillation neural network
CN112989920A
Defective picture recognition system and method based on knowledge distillation, computer and storage medium
CN113592007A
Motor imagery electroencephalogram signal classification method based on iterative learning
CN115034272A
Target detection method based on brain-computer signal fusion
CN116524380A
Auditory attention detection method and system based on electroencephalogram signals
CN116531000A