A deep learning decoding method and system for hand movements driven by neural representation
By obtaining the motion-related cortical potential and event-related synchronized/desynchronized oscillation characteristics of the EEG signal, a time-spectrum-space feature map was constructed, and the attention mechanism and shallow convolutional neural network were used for decoding, the problem of poor bihand motion decoding performance in the existing technology was solved, and more efficient one-hand and two-hand motion decoding was achieved.
Patent Information
- Application Number
- CN202310148622.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-14
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2043-02-14
AI Technical Summary
The existing brain-computer interface based on the ME paradigm faces performance challenges caused by increased brain activity complexity and diversity of motion patterns when decoding hands movements, resulting in poor decoding performance and requiring non-idiopathic hands to remain motionless to reduce motion interference.
A deep learning decoding method for hand movement driven by neural characterization is used to obtain the motion-related cortical potential and event-related synchronized/desynchronized oscillation characteristics of EEG signals, and the time-spectrum-space feature map is constructed, and the attention mechanism and shallow convolutional neural network are used for decoding to improve the decoding performance of hand movement.
The decoding performance of multi-classified one-hand and two-hand movements is improved, the calculation time is shortened, the shallow network structure suitable for EEG data has improved the decoding performance, and the neural characterization of motion-related cortical potential and time-dependent synchronization/desynchronization is integrated to improve the motion decoding performance.
Smart Images

Figure CN116049630B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of neuroscience and deep learning technology, and in particular relates to a deep learning decoding method and system for hand movements driven by neural representation. Background Art
[0002] Brain-computer interfaces (BCIs) have long been a hot topic of interest because they can directly translate brain signals, establishing a communication pathway between the brain and peripheral devices. Brain signals can be recorded invasively or non-invasively. Electroencephalography (EEG) is a major non-invasive recording method. Due to its low cost, portability, and minimal invasiveness, EEG-based BCIs have a wide range of applications and are a key branch of BCI.
[0003] Typical EEG-based brain-computer interfaces include visual brain-computer interfaces, auditory brain-computer interfaces, and motor brain-computer interfaces. Compared to visual and auditory brain-computer interfaces that rely on passively evoked external stimuli, motor brain-computer interfaces can reflect voluntary movement intentions and are therefore more natural and intuitive. Motor brain-computer interfaces are designed to restore or compensate for central nervous system functions. The applications of motor brain interfaces include neurorehabilitation and daily life assistance for patients with motor function loss. For example, combining motor brain-computer interfaces with functional electrical stimulation can help patients actively move their damaged limbs according to their own motor intentions, further promote neural plasticity, and help rebuild neuromuscular circuits.
[0004] Common motor BCI paradigms are motor imagery (MI) and motor execution (ME). MI can be viewed as the mental activity of performing a specific action without any overt motor output, relying more on the repetitive imagining of a motor pattern. Studies have shown that MI of motor actions can produce reproducible brain activation patterns in the supplementary motor area (SMA) and primary motor area (M1). The corresponding neuromodulation is based on sensorimotor rhythms (SMRs), accompanied by event-related and frequency-band-specific decreases and increases in power, termed event-related desynchronization (ERD) and event-related synchronization (ERS). Although MI tasks have shown some value in assisting and restoring impaired motor function in patients, they are limited by the small number of commands decoded, low decoding accuracy, and high mental workload. Furthermore, MI tasks often require imagining movements of different body parts, such as the right hand, left hand, foot, and tongue, resulting in some cases inconsistency between the MI task and the actual output commands. Compared to the MI paradigm, the ME paradigm is more natural because it is goal-oriented and its motor tasks closely match the participant's natural behavior. In addition, the advantage of ME lies in its more obvious brain activity in the temporal and spectral domains, and its better decoding performance for movement, especially multi-classification movement. Unlike other brain-computer interface paradigms (including MI, P300 and SSVEP), the neural activity in the ME paradigm contains both event-related potentials and oscillatory components. Evoked potentials can be captured from the low-frequency band of the EEG signal, called movement-related cortical potentials (MRCP), which are induced during the planning, preparation and execution of movement. The oscillatory component of ME has a similar pattern to MI, manifested as an early ERD of the mu-rhythm before movement initiation and a late ERS of 20-30 Hz after movement execution. Both MRCP and ERS / D oscillations reflect sensorimotor cortical processes, suggesting complementary information related to movement.
[0005] Many studies have explored the use of MRCP or ERD to decode hand movement intentions based on the ME paradigm. Decoding includes movement onset detection, movement direction classification, torque level, speed and movement type, and continuous movement reconstruction. These studies have demonstrated the feasibility of decoding upper limb or hand movement intentions from EEG signals. However, existing studies on hand movement decoding based on the ME paradigm are mostly limited to decoding single-hand movements, requiring the non-dominant hand to remain motionless during the experiment. This setting is intended to reduce motion interference in decoding the dominant hand's movements.
[0006] However, in daily life, coordinating both hands to complete a task is very common, and in neurorehabilitation, bilateral training can promote recovery after stroke. Therefore, decoding bimanual movements is valuable. However, it faces at least two problems. First, brain activity during bimanual movements is more complex than during unimanual movements. Second, bimanual movement patterns are more numerous than unimanual movement patterns. These two issues pose severe challenges to the decoding performance of bimanual movements. Therefore, the problem of bimanual movement decoding urgently needs further exploration. Summary of the Invention
[0007] The purpose of the present invention is to propose a deep learning decoding method and system for hand movements driven by neural representation, so as to improve the decoding performance of multi-classification single-hand and two-hand movements.
[0008] On one hand, to achieve the above-mentioned objectives, the present invention provides a deep learning decoding method for hand movements driven by neural representation, comprising:
[0009] Acquiring an electroencephalogram (EEG) signal, performing feature extraction on the EEG signal, and acquiring movement-related cortical potential features and event-related synchronization / desynchronization oscillation features;
[0010] Acquiring a time-spectrum-space feature map based on the movement-related cortical potential characteristics and the event-related synchronization / desynchronization oscillation characteristics;
[0011] The time-spectrum-space feature map is decoded for single and double-handed movement intentions, a decoding result is obtained, and deep learning decoding of hand movements driven by neural representation is completed.
[0012] Optionally, performing feature extraction on the electroencephalogram signal to obtain movement-related cortical potential features and event-related synchronization / desynchronization oscillation features includes:
[0013] Filtering the electroencephalogram signal using discrete wavelet transform to obtain several frequency bands;
[0014] Obtaining an optimal movement-related cortical potential frequency band based on the mutual information and the plurality of frequency bands;
[0015] Performing continuous wavelet transform on the optimal movement-related cortical potential frequency band to obtain movement-related cortical potential features;
[0016] Filtering the electroencephalogram signal using continuous wavelet transform to obtain spectral power;
[0017] acquiring an optimal event-related synchronization / desynchronization frequency band based on the mutual information and the spectrum power;
[0018] Based on the optimal event-related synchronization / desynchronization frequency band, an event-related synchronization / desynchronization oscillation feature is obtained.
[0019] Optionally, obtaining a time-spectrum-space feature map based on the movement-related cortical potential feature and the event-related synchronization / desynchronization oscillation feature includes:
[0020] fusing the movement-related cortical potential feature and the event-related synchronization / desynchronization oscillation feature to obtain a time-spectrum feature;
[0021] The time-spectrum features are channel-weighted based on the attention mechanism to obtain a time-spectrum-space feature map.
[0022] Optionally, performing channel weighting on the time-spectrum feature based on an attention mechanism to obtain a time-spectrum-space feature map includes:
[0023] Constructing a channel attention model, inputting the time-spectrum features into the channel attention model, and obtaining a channel score;
[0024] Performing a scaled dot product on the channel scores and weighting them based on a normalized exponential function to obtain an attention score;
[0025] Performing a dot product process on the attention score and the time-spectrum feature to obtain channel weighted data;
[0026] The channel weighted data are superimposed and projected to obtain a time-spectrum-space feature map.
[0027] Optionally, performing single-handed and double-handed motion intention decoding on the time-spectrum-space feature map to obtain a decoding result includes:
[0028] Performing a first convolution process on the time-spectrum-space feature map to obtain a plurality of convolved feature maps;
[0029] Performing a second convolution process on the convolved feature maps to obtain a plurality of depth-convolved feature maps;
[0030] Performing a pooling operation on the feature maps after the depth convolution to obtain a plurality of one-dimensional feature maps;
[0031] The one-dimensional feature maps are fully connected in series to obtain decoding results of single-handed and double-handed movements.
[0032] On the other hand, to achieve the above-mentioned object, the present invention also provides a deep learning decoding system for hand movements driven by neural representation, comprising a feature representation module, an attention-based channel weighting module, and a shallow convolutional neural network module, wherein the feature representation module, the attention-based channel weighting module, and the shallow convolutional neural network module are connected in sequence;
[0033] The feature representation module is used to obtain the movement-related cortical potential features and the event-related synchronization / desynchronization oscillation features;
[0034] The attention-based channel weighting module is used to obtain the time-spectrum-space feature map;
[0035] The shallow convolutional neural network module is used to extract features from the time-spectrum-space feature map.
[0036] Optionally, the attention-based channel weighting module includes a query and key submodule, and the query and key submodule is used to weight the channel score.
[0037] The present invention has the following beneficial effects:
[0038] The present invention discloses a deep learning decoding method and system for hand movements driven by neural representations, designs a deep learning model driven by neurophysiological characteristics, and improves the decoding performance of multi-classification single-hand and two-hand movements; the EEG data in the present invention is more suitable for shallower network structures, which can improve decoding performance and shorten calculation time; the present invention is the first to attempt to improve motion decoding performance by fusing motion-related cortical potentials with time-related synchronized / desynchronized neural representations in deep learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0040] Figure 1 This is a schematic diagram of a deep learning decoding method for hand movements driven by neural representation according to an embodiment of the present invention;
[0041] Figure 2 Schematic diagram of the channel attention module proposed in an embodiment of the present invention. DETAILED DESCRIPTION
[0042] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0043] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0044] like Figure 1 As shown in Figure 1, this embodiment provides a deep learning decoding system for hand movements driven by neural representations, including a feature representation module, an attention-based channel weighting module, and a shallow convolutional neural network module. The system uses EEG signals as input and outputs decoding results for single and double-handed movements.
[0045] 1. Neural representation modules of MRCP and ERS / D activities
[0046] The original input EEG signal is expressed as Where i is the ith trial, x is the input EEG sample, n is the total number of trials, C is the number of electrodes, T is the number of time sampling points, R is the real number space, and the corresponding class label is expressed as represents six categories of one-handed and two-handed movement combinations, where y is a subset of the class labels.
[0047] In order to obtain the MRCPs features, the original EEG signal X of the training dataset is first transformed using discrete wavelet transform (DWT). train Decompose into different frequency bands, and then use mutual information (M-Info) to select an optimal frequency band. Wavelet-based methods can well characterize signals in both the time domain and the spectral domain. In wavelet transform, the inner product of the original EEG signal and the basis wavelet function is first discretely calculated, and then the basis to be analyzed is found in a specific band. Finally, a series of signal bases are reconstructed and weighted to obtain the filtered signal. The continuous wavelet transform (CWT) is defined as:
[0048]
[0049] in Corresponding to the EEG signal of electrode c at time t in experiment i, a and b are the scaling parameter and translation parameter of the wavelet basis function, respectively. Ψ is the wavelet function, * represents the complex conjugate, a, b∈R, a≠0. The wavelet transform is performed by discretizing the scale and translation parameters a and b. The present invention uses a binary scale and transform:
[0050] a j =2 j , b j,k =k2 j , k, j∈Z (2)
[0051] Where j and k are discrete sampling points.
[0052] In this case, Ψ j,k (t) = 2 -j / 2 Ψ(2 -j tk),Ψ j,k The set of (t) in the square integrable space L 2 (R). Through multi-resolution decomposition, L can be decomposed 2 (R) into multiple subspaces W j , each subspace can be obtained by the expansion and translation of a single basis function, and then a sequence of closed subspaces V can be found j . Subspace V j All signals included are V j+1 Plus additional high-resolution signals:
[0053] V j =V j+1 ⊕W j+1 (3)
[0054] Where ⊕ represents V j+1 and W j+1 are all subspaces V j and are orthogonal to each other. j By accumulating all the resolution signals in , the original signal f(t) can be obtained as:
[0055]
[0056] Where φ(·) is the scaling function, c k is the scale factor, d j,k is the wavelet coefficient, L is the total number of decomposition layers. Wavelet function Ψ j,k (t) corresponds to a high-pass filter, preserving signal details, while the wavelet function φ(·) corresponds to a low-pass filter, maintaining the signal approximation. By convolving the original signal with a high-pass filter and a low-pass filter, detailed coefficients and approximate coefficients can be obtained. The filtered signal can be reconstructed within a specific frequency band using these coefficients. In this invention, the wavelet family "sym5" is used, and the maximum decomposition level is set to 6.
[0057] After filtering the EEG signal into different frequency bands, the M-Info method is applied to the training data set to select an optimal frequency band. From the perspective of information theory, M-Info can measure the EEG sample X train With category label Y train The statistical dependence between them can be defined as:
[0058]
[0059] In the formula p(x, y) is the joint probability density of continuous random variables, p(x) and p(y) are the marginal probability densities. Shannon entropy H(X train ) and H(Y train ) can measure the information obtained from the variable as follows:
[0060]
[0061]
[0062] The joint entropy is defined as:
[0063]
[0064] Given Y train , X train The conditional entropy of is defined as:
[0065]
[0066] Note that H(X train , Y train )=H(X train |Y train )+H(Y train ). Entropy H(Y train ) measures the Y train The uncertainty of H(X train |Y train ) measures the given Y train Under the condition X train Therefore, in formula (5), I(X train , Y train ) can be equal to:
[0067]
[0068] The present invention calculates the sum of the mutual information of all variables in each frequency band, keeps the frequency band with the maximum mutual information, and then performs DWT processing on the frequency band to obtain MRCPs features.
[0069] To obtain the ERS / D oscillation characteristics, we first perform CWT to obtain the spectral power, and then apply M-Info to select an optimal frequency band. The definition of CWT is shown in Equation (1). CWT is achieved by performing a large number of wavelet transforms on each possible scale and translation. For each EEG electrode in each trial, the calculated spectral power is defined as Power(c, t, f), where c is the electrode number, t is the time sampling point, and f is the frequency sampling point. Then, we calculate the power sum of each frequency band as:
[0070] Sum_Power(c, t) = ∑f Power(c,t,f) (11)
[0071] Then, we perform feature selection based on mutual information on the training dataset to select an optimal ERS / D frequency band. We normalize and concatenate the MRCPs and ERS / D oscillation features to form a feature array of size n×2×C×T.
[0072] 2. Channel Space Weight Calculation Module Based on Attention Mechanism
[0073] Different brain regions have different contributions to different brain activities, so an attention-based spatial channel weighting algorithm is adopted. This paper further improves the attention mechanism based on scaled dot product attention. Specifically, the query (Q) and key (K) modules are used to weight the scores of the channels by dot product, and each module consists of a linear layer, a LayerNorm layer, and a Dropout layer; then, the channel attention score is calculated by Scaling and weighting by using the Softmax function can be expressed as:
[0074]
[0075] Then, the dot product of the scored attention and the input data is performed; finally, the channel weighted data is superimposed and projected through the linear layer, LayerNorm layer, and Dropout layer. The size of the linear layers is C×C, and the dropout rate is set to 0.3. The principle program of the channel attention module is as follows Figure 2 shown.
[0076] 3. Shallow Convolutional Neural Network Module
[0077] In this module, a shallow convolutional neural network architecture is constructed to process 3-D temporal-spectral-spatial feature maps. For classification problems involving small data sets, such as the EEG data in this paper, a shallower network architecture is more suitable, as it can improve decoding performance and reduce computation time. Therefore, a shallow convolutional neural network with only two convolutional neural network layers and a small kernel size is constructed. A detailed description of the network architecture and parameters is shown in Table 1. When the feature map is input into the network, a convolutional layer with a temporal kernel size of 5 and 4 convolutional filters is first used to receive temporal information. A deep convolutional layer with a spatial kernel size of 20 and 8 convolutional filters is then used to centralize the electrodes and aggregate individual feature maps across feature channels. An average pooling layer (kernel size of 16) is then applied to reduce the temporal feature dimensionality. It is noteworthy that in the proposed network, larger pooling kernels are used, rather than larger convolution kernels, to reduce the temporal dimension. This is because larger convolution kernels would destroy the low-frequency information of the EEG signal, which primarily encodes the actions performed. For all convolutional layers, batch normalization is added. After the depthwise convolutional layer, the exponential linear unit (ELU) is used as the activation function, and a dropout layer with a probability of 50% is used to prevent overfitting. Finally, the output feature map is flattened to one dimension and connected to the fully connected layer.
[0078] Table 1
[0079]
[0080] The above description is merely a preferred embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A deep learning decoding method for hand movements driven by neural representation, characterized in that include: Acquiring an electroencephalogram signal, performing feature extraction on the electroencephalogram signal, and acquiring movement-related cortical potential features and event-related synchronization / desynchronization oscillation features, including: Filtering the electroencephalogram signal using discrete wavelet transform to obtain several frequency bands; Obtaining an optimal movement-related cortical potential frequency band based on the mutual information and the plurality of frequency bands; Performing continuous wavelet transform on the optimal movement-related cortical potential frequency band to obtain movement-related cortical potential features; Filtering the electroencephalogram signal using continuous wavelet transform to obtain spectral power; acquiring an optimal event-related synchronization / desynchronization frequency band based on the mutual information and the spectrum power; Acquiring event-related synchronization / desynchronization oscillation characteristics based on the optimal event-related synchronization / desynchronization frequency band; Acquiring a time-spectrum-space feature map based on the movement-related cortical potential characteristics and the event-related synchronization / desynchronization oscillation characteristics; The time-spectrum-space feature map is decoded for single and double-handed movement intentions, a decoding result is obtained, and deep learning decoding of hand movements driven by neural representation is completed.
2. The neural representation driven deep learning decoding method for hand movements according to claim 1, wherein: Acquiring a time-spectrum-space feature map based on the movement-related cortical potential feature and the event-related synchronization / desynchronization oscillation feature includes: fusing the movement-related cortical potential feature and the event-related synchronization / desynchronization oscillation feature to obtain a time-spectrum feature; The time-spectrum features are channel-weighted based on the attention mechanism to obtain a time-spectrum-space feature map.
3. The neural representation driven deep learning decoding method for hand movements according to claim 2, wherein: Performing channel weighting on the time-spectrum features based on the attention mechanism to obtain a time-spectrum-space feature map includes: Constructing a channel attention model, inputting the time-spectrum features into the channel attention model, and obtaining a channel score; Performing a scaled dot product on the channel scores and weighting them based on a normalized exponential function to obtain an attention score; Performing a dot product process on the attention score and the time-spectrum feature to obtain channel weighted data; The channel weighted data are superimposed and projected to obtain a time-spectrum-space feature map.
4. The neural representation driven deep learning decoding method for hand movements according to claim 1, wherein: Decoding the single-handed and double-handed motion intentions on the time-spectrum-space feature map to obtain a decoding result includes: Performing a first convolution process on the time-spectrum-space feature map to obtain a plurality of convolved feature maps; Performing a second convolution process on the convolved feature maps to obtain a plurality of depth-convolved feature maps; Performing a pooling operation on the feature maps after the depth convolution to obtain a plurality of one-dimensional feature maps; The one-dimensional feature maps are fully connected in series to obtain decoding results of single-handed and double-handed movements.
5. A system for the deep learning decoding method of hand movements driven by neural representation according to any one of claims 1 to 4, characterized in that: It includes a feature representation module, an attention-based channel weighting module and a shallow convolutional neural network module, wherein the feature representation module, the attention-based channel weighting module and the shallow convolutional neural network module are connected in sequence; The feature representation module is used to obtain the movement-related cortical potential features and the event-related synchronization / desynchronization oscillation features; The attention-based channel weighting module is used to obtain the time-spectrum-space feature map; The shallow convolutional neural network module is used to extract features from the time-spectrum-space feature map.
6. The neural representation driven hand movement deep learning decoding system of claim 5, wherein: The attention-based channel weighting module includes query and key submodules for weighting the channel scores.
Citation Information
Patent Citations
Limb movement imagination brain-computer interaction method and system
CN112306244A
Lightweight and rapid motor imagery electroencephalogram signal decoding method
CN115316955A