Dynamic spatio-temporal feature enhancement network model for motor imagery classification

Through dynamic spatiotemporal features enhancement network model, combining multi-scale temporal convolution and grouping spatial convolution, the spatiotemporal characteristics of MI-EEG signals are extracted, which solves the problems of low signal-to-noise ratio and high intra-class variability, and significantly improves the accuracy of motion imagination classification.

CN120196926AActive Publication Date: 2025-06-24SHANGHAI SHAONAO SENSING TECH CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510259763.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-24
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

MI-EEG signals have non-stationarity, low signal-to-noise ratio and high intra-class variability, which leads to difficulty in extracting spatiotemporal features, poor generalization of traditional machine learning, and it is difficult for single-scale convolutional networks to fully mine EEG's spatiotemporal information.

Method used

A dynamic spatiotemporal feature enhancement network model is proposed, combining the dynamic spatiotemporal feature enhancement module and the spatiotemporal convolution module to extract the α and β frequency band features through multi-scale temporal convolution layer and grouped spatial convolution layer, and prevent overfitting through weight constraints.

Benefits of technology

On the BCI-IV-2a, OpenBMI, CASIA and stroke patients data sets, the average accuracy of DSTA-Net was improved by 6.29%, 3.05%, 5.26% and 2.25% compared with ShallowConvNet, respectively, significantly improving the classification performance and generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196926A_ABST
    Figure CN120196926A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic spatio-temporal feature enhancement network model for motor imagery classification, and relates to the technical field of motor imagines.The dynamic spatio-temporal feature enhancement network model comprises a dynamic spatio-temporal feature enhancement module and a spatio-temporal convolution module, the dynamic spatio-temporal feature enhancement module comprises a multi-scale time convolution layer and is used for extracting alpha and beta frequency band features in an MI-related electroencephalogram (EEG) signal; an original EEG signal is reserved as a baseline feature layer; the dynamic spatio-temporal feature enhancement module further comprises a grouping space convolution layer which is used for extracting multi-level space features and preventing overfitting through weight constraint; and the space-time convolution module is used for further extracting space-time features and classifying the space-time features. According to the technical scheme, a DSTA module and a space-time convolution (STC) module are combined, and in 10-fold cross validation, the average accuracy rates of DSTA-Net on BCI-IV-2a, OpenBMI, CASIA and apoplexy patient data sets are improved by 6.29% (p < 0.01), 3.05% (p < 0.01), 5.26% (p < 0.01) and 2.25% respectively compared with those of ShallowConvNet, and the average accuracy rates of DSTA-Net on BCI-IV-2a, OpenBMI, CASIA and apoplexy patient data sets are improved by 3.05% (p < 0.01), 3.05% (p < 0.01), 5.26% (p < 0.01) and 2.25% respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of motor imagery, and particularly to a dynamic spatio-temporal feature enhancement network model for motor imagery classification. Background Art

[0002] Brain-computer interface (BCI) acquires central nervous system signals through sensors when a user performs a mental task or receives a stimulus, and encodes these signals into instructions to interact with a computer, thereby realizing the restoration, enhancement, or replacement of nerve functions. Among various physiological signal acquisition technologies, due to low cost, rapid response, and easy deployment, non-invasive electroencephalogram (EEG) has been widely adopted in BCI applications, such as neurofeedback training, wheelchair control, drone navigation, and virtual reality. Compared with steady-state visual evoked potential (SSVEP) and event-related potential (ERP), motor imagery (MI) has attracted much attention due to its potential in enhancing neural plasticity and promoting rehabilitation.

[0003] MI-EEG signals have non-stationarity, low signal-to-noise ratio, and high intra-class variability, resulting in difficulties in spatio-temporal feature extraction. Traditional machine learning relies on handcrafted features (such as band power, common spatial pattern), but requires professional knowledge and has poor generalization. Deep learning breaks through this bottleneck through automatic feature learning: for example, DeepConvNet uses deep convolution to extract complex spatio-temporal features, ShallowConvNet focuses on shallow band power features, and EEGNet combines spatio-frequency convolution, showing high efficiency in small-sample scenarios.

[0004] Single-scale convolutional networks are difficult to fully exploit the spatio-temporal information of EEG. Multi-scale methods enhance feature representation through convolutional kernels of different sizes, but too many convolutional kernels will lead to increased computational complexity, information redundancy, and optimization difficulties, while too few convolutional kernels may limit feature representation, resulting in loss of key information and degradation of generalization ability. Summary of the Invention

[0005] The technical solution of the present invention to solve the above technical problems is to provide a dynamic spatio-temporal feature enhancement network model for motor imagery classification, including a dynamic spatio-temporal feature enhancement module and a spatio-temporal convolution module. Among them, the dynamic spatio-temporal feature enhancement module includes a multi-scale temporal convolution layer for extracting α and β band features in MI-related electroencephalogram (EEG) signals and retaining the original EEG signal as a baseline feature layer;

[0006] The dynamic spatio-temporal feature enhancement module further includes a grouped spatial convolution layer for extracting multi-level spatial features and preventing overfitting through weight constraints;

[0007] The spatio-temporal convolution module is used to further extract spatio-temporal features and perform classification.

[0008] Furthermore, the multi-scale temporal convolutional layer includes one-dimensional temporal convolutional kernels of multiple levels, where the temporal length of the convolutional kernel in the i-th layer is defined as: It is used to dynamically capture spatio-temporal features in different frequency ranges of EEG signals;

[0009] Through and Determine the convolutional kernel size to match the characteristic ranges of the alpha and beta frequency bands in the motor imagery (MI) task;

[0010] Introduce an original signal compensation layer in the multi-scale temporal convolutional layer to reduce the distortion of the feature matrix caused by multi-scale convolution;

[0011] Concatenate the outputs of the original signal compensation layer, the alpha-band feature layer, and the beta-band feature layer to generate an enhanced spatio-temporal feature representation.

[0012] Furthermore, the constrained grouped spatial convolutional layer includes three groups of spatial convolutional modules. Each group of spatial convolutional modules uses grouped spatial convolution operations with a kernel size of (Nc, 1), where Nc represents the number of EEG signal channels;

[0013] Each group of spatial convolutional modules generates 10 convolutional kernels. Concatenate the feature maps output by the three groups along the convolutional kernel dimension to obtain a feature matrix Xspatial with a shape of Bs×Nf×1×T, where Bs represents the batch size taken as 16, Nf = 30 represents the total number of convolutional kernels, and T represents the time dimension;

[0014] Apply a maximum norm constraint to the weight vector of each convolutional kernel, and re-normalize it through the L2 norm so that the L2 norm value of the weight vector is less than 2;

[0015] Perform BatchNorm2d normalization processing and Swish activation function transformation on each feature channel of the feature matrix Xspatial in sequence;

[0016] Reshape the activated feature matrix into a three-dimensional tensor structure, keeping its channel dimension consistent with the original EEG signal spatial dimension to form a feature matrix integrating spatio-temporal features.

[0017] Furthermore, the spatio-temporal convolutional module includes

[0018] Temporal convolutional layer: Use 60 filters to perform convolution on the input EEG signal in the time dimension, with a convolutional kernel size of k = (1, 25) to extract time-domain features;

[0019] Spatial convolutional layer: Perform convolution on the output of the temporal convolutional layer in the spatial dimension, with a convolutional kernel size of k = (60, 1) to expand the number of feature map channels to 120;

[0020] Batch Normalization layer: Performs BatchNorm2d normalization on the output of the spatial convolutional layer to reduce internal covariate shift;

[0021] Square Enhancement layer: Performs element-wise squaring on the normalized features through the custom SquareLayer module to enhance feature separability;

[0022] Average Pooling layer: Performs downsampling on the features in the time dimension using AvgPool2d, with a pooling kernel size of p=(1,100) and a stride of s=(1,10);

[0023] Logarithmic Transformation layer: Performs logarithmic transformation on the pooled features through the LogLayer module to optimize the feature representation;

[0024] Random Dropout layer: Applies the Dropout operation to the features with a probability p = 0.5, randomly deactivating some neurons to prevent overfitting;

[0025] Classification Convolutional layer: Maps 120 input channels to output classes through temporal convolution, with a convolutional kernel size of (1,88), and uses the LogSoftmax activation function to output classification probabilities;

[0026] Loss calculation: Optimizes the model parameters based on the negative log-likelihood loss (NLLLoss).

[0027] Compared with the prior art, the technical solution of the present invention has the following technical effects:

[0028] The proposed Dynamic Spatiotemporal Feature Enhancement Network (DSTA-Net) in this application combines the DSTA and Spatiotemporal Convolution (STC) modules. In the DSTA module, for the α and β frequency bands of MI neurophysiological features, multi-scale temporal convolutional kernels are designed, and the original EEG is used as the baseline feature layer to retain the original information. Grouped spatial convolution extracts multi-level spatial features and combines weight constraints to prevent overfitting. The spatial convolutional kernel maps the EEG channel information to a new spatial domain and further feature extraction is achieved through dimensional transformation. The STC module performs feature extraction and classification.

[0029] DSTA-Net was evaluated on three public datasets and applied to a self-collected dataset of stroke patients. In 10-fold cross-validation, DSTA-Net achieved an average accuracy improvement of 6.29% (p < 0.01), 3.05% (p < 0.01), 5.26% (p < 0.01), and 2.25% compared to ShallowConvNet on the BCI-IV-2a, OpenBMI, CASIA, and stroke patient datasets, respectively. In holdout validation, DSTA-Net achieved an average accuracy improvement of 3.99% (p < 0.01) and 4.2% (p < 0.01) compared to ShallowConvNet on the OpenBMI and CASIA datasets, respectively. Brief Description of the Drawings

[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.

[0031] Figure 1 It is a schematic structural diagram of the dynamic spatio-temporal feature enhancement network model for motor imagery classification according to the present invention;

[0032] Figure 2 It is an experimental paradigm and EEG preprocessing diagram of the stroke dataset according to the present invention;

[0033] Figure 3 It is an individual classification accuracy diagram of the OpenBMI dataset (HO) according to the present invention;

[0034] Figure 4 It is a confusion matrix diagram of different algorithms under HO analysis for the OpenBMI and CASIA datasets according to the present invention;

[0035] Figure 5 It is a DeepLift visualization diagram of different algorithms for Subject 1 in the BCI-IV-2a dataset according to the present invention (with the left hand as the target label and the right hand as the reference label);

[0036] Figure 6 It is a t-SNE visualization diagram of Subject 1 in the BCI-IV 2a dataset under different algorithms according to the present invention (labels 0, 1, 2, and 3 represent the left hand, right hand, foot, and tongue in the imagined movement, respectively);

[0037] Figure 7Ablation experiment results graph of four datasets in the cross-validation (CV) mode of the present invention (based on paired t-tests, statistical significance is marked with **(p<0.01));

[0038] Figure 8 DeepLift visualization graph of the present invention using DSTA-Net for Subject 1 (accuracy 96.33%) and Subject 3 (accuracy 42.67%) of the stroke dataset (with left and right hands as target labels and the resting task as the reference label):

[0039] Figure 9 CSP topographic maps of the left and right hand MI tasks in the α and β frequency bands for Patient 1 and Patient 3 of the present invention. Detailed implementation manners

[0040] The present invention proposes a dynamic spatio-temporal feature enhancement network model for motor imagery classification, aiming to propose a dynamic spatio-temporal feature enhancement network model that combines DSTA and spatio-temporal convolution (STC) modules.

[0041] Next, the specific structure of the dynamic spatio-temporal feature enhancement network model for motor imagery classification proposed by the present invention will be described in specific embodiments:

[0042] Embodiment 1:

[0043] A dynamic spatio-temporal feature enhancement network model for motor imagery classification, including a dynamic spatio-temporal feature enhancement module and a spatio-temporal convolution module, wherein:

[0044] The dynamic spatio-temporal feature enhancement module includes a multi-scale temporal convolution layer for extracting α and β frequency band features in MI-related electroencephalogram (EEG) signals and retaining the original EEG signal as a baseline feature layer;

[0045] The dynamic spatio-temporal feature enhancement module further includes a grouped spatial convolution layer for extracting multi-level spatial features and preventing overfitting through weight constraints;

[0046] The spatio-temporal convolution module is used to further extract spatio-temporal features and perform classification.

[0047] Furthermore, the multi-scale temporal convolution layer includes one-dimensional temporal convolution kernels of multiple levels, where the temporal length of the i-th layer convolution kernel is defined as: For dynamically capturing spatio-temporal features in different frequency ranges of EEG signals;

[0048] Through and Determine the convolution kernel size to match the feature ranges of the α and β frequency bands in the motor imagery (MI) task;

[0049] Introduce an original signal compensation layer in the multi-scale temporal convolutional layer to reduce the distortion of the feature matrix caused by multi-scale convolution;

[0050] Concatenate the outputs of the original signal compensation layer, the α-band feature layer, and the β-band feature layer to generate an enhanced spatio-temporal feature representation.

[0051] Furthermore, the constrained grouped spatial convolutional layer includes three groups of spatial convolutional modules. Each group of spatial convolutional modules uses grouped spatial convolution operations with a kernel size of (Nc, 1), where Nc represents the number of EEG signal channels;

[0052] Each group of spatial convolutional modules generates 10 convolutional kernels. Concatenate the output feature maps of the three groups along the convolutional kernel dimension to obtain a feature matrix Xspatial with a shape of Bs×Nf×1×T, where Bs represents the batch size taken as 16, Nf = 30 represents the total number of convolutional kernels, and T represents the time dimension;

[0053] Apply a maximum norm constraint to the weight vector of each convolutional kernel and perform re-normalization through the L2 norm to make the L2 norm value of the weight vector less than 2;

[0054] Perform BatchNorm2d normalization processing and Swish activation function transformation on each feature channel of the feature matrix Xspatial in sequence;

[0055] Reshape the activated feature matrix into a three-dimensional tensor structure, keeping its channel dimension consistent with the original EEG signal spatial dimension to form a feature matrix integrating spatio-temporal features.

[0056] Furthermore, the spatio-temporal convolutional module includes

[0057] Temporal convolutional layer: Use 60 filters to perform convolution on the input EEG signal in the time dimension, with a convolutional kernel size of k = (1, 25) to extract temporal domain features;

[0058] Spatial convolutional layer: Perform convolution on the output of the temporal convolutional layer in the spatial dimension, with a convolutional kernel size of k = (60, 1) to expand the number of feature map channels to 120;

[0059] Batch normalization layer: Perform BatchNorm2d normalization processing on the output of the spatial convolutional layer to reduce internal covariate shift;

[0060] Square enhancement layer: Perform element-wise squaring operation on the normalized features through a custom SquareLayer module to enhance feature separability;

[0061] Average pooling layer: Use AvgPool2d to downsample the features in the time dimension, with the pooling kernel size p = (1, 100) and the stride s = (1, 10);

[0062] Logarithmic transformation layer: Perform logarithmic transformation on the pooled features through the LogLayer module to optimize the feature representation;

[0063] Random dropout layer: Apply the Dropout operation to the features with a probability p = 0.5, randomly deactivating some neurons to prevent overfitting;

[0064] Classification convolutional layer: Map 120 input channels to output classes through temporal convolution, with the convolutional kernel size (1, 88), and use the LogSoftmax activation function to output classification probabilities;

[0065] Loss calculation: Optimize the model parameters based on the negative log-likelihood loss (NLLLoss).

[0066] Example 2:

[0067] A dynamic spatio-temporal feature enhancement network model for motor imagery classification, including a dynamic spatio-temporal feature enhancement module and a spatio-temporal convolution (STC) module, where:

[0068] The dynamic spatio-temporal feature enhancement module includes a multi-scale temporal convolutional layer for extracting α and β band features in MI-related electroencephalogram (EEG) signals and retaining the original EEG signal as a baseline feature layer;

[0069] The dynamic spatio-temporal feature enhancement module also includes a grouped spatial convolutional layer for extracting multi-level spatial features and preventing overfitting through weight constraints;

[0070] The spatio-temporal convolution module is used to further extract spatio-temporal features and perform classification.

[0071] 1. The architecture of DSTA-Net is shown in Table 1:

[0072] Table 1 shows the architecture of DSTA-Net:

[0073]

[0074]

[0075] 1.1 Dynamic spatio-temporal feature enhancement module:

[0076] The input of the model determines the upper limit of feature extraction and classification performance. Therefore, enhancing the representation of the original EEG data helps improve the model performance. MI contains rich spatio-temporal feature information. · To propose the DSTA module to enhance the spatio-temporal feature representation of MI. In the DSTA module, the dynamic multi-scale temporal convolutional layer aims to improve the EEG signal representation while retaining as much original information as possible. Given that the EEG signal has a high temporal resolution, the first layer of DSTA-Net consists of a dynamic multi-scale temporal convolutional layer along the temporal dimension, which is composed of multi-scale one-dimensional temporal convolutional kernels. The length of each convolutional kernel is set to a specific ratio of the EEG sampling frequency fs, defined as follows: γ i ∈R, where i represents the level in the multi-scale temporal convolution. If the dynamic temporal layer contains N layers, the value range of i is from 1 to N. Therefore, the temporal kernel size of the i-th layer is defined as:

[0077] From the frequency perspective, different temporal kernel sizes can enrich the model's learning representation of the dynamic frequencies in EEG. For example, in the emotion decoding task, Tsception uses dynamic temporal kernels with lengths of [1 / 2, 1 / 4, 1 / 8]×fs and can capture dynamic frequency features above [2, 4, 8] Hz. It is found that MI activities are mainly concentrated in the α and β frequency ranges. To capture the features of these frequency bands, the convolutional kernel size is determined through formula derivation. The sampling frequency is set as:

[0078] Substituting Equation 1 into Equation 2 gives:

[0079] Solving the inequalities for the α and β frequency bands respectively:

[0080] The results show that:

[0081] After comparing the candidate ratios [1 / 2, 1 / 4, 1 / 8, 1 / 16], it is found that matches well with the α frequency band, and matches with the β frequency band. Therefore, [1 / 8, 1 / 16] is selected as the small-scale temporal kernels. In addition, to reduce the distortion of the feature matrix relative to the original EEG signal after multi-scale convolution, an original signal layer is added on top of the multi-scale convolution to compensate for the information loss introduced during the convolution process.

[0082] From the time perspective, the multi-scale temporal kernels can capture short-term and long-term temporal patterns. Let X = [X0, X1,..., X n , where X n ∈R c×t, X represents the EEG signal, n represents the number of trials, c represents the number of EEG channels, and t represents the sampling points of time information. The output of each layer of the multi-scale temporal convolution is defined as:

[0083]

[0084] p i = int(γ i × fs / 2)

[0085] where p i represents padding. After the convolution operation of the dynamic multi-scale temporal convolutional neural network layer, the final output is concatenated along the convolution kernel dimension. The formula is as follows:

[0086]

[0087] where and represent the original signal layer, the α-band feature layer, and the β-band feature layer, respectively.

[0088] The constrained grouped spatial convolution layer combines spatial convolution, batch normalization, and the Swish activation function. The grouped spatial convolution is performed using a kernel size of (Nc, 1), where Nc represents the number of EEG channels. Considering that and each represent different feature layers, three groups of spatial convolutions are designed. Each group generates 10 convolution kernels, and then they are concatenated along the convolution kernel dimension. The shape of the resulting feature matrix Xspatial is Bs × Nf × 1 × T, where Bs represents the batch size of 16, Nf represents the total number of convolution kernels generated by the three groups of convolutions, and T represents the time dimension. To avoid overfitting, the weights are constrained by the maximum norm. The weight vector is renormalized so that its L2 norm is less than 2. Then, BatchNorm2d is used to normalize each feature channel and combined with the Swish activation function to improve the nonlinear expression ability of the model.

[0089] After passing through the constrained grouped spatial convolution layer, the output tensor where Nf represents the number of filters (convolution kernels) learned during the spatial convolution process. Each filter can be regarded as a mapping that transforms the EEG spatial information into a new feature space, effectively capturing rich spatial features from the data. To further enhance the spatio-temporal feature extraction ability, the output is reshaped into This reshaping enables the model to maintain the differences in spatial features between channels while effectively learning time patterns. The data structure of X reshape is consistent with the original EEG, so it is regarded as a feature matrix integrating spatial and time information.

[0090] 1.2 Spatiotemporal Convolution Module:

[0091] The spatiotemporal convolution module adopts a network structure similar to ShallowConvNet, but fine-tunes the number of filters and the size of the convolutional kernel. It consists of a temporal convolution, a spatial convolution, BatchNorm2d, Avgpool2d, and a logarithmic transformation layer.

[0092] The temporal convolution uses 60 filters with a convolutional kernel size of k=(1,25).

[0093] The spatial convolution uses a convolutional kernel of size k=(60,1) to expand the feature map to 120 channels. The output of the spatial convolution is normalized using BatchNorm2d to reduce internal covariate shift and improve training stability. The custom SquareLayer module squares the elements to enhance the separability of the features. Temporal downsampling is performed using AvgPool2d with a convolutional kernel size of p=(1,100) and a stride of s=(1,10). LogLayer applies a logarithmic transformation to the features to enhance representation learning. Dropout is applied with a probability of p = 0.5 to randomly deactivate a portion of the neurons. Finally, a temporal convolution maps the 120 input channels to the output classes with a convolutional kernel size of (1,88). The LogSoftmax activation function is used, followed by the negative log-likelihood loss (NLLLoss) as the loss function.

[0094] 1.3 Evaluation Metrics:

[0095] The main evaluation metrics of this application include classification accuracy, kappa coefficient, and confusion matrix, which are used to evaluate the performance of DSTA-Net. The interpretability metrics include CSP, t-SNE, and DeepLIFT, which are used to analyze the feature representation and decision-making process of the model. Specifically, this application uses the rescaled DeepLift method to explain the feature contribution distribution of DSTA-Net. DeepLIFT calculates the attribution scores through backpropagation to quantitatively evaluate the impact of each input feature on the output decision of the neural network. For example, when analyzing the right-hand MI task, the right-hand MI task is used as the target input, while the left-hand MI task or the rest task is used as the baseline input, and it is implemented through the DeepLift module in the Captum toolkit.

[0096] 2. Experiments:

[0097] Establish an experimental dataset

[0098] The self - collected stroke patient dataset has been approved by the Ethics Committee of Shanghai No. 2 Rehabilitation Hospital (approval number: 2023 - 10 - 01). Before participating in the experiment, each participant signed an informed consent form and completed a detailed questionnaire regarding age, gender, dominant hand, previous BCI experience, and health status. The data was collected from 8 chronic right - handed stroke patients (5 males, 3 females; mean age = 52 years; Brunnstrom stages III - V; right - arm paralysis). The experimental paradigm and EEG pre - processing are as Figure 2 shown, consisting of 4 modules, with each module containing 75 trials. The tasks included left - hand MI, right - hand MI tasks, and rest. The task types were randomly assigned within each module, with 25 trials for each task type. Each trial included a 3 - second cue time, a 4 - second task time, and a 2 - second rest time. Considering that some participants were not familiar with the BCI system, they were encouraged to attempt the tasks multiple times to reach the movement threshold without actual movement.

[0099] EEG data was recorded using a Neuracle 64 - channel wireless amplifier at a sampling rate of 1000 Hz, and the MI tasks were programmed using Eprime 3.0. Data pre - processing was performed using the MATLAB plugin EEGLAB, including downsampling to 250 Hz, 0.5 - 40 Hz FIR band - pass filtering, rereferencing, electrooculogram (EOG) artifact removal based on AAR, and noise removal combining ICA and MARA. Finally, the data was randomly shuffled and divided into two sessions with balanced labels for analysis. Twenty electrodes in the motor area (including FC - 5 / 3 / 1 / z / 2 / 4 / 6, C - 5 / 3 / 1 / z / 2 / 4 / 6, and CP - 5 / 3 / 1 / 2 / 4 / 6) were selected for analysis. Table 2 provides detailed information on various datasets.

[0100] Table 2. Information on the datasets:

[0101]

[0102] The BCIIV - 2a dataset used 22 Ag / AgCl electrodes to collect EEG data from 9 participants at a sampling rate of 250 Hz. Each participant had two sessions on different dates, and each session included four types of MI tasks: left - hand, right - hand, tongue, and feet. Each task had 72 trials, and each trial contained 4 seconds of EEG data.

[0103] The OpenBMI dataset collected EEG data from 54 healthy participants using 62 electrodes. The original sampling rate was 1000 Hz and it was downsampled to 250 Hz. Each participant performed two MI-EEG recording sessions, and each session included left and right hand MI tasks. Each task had 100 trials, and each trial included 4 seconds of task-related EEG data. 20 electrodes in the motor area (including FC-5 / 3 / 1 / 2 / 4 / 6, C-5 / 3 / 1 / z / 2 / 4 / 6, and CP-5 / 3 / 1 / z / 2 / 4 / 6) were selected for analysis.

[0104] The CASIA dataset collected EEG data from 25 healthy participants without MI-BCI experience. Using 64 electrodes with a sampling rate of 1000 Hz, the participants performed three MI tasks: "rest", "hand", and "elbow", and each task had 300 trials. Each trial lasted 4 seconds. Data preprocessing used the EEGLAB toolbox, including common average reference (CAR), 0.1 - 40 Hz band-pass filtering, baseline drift removal, artifact elimination based on AAR, and downsampling to 200 Hz. Focusing on classifying the "hand" and "elbow" tasks using data from 20 electrodes in the motor area, it was consistent with the OpenBMI dataset. This dataset contained 15 sessions, which were divided into two balanced sessions, and each session included 150 trials for the left hand task and 150 trials for the right hand task.

[0105] Ten-fold cross-validation (CV) and holdout (HO) analysis were used to evaluate DSTA-Net. Although some studies only used the data from session 1 to eliminate inter-session variability, considering that limited data in deep learning models might lead to overfitting, the data from both session 1 and session 2 were selected for cross-validation. In ten-fold cross-validation, 9 folds were used for training and 1 fold was used for testing to ensure class balance and improve the generalization ability of the model. In the holdout setting, session 1 was used for training and session 2 was used for testing to evaluate the model's ability to extract generalizable features for cross-session classification.

[0106] 3. Training strategy:

[0107] All algorithms adopt a consistent training process to ensure the robustness and generalization of the model. The model uses the Adam optimizer with default parameters (learning rate = 0.001, β1 = 0.9, β2 = 0.999), and uses logarithmic cross-entropy loss to guide gradient updates. The model is trained using a two-stage training strategy. In stage 1, the training data is divided into a training set and a validation set. If the validation accuracy does not improve for 200 consecutive epochs, an early stopping strategy is applied and the parameters are restored to those at the highest validation accuracy. In stage 2, the optimal model from stage 1 is used to continue training by combining the data from the training set and the validation set, and training stops when the validation loss is lower than the training loss in stage 1. To ensure convergence, the training in stage 1 and stage 2 is limited to 1500 epochs and 600 epochs respectively. For cross-validation (CV), within a 9-fold training framework, 1 fold is selected from the training data as the validation set; for the holdout (HO) analysis, 20% of the training data is reserved as the validation set. In both settings, the test data does not participate in the training process. This comprehensive strategy effectively prevents overfitting and improves the model performance and generalization ability. All models are implemented using the PyTorch framework.

[0108] 4. Results:

[0109] 4.1 Comparison with the baseline:

[0110] Table 3 summarizes the average classification results compared with the baseline methods on four MI datasets (BCI-IV-2a, OpenBMI, CASIA, and stroke patient dataset). In addition, Table 4 shows the average classification accuracy and kappa coefficient on different algorithms and datasets in the HO analysis. In both the CV and HO analyses, DSTA-Net outperforms the baseline algorithms in terms of classification accuracy.

[0111] Table 3. Average classification accuracy (CV) on different algorithms and datasets:

[0112]

[0113] The bolded values represent the highest accuracy, and the significance of the paired t-test is marked with * (p < 0.05) and ** (p < 0.01).

[0114] Table 4. Average classification accuracy and Kappa coefficient on different algorithms and datasets (HO):

[0115]

[0116] This application conducted CV tests on four MI datasets. On the BCI-IV-2a dataset, DSTA-Net achieved an accuracy of 85.70% in four types of tasks, significantly outperforming the baseline algorithm (p < 0.05), especially showing obvious advantages in comparison with ShallowConvNet, DeepConvNet, and EEGITNet (p < 0.01). On the OpenBMI dataset, the average accuracy of DSTA-Net reached 77.18%, significantly exceeding ShallowConvNet and EEGITNet (p < 0.01), and also showing a significant improvement compared with DeepConvNet (p < 0.05). Although the difference from EEGNet was not statistically significant, DSTA-Net still had an improvement of 2.75% (77.18% compared with 74.43%). On the CASIA dataset, DSTA-Net achieved the highest average accuracy of 67.51%, significantly outperforming ShallowConvNet (p < 0.05) and DeepConvNet (p < 0.01). On the stroke patient dataset, the DSTA-Net algorithm achieved the highest classification accuracy of 67.50% in three types of classification tasks, outperforming EEGNet and EEGITNet (p < 0.01). This indicates that DSTA-Net may be more effective in capturing relevant features from the stroke patient dataset, possibly meaning it has better feature extraction capabilities and stronger model robustness.

[0117] Considering the influence of sample size on statistical power and result reliability, HO tests were conducted on the relatively large OpenBMI and CASIA datasets. As Figure 3As shown, the HO analysis results of different algorithms on the OpenBMI dataset are presented. The average accuracy of DSTA-Net reaches 65.57%, exceeding all baseline algorithms, showing a significant improvement compared with ShallowConvNet (p<0.05) and a more significant improvement compared with DeepConvNet (p<0.01). In addition, the kappa coefficient of DSTA-Net reaches 0.3115, the highest among all algorithms, indicating stronger classification consistency compared with ShallowConvNet (kappa = 0.2317), DeepConvNet (kappa = 0.2070), EEGNet (kappa = 0.2804), and EEGITNet (kappa = 0.2756). On the CASIA dataset, the average accuracy of DSTA-Net is 62.28%, continuously outperforming the baseline methods. Its kappa coefficient is 0.2456, also better than other algorithms, including ShallowConvNet (kappa = 0.1616), DeepConvNet (kappa = 0.1949), EEGNet (kappa = 0.1977), and EEGITNet (kappa = 0.2189). The detailed classification results for the four MI datasets are shown in the supplementary materials. As Figure 4 shown (the horizontal axis represents the predicted results, and the vertical axis represents the actual labels: PL / PR (predicted left / right), AL / AR (actual left / right), PE / PH (predicted elbow / hand), AE / AH (actual elbow / hand)), the confusion matrices of HO analysis on the OpenBMI and CASIA datasets are presented. To evaluate the classification performance on these datasets, the present application analyzes the confusion matrices of all subjects. On these datasets, DSTA-Net demonstrates excellent overall performance, achieving the highest correct classification rate and the lowest misclassification rate.

[0118] On the OpenBMI dataset, the correct classification rates of DSTA-Net for the two types of tasks are 64.28% and 66.87% respectively, and on the CASIA dataset, these rates are 63.39% and 61.17% respectively. Compared with other models, DSTA-Net has significantly fewer misclassifications. EEGNet shows certain competitiveness on the OpenBMI and CASIA datasets, but is slightly less robust and has a relatively higher misclassification rate compared with DSTA-Net. Models such as DeepConvNet and ShallowConvNet have lower classification accuracies and higher misclassification rates, which is more obvious on the CASIA dataset, reflecting their limitations in feature extraction and classification. Overall, these results highlight the excellent ability of DSTA-Net in EEG feature extraction and classification, making it the most robust model on these two datasets.

[0119] To further evaluate DSTA-Net, this application analyzed the recognition accuracies of the top 25% and bottom 25% of the subjects in the OpenBMI dataset under the CV strategy. As shown in Table 5, DSTA-Net achieved an accuracy of 95.21% in the group of the top 25% best-performing subjects and 58.38% in the group of the bottom 25% subjects, demonstrating strong classification ability, especially with an obvious advantage in the subjects with excellent performance.

[0120] Table 5. Performance comparison between the top 25% best-performing subjects and the bottom 25% worst-performing subjects on the OpenBMI dataset:

[0121]

[0122] 5. Ablation experiment:

[0123] The DSTA-Net in this application consists of DSTA and STC modules. The DSTA module especially includes multi-scale temporal convolution and grouped spatial convolution. To evaluate the contributions of these components, this application conducted an ablation experiment of HO analysis using the OpenBMI dataset, and the results are shown in Table 6. Scheme A represents the complete DSTA-Net, Scheme B replaces the dynamic multi-scale layer with the original signal as the input of DSTA-Net, and Scheme C uses the original signal as the input and only extracts features and classifies through the STC module. Under HO analysis, the average classification accuracy of Scheme A is 65.57%, that of Scheme B is 62.19%, and that of Scheme C is 61.58%. Scheme A has an accuracy 3.38% higher than Scheme B (p<0.01), highlighting the key role of the dynamic multi-scale layer. Scheme B has an accuracy 0.61% higher than Scheme C, indicating that the constrained spatial grouped convolution including dimensional transformation enhances the model's ability to extract multi-scale temporal features. In addition, the performance of Scheme A is improved by 3.99% compared with Scheme C (p<0.01), proving that the DSTA module helps to enhance feature extraction from the original EEG signal, thus improving the classification performance.

[0124] Table 6. Subject-level classification accuracies of the OpenBMI dataset (HO):

[0125]

[0126] MTC refers to Multi-Scale Temporal CNN, GSC represents Grouped Spatial CNN, and STC stands for SpatioTemporalCNN. A tick (√) indicates including the module, while a cross (×) indicates excluding the module.

[0127] 6. Interpretability and Visualization:

[0128] To evaluate the discriminative ability of the features extracted by DSTA-Net and the baseline methods, this study adopted the t-distributed Stochastic Neighbor Embedding (t-SNE) algorithm to visualize the high-dimensional latent features of the classification layer in a two-dimensional space. Taking the multi-class BCI-IV-2a dataset as an example, this application used the data of Subject 1 for experiments to illustrate the effectiveness of t-SNE in evaluating feature separability. In addition, this application applied the Higher-Order (HO) analysis method to evaluate the classification performance of multiple models. The classification accuracies of Subject 1 are as follows: 81.25% for DSTA-Net, 79.86% for ShallowConvNet, 77.43% for DeepConvNet, 75.35% for EEGNet, and 75.34% for EEGITNet.

[0129] Figure 5 Visualization of the DeepLift results of different algorithms on Subject 1 of the BCI-IV-2a dataset is shown.

[0130] To facilitate cross-feature dimension comparison, the attribution scores were normalized. The attribution distribution of the DSTA-Net model mainly concentrated near the right C4 region and showed obvious asymmetry, which is consistent with the neurophysiological characteristics of the left hand MI activating the right brain. In contrast, although the attribution distributions of the comparison algorithms made certain contributions near the C4 region, they were more dispersed in other brain regions, which may have reduced the classification performance of these models.

[0131] Figure 6 The t-SNE visualization results of several techniques are shown, reflecting the discriminative potential of the extracted features. The inter-class separation effect of DSTA-Net and ShallowConvNet is the best, with a significant reduction in inter-class overlap, demonstrating excellent discriminative ability. DeepConvNet and EEGNet perform slightly worse, with a certain degree of overlap in the point cloud distribution, but still maintain reasonable inter-class separability. In contrast, the inter-class overlap of EEGITNet is the most obvious, resulting in relatively weak feature discriminative ability and clustering performance.

[0132] 7. Discussion:

[0133] The Effect of DSTA-Net:

[0134] MI-BCI has the potential to promote neuroplasticity, bringing great hope for stroke rehabilitation. However, non-stationarity and high within-class variability pose huge challenges to MI decoding. Traditional single-scale CNN methods have achieved some success in MI decoding, but the performance still needs to be further improved. The introduction of multi-scale CNN has enhanced the decoding performance, but the selection of the number and size of convolutional kernels will affect the decoding results. Considering the neurophysiological and spatio-temporal characteristics of the α and β frequency bands of MI, this application proposes DSTA-Net from the perspective of enhancing spatio-temporal feature representation. To verify the performance of DSTA-Net, this application conducts tests on three MI public datasets and a self-collected stroke patient dataset.

[0135] In terms of decoding performance, Table 3 shows that in the CV analysis, compared with ShallowConvNet, the average accuracy of DSTA-Net on the BCI-IV-2a, OpenBMI, CASIA, and stroke patient datasets increased by 6.29% (p<0.01), 3.05% (p<0.01), 5.26% (p<0.01), and 2.25% respectively. Table 4 shows that in the HO analysis, compared with ShallowConvNet, the average accuracy of DSTA-Net on the OpenBMI and CASIA datasets increased by 3.99% (p<0.01) and 4.2% (p<0.01) respectively. According to Table 4 and Figure 4 , DSTA-Net is superior to other algorithms in terms of kappa coefficient and precision, indicating that it has better classification consistency and can provide more stable and reliable results. Figure 5 The performance differences of DSTA-Net on data of different qualities were further evaluated. For the top 25% of the subjects, DSTA-Net performed excellently, with an average recognition accuracy of 95.21%. However, for the bottom 25% of the subjects, the accuracy of DSTA-Net was 58.38%, which was 1.64% lower than that of EEGNet. This is because it is difficult to extract motor intention features from the data of 25% of the subjects, and the noise interference increases, making it difficult for the DSTA module to achieve significant improvement.

[0136] In the ablation experiment, Table 6 highlights the important contributions of the key modules of DSTA-Net, namely the dynamic multi-scale temporal convolutional neural network, the constrained spatial grouping convolution containing dimensional transformation, and the spatio-temporal convolution. On the OpenBMI dataset, DSTA-Net had the highest average accuracy (65.57%), which was better than Scheme B (62.19%) and Scheme C (61.58%), and was 3.99% higher than Scheme C (p<0.01). This emphasizes the synergistic effect of its modules in enhancing EEG feature representation. Figure 5 and Figure 6It shows that the inter-class separability of DSTA-Net is the most obvious and the overlap is the smallest, reflecting its excellent performance in feature extraction and classification.

[0137] To further study the statistical differences in CV analysis, this application conducted a paired t-test on the results, as Figure 7 shown. On all four MI datasets, the average recognition rate of DSTA-Net was consistently higher than that of Scheme C (P<0.01). Although the improvement on the patient dataset did not reach statistical significance, the recognition performance still increased by 2.25%.

[0138] Interpretability and visualization of stroke patient data:

[0139] In clinical applications, analyzing the brain activity patterns of stroke patients helps to deeply understand the neural mechanisms in the MI task. This understanding is crucial for improving the accuracy and reliability of the BCI system in the clinical environment. In this study, this application used DeepLift to evaluate the contribution of each EEG channel to the model's decision-making, comparing the target task with a reference task (such as the resting state). In addition, this application used CSP to study the activation patterns of patients. The analysis objects were Subject 1 with the highest classification accuracy (96.33%) and Subject 3 with the lowest (42.67%) in the stroke patient dataset.

[0140] Figure 8 Shows the DeepLift results of Stroke Patient 1 and Patient 3 using the DSTA-Net algorithm (both patients are male and have right-sided hemiplegia; the upper limb function assessment scale (UFMA) score of Patient 1 is 61 and that of Patient 3 is 33). In this study, the left-hand task and the right-hand task were designated as the target tasks, and the resting task was used as the reference task. The results showed that Patient 1 showed significant discriminative ability between tasks, with a classification accuracy of 96.33%, while the accuracy of Patient 3 was only 42.67%. As Figure 8 shown, during the left-hand MI task, the right motor area of Patient 1 showed significant and concentrated activation, while partial compensatory activation occurred in the left hemisphere. In contrast, during the right-hand MI task, the left brain region showed clear and concentrated activation, and partial compensatory activation occurred in the right brain region. In contrast, Patient 3 had lower discriminability between the left-hand and right-hand MI tasks, with activation in both hemispheres. The DeepLift results showed that the right hemisphere of Patient 3 was in a more active state during the left-hand task.

[0141] Since the activation levels and regions identified by DeepLift largely depend on the model's feature extraction ability, this application further analyzed the activation levels and regions during the MI task from the perspective of the original EEG signals. This analysis used the classic CSP features. Considering that the CSP variances are relatively high and vulnerable to noise artifacts, this application performed meticulous preprocessing on the original signals, such as Figure 2 shown, performing CSP analysis on 100 left-hand tasks and 100 right-hand tasks for each patient, aiming to maximize the variance difference between the two types of tasks.

[0142] Figure 9 shows the topographic activation patterns in the α (8 - 13 Hz) and β (13 - 30 Hz) frequency bands during the left-hand and right-hand MI tasks for Patient 1 and Patient 3. For Patient 1, during the left-hand MI task, the left hemisphere showed significant α-band suppression (represented by blue), while the right hemisphere had obvious β-band activation (represented by red). Conversely, during the right-hand MI task, there was no obvious α-band suppression in the right hemisphere, and the left hemisphere showed clear β-band activation. These results indicate that Patient 1 exhibited obvious event-related desynchronization / event-related synchronization (ERD / ERS) phenomena in both the α and β frequency bands. In contrast, Patient 3 had lower discriminability between the left-hand and right-hand MI tasks. Specifically, during the left-hand MI task, there was no significant α-band suppression in the left hemisphere, and both β-band activation (red) and α-band suppression (blue) occurred simultaneously in the right hemisphere. Similarly, during the right-hand MI task, there was no obvious α-band suppression in the right hemisphere, and the left hemisphere showed clear β-band activation. Therefore, Patient 3 exhibited ERD / ERS phenomena in the β frequency band, while the influence in the α frequency band was not significant. Finally, based on Figure 5 , Figure 8 and Figure 9 the brain region activation positions and regions shown in, this application preliminarily found that DeepLIFT has more stable feature interpretability than common spatial pattern (CSP) and shows more concentrated brain region activation in high-precision data.

[0143] Conclusion: In this study, this application proposed DSTA-Net for MI classification. The DSTC module performs dynamic multi-scale feature extraction and representation in the time and space domains, achieving satisfactory recognition results on multiple datasets and demonstrating strong interpretability.

[0144] The above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A dynamic spatiotemporal feature enhanced network model for motor imagery classification, characterized in that: Including dynamic spatiotemporal feature enhancement module and spatiotemporal convolution module, The dynamic spatiotemporal feature enhancement module includes a multi-scale temporal convolution layer for extracting alpha and beta frequency band features in MI-related electroencephalogram signals and retaining the original EEG signal as a baseline feature; The dynamic spatiotemporal feature enhancement module also includes a grouped spatial convolution layer for extracting multi-level spatial features and preventing overfitting through a weight constraint mechanism; The spatiotemporal convolution module is used to further extract spatiotemporal features and perform classification.

2. The dynamic spatiotemporal feature enhanced network model for motor imagery classification according to claim 1, characterized in that: The multi-scale temporal convolution layer includes multiple levels of one-dimensional temporal convolution kernels, where the time length of the i-th convolution kernel is defined as: Used to dynamically capture the spatiotemporal characteristics of different frequency ranges in EEG signals; pass and Determine the convolution kernel size to match the characteristic range of the alpha and beta frequency bands in the motor imagery task; An original signal compensation layer is introduced into the multi-scale temporal convolution layer to reduce the feature matrix distortion caused by multi-scale convolution; The outputs of the original signal compensation layer, the α-band feature layer, and the β-band feature layer are Perform splicing to generate enhanced spatiotemporal feature representation.

3. The dynamic spatiotemporal feature enhanced network model for motor imagery classification according to claim 1, characterized in that: The constrained grouped spatial convolution layer includes three groups of spatial convolution modules, each group of spatial convolution modules adopts a grouped spatial convolution operation with a kernel size of (Nc, 1), where Nc represents the number of EEG signal channels; Each group of spatial convolution modules generates 10 convolution kernels, and the feature maps of the three groups of outputs are spliced ​​along the convolution kernel dimension to obtain a feature matrix Xspatial with a shape of Bs×Nf×1×T, where Bs represents the batch size of 16, Nf=30 represents the total number of convolution kernels, and T represents the time dimension; The weight vector of each convolution kernel is constrained by the maximum norm, and the L2 norm value of the weight vector is made less than 2 through L2 norm renormalization. Each feature channel of the feature matrix Xspatial is subjected to BatchNorm2d normalization and Swish activation function transformation in turn; The activated feature matrix is ​​reshaped into a three-dimensional tensor structure, keeping its channel dimension consistent with the spatial dimension of the original EEG signal, forming a feature matrix that integrates spatiotemporal features.

4. The dynamic spatiotemporal feature enhanced network model for motor imagery classification according to claim 1, characterized in that: The spatiotemporal convolution module includes Temporal convolution layer: 60 filters are used to perform time-dimensional convolution on the input EEG signal, with a convolution kernel size of k = (1, 25) to extract time domain features; Spatial convolution layer: convolve the output of the temporal convolution layer in the spatial dimension, with a convolution kernel size of k = (60, 1), expanding the number of feature map channels to 120; Batch Normalization Layer: Perform BatchNorm2d normalization on the output of the spatial convolutional layer to reduce internal covariate shift; Square enhancement layer: The normalized features are squared element by element through the custom SquareLayer module to enhance feature separability; Average pooling layer: Use AvgPool2d to downsample the features in the time dimension, with a pooling kernel size of p = (1, 100) and a step size of s = (1, 10); Logarithmic transformation layer: Logarithmic transformation is performed on the pooled features through the LogLayer module to optimize the feature representation; Random dropout layer: Dropout operation is applied to the features with probability p = 0.5, randomly deactivating some neurons to prevent overfitting; Classification convolution layer: 120 input channels are mapped to output categories through temporal convolution, the convolution kernel size is (1,88), and the LogSoftmax activation function is used to output the classification probability; Loss calculation: Optimize model parameters based on negative log-likelihood loss.

Citation Information

Patent Citations

  • Parallel convolutional neural network motor imagery electroencephalogram classification method based on spatial-temporal feature fusion

    CN111012336A

  • Feature fusion method based on multilevel electroencephalogram signal expression

    CN113128459A

  • Motor imagery electroencephalogram signal recognition method based on EMD data enhancement and parallel SCN

    CN115221969A

  • Motor imagery electroencephalogram signal processing method based on FA-CNN and FA-CNN model

    CN117390543A

  • Analysis method of convolutional neural network based on Wavelet transform for identifying motor imagery brain waves

    KR102096565B1