EEG (electroencephalogram) signal decoding method, device and equipment based on aggregation perception enhanced convolution Transform network and medium
Through the method of augmented convolutional Transformer network based on aggregation perception, the problem of insufficient multi-scale spatiotemporal feature capture in EEG signal decoding is solved, and higher decoding accuracy and feature correlation are achieved.
Patent Information
- Application Number
- CN202510550978.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The prior art is difficult to effectively capture multi-scale spatiotemporal features in EEG signal decoding, resulting in limited decoding accuracy.
The EEG signal decoding method based on the aggregation perception enhancement convolution Transformer network is adopted. Multi-scale shallow local features are extracted through the spatiotemporal convolution module, the adaptive feature recalibration module strengthens key features, the position perception enhancement module refines features, and the sparse information aggregation Transformer module extracts long-range dependencies and local associations.
The multi-scale feature interaction fusion of EEG signals is realized, the decoding accuracy is improved, the correlation between local and global features is balanced, and the depth and breadth of EEG data analysis is improved.
Smart Images

Figure CN120067843A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of EEG signal decoding, and in particular to an EEG signal decoding method, device, equipment and medium based on an aggregated perception enhanced convolutional Transformer network. Background Art
[0002] The rapid development of brain-computer interface technology has put forward higher requirements for the accuracy and efficiency of EEG signal decoding. As a typical non-stationary and nonlinear physiological signal, EEG signals have the characteristics of low signal-to-noise ratio, large individual differences, and complex spatiotemporal dynamic changes. How to effectively extract its spatiotemporal dynamic features has become the key to improving decoding accuracy. Especially in motor imagery tasks, the characteristic expression of EEG signals needs to take into account both the rhythmic changes in the time dimension and the cortical activation patterns in the spatial dimension, which poses a severe challenge to the spatiotemporal modeling capabilities of feature extraction methods.
[0003] The current EEG signal decoding method mainly adopts the technical route of combining traditional feature engineering with deep learning. Traditional methods rely on manually designed time-frequency domain feature extraction, such as wavelet transform and power spectrum analysis, which have the limitations of high feature dimension and poor generalization ability. The deep learning model based on convolutional neural network captures spatiotemporal features through local receptive fields, but the convolution kernel of fixed scale is difficult to adapt to the multi-scale characteristics of EEG signals, resulting in the loss of fine-grained features. Some studies have attempted to introduce multi-scale convolutional structures to enhance feature expression capabilities, but the interaction mechanism between features of each scale is missing, making it difficult to establish effective cross-scale associations. In addition, methods that rely solely on convolutional neural networks have bottlenecks in modeling long-term dependencies, and although the model using the Transformer architecture can capture global relationships, it ignores the fine characterization of local features, resulting in the loss of important detail information.
[0004] Existing technologies still have obvious deficiencies in the joint modeling of multi-scale spatiotemporal features. Traditional convolutional neural networks are limited by local receptive fields and cannot effectively capture dynamic cross-channel and cross-time correlations in EEG signals. Although multi-scale convolutional structures can extract features of different granularities, each branch feature lacks an effective information interaction and recalibration mechanism, making it difficult to achieve enhanced expression of key features. At the same time, there is a contradiction between computational complexity and local feature retention in the global modeling method based on the attention mechanism. Excessive attention to global dependencies will weaken the focus on important local features, while simply emphasizing local features will lead to a lack of contextual associations. This problem of insufficient capture of spatiotemporal features, loss of fine-grained information, and imbalance in global and local relationships has seriously restricted the further improvement of EEG signal decoding accuracy. Summary of the invention
[0005] The present invention provides an EEG signal decoding method, device, equipment and medium based on an aggregation-aware enhanced convolutional Transformer network to improve at least one of the above technical problems.
[0006] In a first aspect, the present invention provides an EEG signal decoding method based on an aggregation-aware enhanced convolutional Transformer network, which includes steps S1 to S9.
[0007] S1. Obtain the offline data of the EEG signal.
[0008] S2. Extract multi-scale shallow local features through a spatio-temporal convolutional module according to the offline data.
[0009] S3. Perform interactive reinforcement learning on the multi-scale shallow local features through an adaptive feature recalibration module to obtain multiple recalibrated features.
[0010] S4. Fuse the multiple recalibrated features to obtain a first fused feature.
[0011] S5. Input the first fused feature into a position-aware enhancement module to extract deep fine-grained features through parallel enhanced convolutions, and perform adaptive encoding on the deep fine-grained features to obtain position-aware enhanced features.
[0012] S6. Extract long-range dependencies and local associations through a sparse information aggregation Transformer module to obtain globally refined features.
[0013] S7. Input the globally refined features into a classifier to complete the training of the model and obtain a pre-trained model that can be used for real-time prediction.
[0014] S8. Obtain the real-time data of the EEG signal.
[0015] S9. Input the real-time data into the pre-trained model to obtain the decoding result of the real-time EEG signal.
[0016] As a preferred aspect of the present invention, step S1 specifically includes steps S11 to S15.
[0017] S11. Obtain the stored EEG signal data, eliminate 50Hz power frequency and environmental noise interference, and screen out the EEG signals in the key frequency band of 0.5 - 30Hz.
[0018] S12. Perform segmented cutting on the continuous EEG signals, remove the rest segment data, and only retain the MI segment data.
[0019] S13. Using the average value of the rest segment data in the first 300 milliseconds of the MI segment data as the baseline, perform independent baseline correction on each MI segment data to eliminate the potential impact of baseline shift.
[0020] S14. After normalizing the data, remove the artifacts mixed in the EEG signal to obtain the preprocessed EEG data.
[0021] S15. Divide and reconstruct the training data according to the label, and additionally add random values following the Gaussian distribution to simulate batch data. The newly generated data is shuffled and mixed with the original batch data to obtain the data-augmented EEG data, and then jointly input into the model to learn features.
[0022] As a preferred aspect of the present invention, step S8 specifically includes steps S81 to S83.
[0023] S81. Store the collected real-time EEG data stream through a buffer.
[0024] S82. Extract real-time segment data from the buffer through a sliding time window.
[0025] S83. Preprocess the real-time segment data; wherein, the preprocessing steps of the real-time data are less than those of the offline data by data augmentation processing, and the baseline correction operation is based on the overall mean of the segment.
[0026] As a preferred aspect of the present invention, the spatio-temporal convolution module is provided with three branches. Each branch is sequentially provided with a temporal convolution layer, a spatial convolution layer, a batch normalization layer, an ELU activation function layer, and an average pooling layer. The temporal convolution layer of the first branch has 16 (1, ) convolutional kernels. The temporal convolution layer of the second branch has 16 (1, ) convolutional kernels. The temporal convolution layer of the third branch has 16 (1, ) convolutional kernels. Fs is the sampling frequency. The spatial convolution layer uses 32 convolutional kernels of size (C, 1), where C is the number of channels. Among them, during model training, a dropout layer is also set.
[0027] As a preferred aspect of the present invention, step S2 specifically includes steps S21 to S22.
[0028] S21. Input the EEG signal into the three branches of the spatio-temporal convolution module respectively to obtain multi-scale shallow local features.
[0029] S22. The three branches of the spatio-temporal convolution module respectively perform convolution operations through a temporal convolution layer and a spatial convolution layer in sequence to extract features, and input the features extracted by the convolution operation into a batch normalization layer, an ELU activation function layer, and an average pooling layer connected in sequence.
[0030] As a preferred aspect of the present invention, step S3 specifically includes step S31 and step S32.
[0031] S31. Feature superposition is performed on the features extracted from each branch of the spatio-temporal convolution module in an interactive connection structure to obtain a plurality of interactive superposition features with the same number as the number of branches of the spatio-temporal convolution module.
[0032] S32. Through a convolutional attention mechanism, the features of the plurality of interactive superposition features along the channel and space are respectively calculated to obtain a plurality of recalibrated features.
[0033] As a preferred aspect of the present invention, step S5 specifically includes step S51 to step S54.
[0034] S51. According to the first fusion feature, features are further extracted in the time direction through a parallelly arranged 1x3 convolution and 1x7 convolution respectively. The extracted features are processed through batch normalization and then feature fusion is performed.
[0035] S52. The features after feature fusion are processed using an ELU activation function and an average pooling layer. Among them, a dropout layer is also set in the training stage.
[0036] S53. Through skip connection, the features processed by the average pooling layer and the first fusion feature are added together.
[0037] S54. The added features are subjected to dimensional conversion and then adaptive coding is performed to obtain a position perception enhanced feature.
[0038] As a preferred aspect of the present invention, step S6 specifically includes step S61 to step S64.
[0039] S61. The position perception enhanced feature is divided into blocks through a sliding window to obtain a plurality of blocks.
[0040] S62. The plurality of blocks are respectively averaged, and the continuous Tokens within the blocks are averaged and aggregated into a single block representation to obtain an aggregated block.
[0041] S63. Through the highest attention mechanism, the attention scores of each aggregated block are calculated, k important blocks are selected, and the original Tokens of the important blocks are restored to obtain the highest attention block.
[0042] S64. Combine the aggregated block and the highest attention block through a gating mechanism to obtain the global refined features.
[0043] As a preferred aspect of the present invention, step S7 specifically includes steps S71 to S73.
[0044] S71. Use a flattening layer to perform a flattening and dimensionality reduction operation on the features, converting the multi-dimensional features into one-dimensional for feature integration.
[0045] S72. Process the flattened features through two fully connected layers. Among them, a dropout layer is inserted into the two fully connected layers during the training phase.
[0046] S73. Pass the features processed by the fully connected layers through the softmax function to calculate the prediction probability of each category, thereby completing the training of the model and obtaining a pre-trained model that can be used for real-time prediction.
[0047] As a preferred aspect of the present invention, step S9 specifically includes steps S91 to S93.
[0048] S91. At the first prediction, after the buffer accumulates a complete window of data, use the pre-trained model to carry out the prediction.
[0049] S92. For each newly acquired data volume of the sliding distance, use the pre-trained model to predict the data within this window.
[0050] S93. Map the prediction result of each time to a control instruction.
[0051] Second aspect, the present invention provides an EEG signal decoding device based on an aggregation-aware enhanced convolutional Transformer network, which includes a signal acquisition module, a spatio-temporal convolution module, a recalibration module, a fusion module, a position perception module, an information aggregation module, and a classifier module.
[0052] An offline data acquisition module for acquiring offline data of EEG signals.
[0053] The spatio-temporal convolution module is used to extract multi-scale shallow local features through the spatio-temporal convolution module according to the offline data.
[0054] The recalibration module is used to perform interaction sharing and enhance key features on the multi-scale shallow local features through an adaptive feature recalibration module to obtain multiple recalibrated features.
[0055] The fusion module is used to fuse the multiple recalibrated features to obtain the first fusion feature.
[0056] A position perception module, configured to input the first fused feature into a position perception enhancement module to extract deep fine-grained features through parallel enhancement convolutions, and perform adaptive encoding on the deep fine-grained features to obtain position perception enhancement features.
[0057] An information aggregation module, configured to extract long-range dependencies and local associations through a sparse information aggregation Transformer module to obtain global refined features.
[0058] A classifier module, configured to input the global refined features into a classifier to complete the training of the model, and obtain a pre-trained model that can be used for real-time prediction.
[0059] A real-time data acquisition module, configured to acquire real-time data of EEG signals.
[0060] A real-time decoding module, configured to input the real-time data into the pre-trained model to obtain a decoding result of the real-time EEG signal.
[0061] In a third aspect, the present invention provides an EEG signal decoding device based on an aggregation-aware enhanced convolutional Transformer network, which is characterized by including a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement an EEG signal decoding method based on an aggregation-aware enhanced convolutional Transformer network as described in any paragraph of the first aspect.
[0062] In a fourth aspect, the present invention provides a computer-readable storage medium, which is characterized in that the computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute an EEG signal decoding method based on an aggregation-aware enhanced convolutional Transformer network as described in any paragraph of the first aspect.
[0063] By adopting the above technical solutions, the present invention can achieve the following technical effects: The EEG signal decoding method based on the aggregation-aware enhanced convolutional Transformer network of the present invention performs excellently in the EEG classification and recognition task, and effectively solves some defects of traditional networks in feature extraction. The network is a fusion network structure with multi-scale feature interaction, and gradually refines features in a hierarchical manner. The spatio-temporal convolution and the adaptive feature recalibration module capture roughly coarse-grained features for the network, refine the features through the position perception enhancement module, and input position encoding information to strengthen the internal connection between features. The sparse information aggregation Transformer module strengthens long-range dependencies and local associations, balances local and global correlations, and comprehensively improves the depth and breadth of EEG data analysis. Description of the Drawings
[0064] To more clearly illustrate the technical solution of the present invention, the following will briefly introduce the drawings required for the specific implementation of the present invention. It should be understood that the following drawings only show certain specific implementation manners of the present invention, and thus should not be regarded as a limitation on the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0065] Figure 1 It is a schematic flowchart of the EEG signal decoding method.
[0066] Figure 2 It is a network overall architecture diagram of the EEG signal decoding method.
[0067] Figure 3 It is a flowchart of EEG signal preprocessing.
[0068] Figure 4 It is a structural diagram of a single branch of the spatio-temporal convolution module.
[0069] Figure 5 It is a structural diagram of the position perception enhancement module.
[0070] Figure 6 It is a structural diagram of the sparse information aggregation Transformer.
[0071] Figure 7 It is a flowchart of real-time data processing.
[0072] Figure 8 It is a logic diagram of real-time data processing. Specific Embodiments
[0073] The following will refer to the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0074] Embodiment 1. Please refer to Figures 1 to 8 , the first embodiment of the present invention provides an EEG signal decoding method based on an aggregation perception enhancement convolutional Transformer network (abbreviation: APCformer), which can be executed by an EEG signal decoding device (hereinafter referred to as: decoding device). In particular, it is executed by one or more processors in the decoding device. It can be understood that the decoding device can be a portable notebook computer, a desktop computer, a server, a smart phone, a tablet computer or other electronic devices with computing performance.
[0075] EEG represents brain waves, and the EEG signal is an electroencephalogram signal. In exploring deep learning methods for EEG data analysis, there is a challenge of how to effectively decode spatio-temporal dynamic information. Therefore, the analysis model needs to have the ability of precise multi-domain feature joint learning and powerful sequence processing ability. For this purpose, the present invention proposes a new EEG decoding network: APCformer. The overall architecture of the network is as Figure 2 shown. The network mainly includes five parts: a spatio-temporal convolution module, an adaptive feature recalibration module (abbreviation: AFR), a position-aware enhancement module (abbreviation: PAE), a sparse information aggregation Transformer module (abbreviation: SAT), and a classifier module.
[0076] A brief description of the EEG signal decoding method based on the aggregated perception-enhanced convolutional Transformer network of the present invention is as follows: Set the preprocessed EEG data as , where is a real number, C represents the number of EEG channels, and T represents the number of sampling points along the time dimension. Input the spatio-temporal convolution module with a batch size of N = 32 to extract multi-scale shallow local features, then focus on key spatio-temporal features through the AFR module, input the fused features into the PAE module to extract deep fine-grained features and perform adaptive encoding on the features, and then refine the features through the SAT module to further extract long-range dependencies and local associations, realizing the effective aggregation of local features and global features. Finally, the classifier outputs the classification result. Through the previous steps, the model is trained to obtain a pre-trained model that can be used for real-time prediction. Then, a sliding window is used to extract real-time data streams for real-time preprocessing, and the data is predicted through the pre-trained model to obtain the final prediction result and map it to an instruction.
[0077] Table 1 APCformer network architecture parameters
[0078] In Table 1, N is the number of batch samples. C is the number of channels. T is the number of sampling points. F 1 and F 2 respectively represent the number of filters of the corresponding convolutional layers. F 1 = 16, F 2 = 32. Fs is the sampling frequency. axis is the dimension to which the normalization operation is applied. ELU represents the activation function used for convolution. Softmax represents the activation function used for classification.
[0079] S1. Obtain the offline data of the EEG signal.
[0080] In this embodiment, step S1 is the preprocessing of offline data before the training phase. Specifically, it includes: data filtering, segmentation cutting, independent baseline correction, data normalization, artifact removal, and data augmentation. Among them, data augmentation includes segmentation reconstruction and adding Gaussian noise.
[0081] Specifically, first, eliminate the 50Hz power frequency and environmental noise interference, and screen out the 0.5 - 30Hz key frequency band EEG signals as the basis for subsequent analysis. Then, perform segmentation cutting on the continuous EEG signals, only retain the MI segment data closely related to the task, and splice these segments. Among them, MI represents Motor Imagery, and the MI segment data is the motor imagery segment data obtained after segmentation cutting of the EEG signals.
[0082] To further improve the signal quality, according to the time characteristics of human response to stimuli, select the average value of the 300 - millisecond rest segment data before the start of each MI segment as the baseline, and then perform independent baseline correction on each MI segment to eliminate the potential impact of baseline shift. Subsequently, use Z - score to normalize the data and then remove the artifacts that are often confused with the EEG signals.
[0083] Since small - sample EEG data is prone to causing model over - fitting, the present invention performs data augmentation on the training data. Specifically, in each batch training process, randomly select samples by category, evenly divide these samples into multiple segments, and randomly select one segment from each sample to form new data, but the time - order principle needs to be followed during combination. Then, data augmentation is achieved by adding random values following a Gaussian distribution to the new data, so as to improve the model's robustness and generalization ability. Finally, add the new data to the original training data and shuffle the order to jointly complete the training.
[0084] Data augmentation is performed before the model trains and learns each batch of data. The batch size of the model is set to 32. After introducing the data augmentation technology, the system will perform augmentation processing on an input batch of data before the model starts learning data features. Specifically, the system will segment and reconstruct the batch data according to the labels of the data, and generate simulated data by adding Gaussian noise. These newly generated data will be mixed with the original batch data and randomly shuffled, thus doubling the data volume. Finally, the processed batch data is jointly input into the model to learn features.
[0085] S2. Extract multi - scale shallow local features through a spatio - temporal convolution module according to the offline data.
[0086] To enable the model to no longer be limited to a single convolutional perception range but be able to flexibly process feature information of different scales and levels, the present invention uses multi-scale spatio-temporal convolution to extract local features of EEG data, enhancing the model's feature recognition ability without increasing the network depth.
[0087] The first core component of the network structure of the EEG signal decoding method based on the aggregation perception enhanced convolutional Transformer network of the present invention is the spatio-temporal convolution module based on the CNN architecture. As Figure 2 shown, the spatio-temporal convolution module is provided with three branches. The structure of each branch is as Figure 4 shown, and in sequence, it is provided with a temporal convolution layer, a spatial convolution layer, a batch normalization layer, an ELU activation function layer, an average pooling layer, and a dropout layer.
[0088] Specifically, in the spatio-temporal convolution module, the structural composition of each branch is roughly the same, and double-layer convolution is used as the shallow feature encoder. However, the temporal convolution of the first layer has differences in design, and 16 different-scale large convolution kernels are respectively set, specifically (1, ), (1, ), and (1, ). The purpose is to obtain richer temporal dimension features through convolution kernels with different receptive fields. The spatial convolution of the second layer uses 32 convolution kernels of size (C, 1), and the size of this convolution kernel is the same as the device sampling channel number, which helps to process the spatial features of the data. After the convolution operation is completed, batch normalization is connected to alleviate the problem of covariate shift. Then, the nonlinear expression ability of the model is enhanced through the ELU activation function. Subsequently, the feature dimension is reduced through an average pooling layer of size (1, 75) to reduce redundant features. During model training, the dropout technique is used to accelerate convergence to reduce the risk of overfitting, and finally, more representative shallow spatio-temporal features are obtained.
[0089] S3. Through the adaptive feature recalibration module, interactively share and strengthen key features of the multi-scale shallow local features to obtain multiple recalibrated features.
[0090] Step S3 specifically includes step S31 and step S32.
[0091] S31. Feature stack the features extracted by each branch of the spatio-temporal convolution module with an interactive connection structure to obtain multiple interactive stacked features with the same number as the number of branches of the spatio-temporal convolution module.
[0092] S32. Through the convolutional attention mechanism, calculate the features of multiple interactive stacked features along the channel and space respectively to obtain multiple recalibrated features.
[0093] Specifically, the present invention creates an Adaptive Feature Recalibration (AFR) module to re-identify and calibrate the learned spatio-temporal features, emphasizing key partial features, enhancing the model's sensitivity to important information, and improving the accuracy of the decoding task for small datasets.
[0094] The Adaptive Feature Recalibration module first realizes the sharing of feature information between branches through an interactive connection structure. After the feature interaction between branches, a convolutional attention mechanism is introduced to strengthen the learning of key features with its excellent precise focusing ability.
[0095] S4. Effectively fuse the features passing through the Adaptive Feature Recalibration module in each branch to obtain a more comprehensive feature representation.
[0096] The fusion model is as follows: 。
[0097] In the formula, is the fused feature of the branch network, represents the feature concatenation operation, respectively represent the features output by the first branch, the second branch, and the third branch.
[0098] S5. Input the first fused feature into the Position Aware Enhancement (PAE) module to extract deep fine-grained features through parallel enhanced convolutions, and perform adaptive encoding on the deep fine-grained features to obtain position-aware enhanced features.
[0099] In previous networks, there was a problem of loss of fine-grained information in the spatial and temporal dimensions. The present invention proposes a Position Aware Enhancement (PAE) module, which can take into account both fine-grained and coarse-grained information, enhancing the APCformer network's feature learning ability, enabling subsequent SAT to optimize global and local features based on richer and more relevant features. The Position Aware Enhancement module designs a structure of parallel small convolutional kernels. By reducing the receptive field of the convolutional kernel, it helps to capture fine-grained local features. Its structure is as Figure 5 shown. Transmit the first fused feature to the PAE module to extract deep fine-grained EEG features.
[0100] The Position Aware Enhancement module is as Figure 5 shown. After the features pass through parallel enhanced convolutions, the dimension of the features is transformed, and position encoding information is added to the sequence.
[0101] Specifically, step S5 specifically includes steps S51 to S54.
[0102] S51. According to the first fusion feature, features are further extracted in the time direction through a 1x3 convolution and a 1x7 convolution set in parallel respectively. After the extraction, the features are processed by batch normalization and then feature fusion is performed.
[0103] S52. The features after feature fusion are processed using the ELU activation function and an average pooling layer. Among them, a dropout layer is also set in the training stage.
[0104] S53. Through skip connection, the features processed by the average pooling layer and the first fusion feature are added together.
[0105] S54. The features after addition are subjected to dimensional transformation and then adaptive coding is performed to obtain position-aware enhanced features.
[0106] Specifically, the input first fusion feature further extracts features in the time direction through 32 convolutional kernels with sizes of (1, 3) and (1, 7) respectively set in parallel. The features obtained from the two convolutions are batch-normalized, and then feature fusion is performed to enrich the extracted information.
[0107] Subsequently, the ELU activation function is used as the activation layer, and the model is optimized through an average pooling layer with a size of (1, 3) and a dropout layer. In order to address the problem of insufficient adaptability of EEG signals in deep networks, skip connections are added, which can reduce the risk of model overfitting to a certain extent and improve the stability of the model. Then, the feature dimension is linearly transformed, a learnable positional encoder (PE) is added, a trainable matrix with the same size as the input feature dimension is set, and the matrix is randomly initialized. The parameters of the positional encoder will be continuously updated as the network is trained, autonomously mining the correlation patterns between different positions, so that the network can perceive the position information of each feature.
[0108] S6. The position-aware enhanced features are refined through a sparse information aggregation Transformer module to further extract long-range dependencies and local correlations, realize the effective aggregation of local features and global features, and obtain globally refined features. Preferably, step S6 specifically includes steps S61 to S64.
[0109] It can be understood that EEG signals often have cross-period correlations in the time series and it is difficult to effectively capture the long-term dependence relationship in EEG signals, which is extremely crucial for accurate decoding. And Transformer demonstrates excellent characteristics when processing sequence data, and each element in the sequence can pay attention to all other elements, so as to accurately capture the long-term dependence relationship.
[0110] Although convolution has unique advantages, it mainly focuses on local information and has obvious deficiencies in dealing with global correlations. Simply integrating the Transformer architecture can capture long-range dependencies but will ignore local fine-grained features in the temporal order.
[0111] Therefore, the present invention designs a Sparse Information Aggregation Transformer module (abbreviation: SAT) to break the bottleneck that traditional networks have difficulty in simultaneously considering global and local features when processing EEG signals. At the same time, the idea of sparse attention is introduced to accelerate the efficiency in processing large-scale and high-dimensional EEG data, so as to efficiently capture all relationships within the sequence. By fully exploiting the complex relationships and dependencies between features through SAT, the performance of the model is further improved.
[0112] As Figure 6 shown, the SAT structure includes: a sliding window, an aggregation attention, a top attention, and a gating mechanism.
[0113] S61. Divide the position-aware enhanced features into blocks through a sliding window.
[0114] Specifically, is the feature sequence input to SAT , and are the first feature, the second feature, and the th feature respectively.
[0115] The sliding window divides the input sequence into blocks and then processes these blocks. Each block can be regarded as a local area, and finally these local information will be integrated into global features through the attention mechanism.
[0116] Specifically, the present invention uses a sliding window with a fixed length of to divide the input feature sequence of length n into m blocks, and the sliding interval between adjacent blocks is set to , and non-overlapping blocks can be divided. Among them, . In this way, each block contains consecutive Tokens, and a total of Tokens will be generated, which is greater than n.
[0117] Since the present invention sets a sliding interval instead of directly dividing the entire sequence into multiple blocks, the computational amount is increased compared to before the division into blocks. However, this can supplement the connection between blocks to a certain extent, and the degree of information overlap between blocks can be adjusted by controlling the sliding interval, that is, the sparsity ratio is controlled. While maintaining the independence of local features, the global context association is enhanced, making the entire network more flexible.
[0118] S62. Average each of the multiple blocks respectively, perform an average operation on consecutive Tokens within the block to aggregate them into a single block representation, and obtain the aggregated feature.
[0119] In the initial stage of the aggregated attention, considering the global perspective, the concept of sparse attention is introduced to quickly capture the global pattern of the sequence. It is not necessary for each position in the sequence to perform attention calculation with all other positions. Instead, the m divided blocks are averaged respectively, and consecutive Tokens within the block are averaged to aggregate into a single block representation to achieve feature capture of the global pattern.
[0120] Unfold the Tokens within all blocks in parallel in order. The position information of the first Token in any block can be expressed as The position information of any Token within the block can be expressed as . and respectively represent the key-value pairs of the i-th Token in the j-th block, and the corresponding block key and block value vectors are respectively , .
[0121] , is defined as: .
[0122] .
[0123] In the formula, represents the -th block, and its value is represents the sequence number of the -th Token within the block, and its value is .
[0124] Perform a dot product operation on each query in the sequence and all aggregated block keys to obtain the attention score, and then obtain the weighted sum through the block value . Among them, is the query weight, is the weight of the block key, is the weight of the block value.
[0125] The output of the aggregated attention path can be expressed as: .
[0126] In the formula, , represents the position of any Token, and is the block key and the block value set, represents transpose, represents the feature vector dimension.
[0127] By aggregating the blocks, the global features of the entire sequence can be obtained, and the computational complexity can be reduced from the original to , , significantly improving the processing efficiency of long sequences.
[0128] S63. Through the highest attention mechanism, calculate the attention scores of each aggregated block, select the k important blocks, restore the original Tokens of the important blocks, and obtain the highest attention block.
[0129] The aggregated attention based on blocks can reduce the computational complexity and achieve efficient long-range dependence modeling. However, simply relying on block representations will make it difficult to restore the fine-grained information of the original sequence, and this operation will inevitably cause loss of feature information. To balance efficiency and accuracy, the present invention discriminates and selects block information, retains the key and representative local features therein, so as to achieve efficient long-range dependence modeling without loss of accuracy. For this purpose, the present invention incorporates the highest attention mechanism on the basis of aggregated attention, selects the blocks with significant contributions, and constructs a hierarchical attention structure.
[0130] Specifically, the present invention evaluates the importance of the m blocks segmented in the aggregated attention. The evaluation criterion is the attention score obtained by the block. Based on this criterion, there is no need to recalculate, which can avoid generating new computational overhead. Subsequently, compare the attention scores of all blocks, select the k positions with the largest scores, mark them as important, and mark the remaining positions as unimportant, obtain the largest k indices, and generate a mask matrix depending on the attention score matrix , each row corresponds to a query, each column corresponds to a block, fill 1 at the selected k index positions, and fill at the remaining positions to make the weights of non-Top-k positions approach 0.
[0131] The attention score of the important block is expressed as: 。
[0132] Wherein, k represents the top k blocks with the highest importance.
[0133] The k important blocks screened by the mask matrix M are used to restore s original Tokens in each block and rearrange them in the order of their positions in the original sequence. Relative to the Token position corresponding to the th important block is After unfolding the original Tokens of all k blocks in order, the total number of Tokens is and 。
[0134] The restored key-value pairs will be used for attention calculation with the query vector as follows: 。
[0135] Wherein, and are the sets of the restored key and value , is the restored key of the i-th Token in the j-th important block, is the restored value of the i-th Token in the j-th important block.
[0136] The highest attention completes feature refinement by screening important blocks and selecting strongly relevant features. It only calculates for a few Tokens within the important blocks, and the computational complexity of this step is much smaller than that for the original input sequence. While maintaining the computational complexity, through the sparse selection of k << m, the complexity is further optimized to , which can greatly reduce computational redundancy while retaining key fine-grained features.
[0137] S64. Combine the aggregated block and the highest attention block through a gating mechanism to obtain the global refined feature.
[0138] Specifically, the present invention creates a learnable gating mechanism on the pathways of aggregated attention and highest attention, and combines them through the information gate of the sigmoid activation function to dynamically adjust the information flow. This mechanism can learn and identify the importance of each pathway in the data, map the value to the range of 0 to 1, and adaptively determine the amount of information passing through.
[0139] The output of the final SAT is composed of the results of aggregated attention and the highest attention, expressed as: .
[0140] .
[0141] Wherein, is the output aggregated feature, is the sigmoid activation function, represents the output of the aggregated attention, is the output of the highest attention.
[0142] Through the calculation of their weights, local and global perception of EEG information is realized. The hierarchical attention mechanism captures global patterns through aggregated attention and restores key local details through the highest attention, achieving a balance between computational efficiency and feature integrity. By adjusting the sparsity ratio r and the selection coefficient k, the computational amount and feature granularity of the model can be flexibly controlled, which is particularly suitable for processing high-dimensional temporal EEG signals.
[0143] S7. Input the global refined features into the classifier to complete the training of the model and obtain a pre-trained model that can be used for real-time prediction.
[0144] The classifier module receives the features processed by SAT, flattens and reduces the dimensions of the features using a flattening layer, converts the multi-dimensional features into one-dimensional for feature integration. Subsequently, it is processed through two fully connected layers, and a dropout layer is inserted in the middle to reduce the risk of overfitting. Then these features pass through the softmax function to calculate the prediction probability of each class, thereby completing the training of the model and obtaining a pre-trained model that can be used for real-time prediction.
[0145] The model is trained through steps S1 to S7, and a pre-trained model that can be used for real-time prediction is obtained. In this embodiment, the classifier is trained using the cross-entropy loss function, which can quantify the difference between the model's predicted probability distribution and the true label, and minimize this difference through the optimization process.
[0146] S8. Obtain the real-time data of the EEG signal. In this embodiment, the real-time data needs to be preprocessed. In this embodiment, step S8 is the preprocessing of the real-time data before the prediction stage. Step S8 specifically includes steps S81 to S83.
[0147] S81. Store the collected real-time EEG data stream through a buffer.
[0148] S82. Extract real-time segment data from the buffer through a sliding time window.
[0149] S83. Preprocess the real-time segment data; wherein, the preprocessing steps for real-time data have one less data augmentation process than those for the offline data, and the baseline correction operation is based on the overall segment mean.
[0150] Specifically, real-time EEG signals are obtained through a real-time data stream, and data is extracted through a sliding time window for preprocessing to obtain preprocessed EEG signals. The goal of real-time EEG signal processing is to achieve precise control of external devices through a brain-computer interface. Different from the preprocessing in the offline training stage, the real-time process completely relies on the subject's active imagination without visual or auditory cues, and continuous imagination incentives are provided throughout the process without rest segment data. Therefore, in real-time preprocessing, there is no need to perform the segmentation operation for the incentive interval, and the baseline correction is also adjusted from being based on the mean of the first 300 ms rest segment data of the segment to being based on the overall segment mean. In addition, the real-time preprocessing stage does not involve data augmentation operations, and the remaining steps are in the same order as those in the preprocessing steps of the offline training stage.
[0151] Specifically, as Figure 7 and Figure 8 shown, the present invention sets up a buffer for the real-time data stream to store all the collected real-time data. The sampling frequency of the device is 256 Hz, and a sliding time window with a size of 1024 sampling points is set in the buffer, and the window slides a distance of 256 sampling points. Segment data is extracted from the buffer through the sliding time window for real-time preprocessing. The setting of the buffer ensures the integrity and continuity of the window data during each prediction, and at the same time provides connection support for the subsequent access of data.
[0152] To address the impact caused by device and data transmission delays, instead of using absolute time as the division basis, fixed sampling points related to the sampling frequency are used as the window positioning flag. The method of selecting data using a sliding time window has the following advantages: First, by using fixed sampling points as the window division basis, prediction errors caused by device delays or time drifts can be effectively avoided. Second, the design of the sliding time window can make full use of the data continuity, improve data utilization rate and reduce system latency. Finally, the gradual update mechanism of the sliding time window can smooth the prediction results, enhance the real-time performance and stability of the system, thus ensuring precise control of external devices.
[0153] S9. Input the data after real-time preprocessing into the pre-trained model to obtain the decoding result of the real-time EEG signal. Preferably, step S9 specifically includes steps S91 to S93.
[0154] S91. At the first prediction, after the buffer accumulates a complete window of data, use the pre-trained model to conduct the prediction.
[0155] S92. For each newly collected amount of data on the sliding distance, use the pre-trained model to perform predictions on the data within this window.
[0156] S93. The prediction result of each time will be mapped into a control instruction in real time.
[0157] Specifically, as Figure 7 and Figure 8 shown, the sliding time window will be set according to the window described in step S8. At the first prediction, when a complete window of data has been accumulated in the buffer, the system will perform real-time preprocessing on the window data, and then perform predictions through the pre-trained model obtained in the offline training stage. The window data used for prediction will be retained in the buffer. Then, whenever a new amount of data on the sliding distance is collected, the sliding time window will intercept this new data, and combine it with the data retained in the buffer after the previous round of prediction to form a new window of data. The system will immediately process the new window data and perform predictions on the preprocessed data within the window through the pre-trained model. Finally, all prediction results will be immediately mapped into control instructions and sent to the target device.
[0158] The present invention provides a new EEG signal decoding network: APCformer, which performs excellently in EEG classification and recognition tasks and effectively solves some defects of traditional networks in feature extraction. This network is a fusion network structure with multi-scale feature interaction, gradually refining features in a hierarchical manner. Spatiotemporal convolution and the adaptive feature recalibration module capture roughly coarse-grained features for the network. The features are refined through the position-aware enhancement module, and position encoding information is input to strengthen the internal connection between features. The sparse information aggregation Transformer module strengthens long-range dependencies and local associations, balances local and global correlations, and comprehensively improves the depth and breadth of EEG data analysis.
[0159] Embodiment 2. The present invention provides an EEG signal decoding device based on an aggregation-aware enhanced convolutional Transformer network, which includes a signal acquisition module, a spatiotemporal convolution module, a recalibration module, a fusion module, a position-aware module, an information aggregation module, and a classifier module.
[0160] The offline data acquisition module is used to acquire offline data of EEG signals.
[0161] The spatiotemporal convolution module is used to extract multi-scale shallow local features through the spatiotemporal convolution module according to the offline data.
[0162] The recalibration module is used to perform interaction sharing and strengthen key features on the multi-scale shallow local features through the adaptive feature recalibration module to obtain multiple recalibrated features.
[0163] A fusion module for fusing multiple recalibrated features to obtain a first fused feature.
[0164] A position perception module for inputting the first fused feature into a position perception enhancement module to extract deep fine-grained features through parallel enhanced convolutions and adaptively encoding the deep fine-grained features to obtain position perception enhanced features.
[0165] An information aggregation module for extracting long-range dependencies and local associations through a sparse information aggregation Transformer module to obtain globally refined features.
[0166] A classifier module for inputting the globally refined features into a classifier to complete the training of the model and obtain a pre-trained model that can be used for real-time prediction.
[0167] A real-time data acquisition module for acquiring real-time data of EEG signals; A real-time decoding module for inputting the real-time data into the pre-trained model to obtain the decoding result of the real-time EEG signal.
[0168] Embodiment 3. The present invention provides an EEG signal decoding device based on an aggregation-aware enhanced convolutional Transformer network, which is characterized by including a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement an EEG signal decoding method based on an aggregation-aware enhanced convolutional Transformer network as described in any paragraph of Embodiment 1.
[0169] Embodiment 4. The present invention provides a computer-readable storage medium, which is characterized in that the computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute an EEG signal decoding method based on an aggregation-aware enhanced convolutional Transformer network as described in any paragraph of Embodiment 1.
[0170] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the part of the module, program segment, or code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0171] In addition, each functional module in various embodiments of the present invention may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.
[0172] If the above functions are implemented in the form of software function modules and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs that can store program codes. It should be noted that in the present invention, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.
[0173] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise.
[0174] It should be understood that the term "and / or" used in the present invention is merely an associative relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally represents an "or" relationship between the associated objects before and after.
[0175] Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" may be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".
[0176] The "first / second" mentioned in the embodiments is only to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged in a specific order or sequence when allowed. It should be understood that the objects distinguished by the "first / second" can be interchanged appropriately so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.
[0177] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An EEG signal decoding method based on aggregated perception enhanced convolutional Transformer network, characterized in that: Include: Obtain offline data of EEG signals; Extracting multi-scale shallow local features through a spatiotemporal convolution module according to the offline data; Interactively sharing the multi-scale shallow local features and strengthening key features through an adaptive feature recalibration module to obtain multiple recalibrated features; Fusing the recalibrated multiple features to obtain a first fused feature; The first fused feature is input into the location-aware enhancement module to extract deep fine-grained features through parallel enhanced convolution, and the deep fine-grained features are adaptively encoded to obtain location-aware enhanced features; The location-aware enhanced features are refined through a sparse information aggregation Transformer module to extract long-range dependencies and local associations to obtain global refined features; Inputting the global refined features into a classifier to complete the training of the model and obtain a pre-trained model that can be used for real-time prediction; Get real-time data of EEG signals; The real-time data is input into the pre-trained model to obtain the decoding result of the real-time EEG signal.
2. According to claim 1, the EEG signal decoding method based on the aggregated perception enhanced convolutional Transformer network is characterized in that: Obtain offline data of EEG signals, including: Obtain the stored EEG signal data, eliminate the interference of 50Hz power frequency and environmental noise, and filter out the EEG signal in the key frequency band of 0.5-30Hz; The continuous EEG signal is segmented and processed, the rest segment data is removed, and only the MI segment data is retained; The mean of the rest segment data 300 milliseconds before the MI segment data was used as the baseline, and each MI segment data was independently baseline corrected to eliminate the potential influence of baseline shift; After standardizing the data, the artifacts that are confused in the EEG signal are removed to obtain the preprocessed EEG data; The training data is segmented and reconstructed according to the label division, and random values following the Gaussian distribution are added to simulate batch data. The newly generated data is mixed with the original batch data to obtain the EEG data after data enhancement, and then they are input into the model to learn features. Get real-time data of EEG signals, including: The real-time EEG data stream collected is stored in a buffer; Extract real-time segment data from the buffer using a sliding time window; The real-time segment data is preprocessed; wherein the preprocessing step of the real-time data lacks data enhancement processing compared to the preprocessing step of the offline data, and the baseline correction operation is based on the overall mean of the segment.
3. The EEG signal decoding method based on the aggregated perception enhanced convolutional Transformer network according to claim 1 is characterized in that: The spatiotemporal convolution module is provided with three branches; each branch is provided with a temporal convolution layer, a spatial convolution layer, a batch normalization layer, an ELU activation function layer, and an average pooling layer in sequence; the temporal convolution layer of the first branch is 16 (1, ) convolution kernel; the second branch has 16 temporal convolution layers (1, ) convolution kernel; the third branch has 16 temporal convolution layers (1, ) convolution kernel; Fs is the sampling frequency; the spatial convolution layer uses 32 convolution kernels of size (C, 1), where C is the number of channels; during model training, a random dropout layer is also set to reduce the risk of overfitting; According to the EEG signal, multi-scale shallow local features are extracted through the spatiotemporal convolution module, specifically including: The EEG signal is input into three branches of the spatiotemporal convolution module respectively to obtain multi-scale shallow local features; The three branches of the spatiotemporal convolution module perform convolution operations through the temporal convolution layer and the spatial convolution layer respectively to extract features; and the features extracted by the convolution operation are input into the batch normalization layer, the ELU activation function layer and the average pooling layer connected in sequence.
4. The EEG signal decoding method based on the aggregated perception enhanced convolutional Transformer network according to claim 1, characterized in that: The multi-scale shallow local features are interactively shared and key features are enhanced through an adaptive feature recalibration module to obtain multiple recalibrated features, including: The features extracted from each branch of the spatiotemporal convolution module are superimposed with an interactive connection structure to obtain multiple interactive features with the same number of branches as the spatiotemporal convolution module; Through the convolutional attention mechanism, the features of multiple interactive features along the channel and space are calculated separately to obtain multiple recalibrated features.
5. The EEG signal decoding method based on the aggregated perception enhanced convolutional Transformer network according to claim 1 is characterized in that: The first fusion feature is input into the location-aware enhancement module to extract deep fine-grained features through parallel enhanced convolution, and the deep fine-grained features are adaptively encoded to obtain location-aware enhanced features, specifically including: According to the first fusion feature, further extract features in the time direction through parallel 1x3 convolution and 1x7 convolution, and the extracted features are processed by batch normalization and then feature fusion is performed; The features after feature fusion are processed using the ELU activation function and the average pooling layer; a random dropout layer is also set in the training stage; Through a skip connection, the features processed by the average pooling layer are added to the first fused features; The added features are dimensionally transformed and then adaptively encoded to obtain location-aware enhanced features.
6. The EEG signal decoding method based on the aggregated perception enhanced convolutional Transformer network according to any one of claims 1 to 5, characterized in that: The location-aware enhanced features are refined through the sparse information aggregation Transformer module to extract long-range dependencies and local associations and obtain global refined features, including: Dividing the location-aware enhanced feature into blocks through a sliding window to obtain a plurality of blocks; Average multiple blocks separately, perform average calculation on consecutive tokens in the block and aggregate them into a single block representation to obtain the aggregated block; Through the highest attention mechanism, the attention scores of each aggregated block are calculated, k important blocks are screened out, the original tokens of the important blocks are restored, and the highest attention blocks are obtained; The aggregated block and the highest attention block are combined through a gating mechanism to obtain the global refined features.
7. The EEG signal decoding method based on the aggregated perception enhanced convolutional Transformer network according to any one of claims 1 to 5, characterized in that: The global refined features are input into the classifier to complete the training of the model and obtain a pre-trained model that can be used for real-time prediction, which specifically includes: Use the flattening layer to flatten and reduce the dimension of features, converting multi-dimensional features into one dimension for feature integration; The flattened features are processed through two fully connected layers; in the training phase, random dropout layers are inserted into the two fully connected layers; The features processed by the fully connected layer are passed through the softmax function to calculate the prediction probability of each category, so as to complete the training of the model and obtain a pre-trained model that can be used for real-time prediction; Inputting the real-time preprocessed data into the pre-trained model to obtain the decoding result of the real-time EEG signal, specifically including: During the first prediction, after a complete window of data is accumulated in the buffer, the pre-trained model is used to make predictions. Every time a new amount of data of a sliding distance is collected, the pre-trained model is used to predict the data in this window; Each prediction result is mapped into control instructions.
8. An EEG signal decoding device based on an aggregated perception enhanced convolutional Transformer network, characterized in that: Include: An offline data acquisition module, used to acquire offline data of EEG signals; A spatiotemporal convolution module, used to extract multi-scale shallow local features through the spatiotemporal convolution module according to the offline data; A recalibration module, used for interactively sharing the multi-scale shallow local features and strengthening key features through an adaptive feature recalibration module to obtain a plurality of recalibrated features; A fusion module, used for fusing the recalibrated multiple features to obtain a first fused feature; A position perception module, used for inputting the first fusion feature into the position perception enhancement module to extract deep fine-grained features through parallel enhanced convolution, and adaptively encoding the deep fine-grained features to obtain position perception enhanced features; The information aggregation module is used to extract long-range dependencies and local associations through the sparse information aggregation Transformer module to obtain global refined features; A classifier module, used for inputting the global refined features into a classifier to complete the training of the model and obtain a pre-trained model that can be used for real-time prediction; A real-time data acquisition module, used to acquire real-time data of EEG signals; The real-time decoding module is used to input the real-time data into the pre-training model to obtain the decoding result of the real-time EEG signal.
9. An EEG signal decoding device based on an aggregated perception enhanced convolutional Transformer network, characterized in that: It includes a processor, a memory, and a computer program stored in the memory; the computer program can be executed by the processor to implement an EEG signal decoding method based on an aggregated perception enhanced convolutional Transformer network as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the EEG signal decoding method based on the aggregated perception enhanced convolutional Transformer network as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Brain activity state recognition method and device, equipment and storage medium
CN115736948A
Electroencephalogram cognitive load assessment method and system based on multi-feature-domain attention network
CN116584955A
Multi-modal gesture recognition method based on data enhancement electroencephalogram and electromyographic signal fusion
CN118445747A
Time Domain-Based Methods for Noninvasive Brain-Machine Interfaces
US20140058528A1
Cited By
Pancerous cancer cell detection method and system and storage medium
CN120598959A