Electroencephalogram attention classification model and method based on spatial-temporal feature fusion
By adopting a model of spatiotemporal and spatial feature fusion in the classification of EEG signals, combining convolutional networks and attention mechanisms, the problems of low classification accuracy and incomplete feature extraction in the existing technology are solved, and a more efficient classification of EEG attention states is achieved.
Patent Information
- Application Number
- CN202510203559.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-24
AI Technical Summary
The prior art has problems in the classification of EEG signals with low classification accuracy, incomplete feature extraction, and limited ability to handle complex EEG signals.
An EEG attention classification model based on spatiotemporal feature fusion is proposed. Through the combination of convolutional blocks, graph convolution networks, timing convolution networks, attention feature fusion networks and classifiers, the temporal and spatial characteristics of EEG signals are captured, and the feature fusion ability is enhanced through attention mechanisms.
It significantly improves the accuracy of EEG attention classification, automatically identify key feature channels by calculating attention weights, enhances attention to space-time features, and optimizes classification performance.
Smart Images

Figure CN120154342A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electroencephalogram signal processing, and particularly relates to an electroencephalogram attention classification model based on spatio-temporal feature fusion and a method thereof. Background Art
[0002] Electroencephalography (EEG) is a key technology for studying the functions and states of the brain by recording the electrical activities of the cerebral cortex. Due to its non-invasive nature, high temporal resolution, and relatively low cost, EEG technology has been widely applied in the fields of neuroscience, clinical diagnosis, Brain-Computer Interface (BCI), and cognitive science. In the research of attention detection and classification, EEG signals have become an important research means because they can reflect the dynamic changes of brain activities in real time.
[0003] In recent years, the development of deep learning technology has brought new opportunities for the automatic classification of EEG signals. Traditional EEG classification methods usually rely on manual feature extraction and traditional machine learning models, but these methods face multiple challenges when dealing with complex EEG signals. For example, traditional methods are usually sensitive to noise, unable to fully utilize the potential information in multi-channel signals, and have low feature extraction efficiency. In addition, when facing complex EEG signal data, these methods are difficult to capture the complex interactions between temporal dynamic changes and spatial features, resulting in low recognition accuracy. With the rapid development of deep learning technology, models based on convolutional neural networks (CNNs) and recurrent neural networks (RNNs) have shown strong advantages in automatic feature extraction and multi-level information capture, greatly improving the performance of EEG signal classification. However, since EEG signals are essentially multi-channel spatio-temporal sequence data, it is difficult to comprehensively characterize the attention state of the brain only relying on a single feature extraction method. In existing research, how to further enhance the spatio-temporal feature fusion ability while ensuring the efficiency of the model, and how to effectively combine the attention mechanism to improve the attention classification effect are still urgent problems and research challenges to be solved. Summary of the Invention
[0004] The present invention aims to provide an electroencephalogram attention classification model and method based on spatio-temporal feature fusion to effectively solve the problems of low classification accuracy, incomplete feature extraction, and limited ability to process complex EEG signals in the prior art.
[0005] In a first aspect, a brain - electroencephalogram (EEG) attention classification model based on spatio - temporal feature fusion is proposed, which includes: a convolutional block configured to perform convolutional operations on EEG signals in both time and space dimensions; a graph convolutional network configured to perform graph convolutional operations on the feature maps output by the convolutional block to capture the spatial dependence relationships between electrodes; a temporal convolutional network configured to perform dilated causal convolutions on the feature maps output by the graph convolutional network in the temporal dimension to capture the dynamic changes of EEG signals over time; an attention feature fusion network configured to perform preliminary fusion on the spatial feature maps extracted by the graph convolutional network and the temporal feature maps extracted by the temporal convolutional network to form an initial fusion feature map F fusion , using an attention weight α to weight the initial fusion feature map F fusion to obtain a weighted feature map F attention , and performing weighted summation on the initial fusion feature map F fusion and the weighted feature map F attention to obtain a final fusion feature F final ; and a classifier configured to map the final fusion feature F final to a vector of the number of classes through a fully - connected layer, then calculate class probabilities through a Softmax function, and finally output an attention classification result.
[0006] In some examples, the convolutional block includes: a first two - dimensional convolutional layer configured to perform independent convolutional operations on EEG signals of each channel to capture the dynamic characteristics across multiple time points and generate an initial time feature map reflecting the temporal correlation between channels; a first batch normalization layer configured to perform batch normalization on the initial time feature map generated by the first two - dimensional convolutional layer; a second two - dimensional convolutional layer configured to perform convolution on the initially batch - normalized time feature map in the spatial dimension to extract the relationships between different channels and generate an initial spatial feature map; a second batch normalization layer configured to perform batch normalization on the initial spatial feature map; an activation function configured to apply a non - linear transformation to the initially batch - normalized spatial feature map; a pooling layer configured to perform average pooling operations on the initial spatial feature map after passing through the activation function; and a depth - wise separable convolutional layer configured to perform depth - wise separable convolutional operations on the initially average - pooled spatial feature map.
[0007] In some examples, the depth - wise separable convolutional layer includes: a depth convolution that performs spatial convolution on each input channel while maintaining the same number of channels as the input; and a point - wise convolution that performs cross - channel feature fusion.
[0008] In some examples, after the depth convolution performs spatial convolution on each input channel, the resulting feature maps of each channel are sequentially passed through batch normalization and ReLU non - linear activation.
[0009] In some examples, the ReLU activation function is applied after pointwise convolution.
[0010] In some examples, an average pooling layer is configured after the ReLU activation function, and the average pooling layer performs spatial downsampling to reduce the spatial dimension of the feature map input to the graph convolutional network.
[0011] In some examples, the initial fused feature map F fusion is subjected to global average pooling to obtain a global feature vector, and the global feature vector is input into a fully connected layer. After passing through the non-linear activation function ReLU, the attention weight α is generated.
[0012] In a second aspect, an electroencephalogram (EEG) attention classification method is proposed, including: collecting the EEG signals of a subject; inputting the EEG signals into the EEG attention classification model based on spatio-temporal feature fusion, and the model outputs an attention classification result.
[0013] In a third aspect, a computer system is proposed, including: a processor; a memory including one or more computer program modules; wherein, the one or more computer program modules are stored in the memory and configured to be executed by the processor, and the one or more computer program modules include instructions for implementing the EEG attention classification method.
[0014] In a fourth aspect, a computer-readable storage medium is proposed for storing non-transitory computer-readable instructions, which can implement the EEG attention classification method when executed by a computer.
[0015] The present invention can effectively extract and fuse spatio-temporal features in EEG signals. Through the synergistic effect of each part of the model, key features are highlighted while redundant information is suppressed, thereby significantly improving the accuracy of EEG attention classification.
[0016] By calculating the attention weight, the present invention automatically identifies and gives higher weights to the feature channels that are most important for the classification task, thereby enhancing the attention to key spatio-temporal features and further optimizing the classification performance. This mechanism enables the model to more accurately focus on the features related to the task, improving the accuracy of EEG attention classification.
[0017] The EEG attention classification model of the present invention adopts a reasonable and modular design, which is convenient for expansion and adjustment in practical applications and can flexibly adapt to different application scenarios and requirements. By applying deep learning to EEG signal analysis and using artificial intelligence algorithms for feature extraction and classification, the present invention not only improves the classification accuracy but also makes it easy to be popularized and applied in fields such as daily mental health monitoring and attention state assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a flowchart of an electroencephalogram (EEG) attention classification method based on spatio-temporal feature fusion according to an embodiment of the present invention.
[0019] Figure 2 is a block diagram of an EEG attention classification model based on spatio-temporal feature fusion according to an embodiment of the present invention.
[0020] Figure 3 is a structural diagram of an EEG attention classification model based on spatio-temporal feature fusion according to an embodiment of the present invention.
[0021] Figure 4 is an electrode distribution diagram for collecting EEG signals according to an embodiment of the present invention. Detailed implementation manners
[0022] The present invention can achieve efficient and accurate classification of attention states, and is particularly suitable for application scenarios such as education and medical care that require real-time monitoring and evaluation of attention states. The core of the present invention lies in enhancing the ability of the classification model to capture multi-dimensional information through deep fusion of spatio-temporal features, thereby significantly improving the accuracy of EEG signal classification. Compared with traditional methods, the present invention can more comprehensively represent the attention state of the brain, and by introducing an attention mechanism, the model can focus on the most distinguishable feature signals, further improving the classification effect.
[0023] Figure 1 and Figure 2 and Figure 3 show the process of an EEG attention classification method based on spatio-temporal feature fusion. Next, we will combine Figure 1 and Figure 2 and Figure 3 to explain this method in detail.
[0024] 1. EEG signal acquisition
[0025] Collect multi-channel EEG signals of the subject in different attention states to provide a data basis for subsequent EEG classification. For example, when collecting the EEG signals of the subject, the sampling rate is set to 1000 Hz, and 32 channels covering the prefrontal, parietal, central, occipital, and temporal regions of the brain are selected, specifically including Fz, Fp1, Fp2, F3, F4, F7, F8, FC1, FC2, FC5, FC6, FT9, FT10, Cz, C3, C4, T7, T8, CP1, CP2, CP5, CP6, TP9, TP10, Pz, P3, P4, P7, P8, Oz, O1, and O2. Among these channels, Fz is used as the reference electrode, and the ground electrode (GND) is placed on the forehead of the subject, as shown in Figure 4 shown.
[0026] 2. EEG Signal Preprocessing
[0027] Preprocess the collected EEG signals to improve data quality and consistency and ensure the accuracy of subsequent analysis.
[0028] 2.1 Band - Pass Filtering
[0029] Filter the original EEG signals using a band - pass filter of 0.5 - 30 Hz to remove signals in the unwanted frequency range, including high - frequency EMG signals and low - frequency ECG signals, EDA signals, and respiratory signals.
[0030] 2.2 Independent Component Analysis (ICA) Denoising
[0031] Use independent component analysis (ICA) to remove electro - oculogram (EOG) signals from the original signals. First, decompose the filtered EEG signal X filtered into independent components. The calculation formula is as follows:
[0032] S = WX filtered ;
[0033] where S is the independent component matrix and W is the inverse matrix of the mixing signal matrix A found through the ICA algorithm.
[0034] Then identify the EOG signal components and remove these components from the independent component matrix S. The calculation formula is as follows:
[0035] S EEG = S - S EOG ;
[0036] where S EOG is the EOG signal component and S EEG is the remaining EEG component.
[0037] Finally, use the remaining independent component S EEG and the mixing signal matrix A to reconstruct the EEG signal. The calculation formula is as follows:
[0038] X clean = AS EEG ;
[0039] where X clean is the reconstructed / pre - processed EEG signal. Through this step, the noise signals are effectively removed, improving the accuracy of subsequent analysis and classification.
[0040] 3. EEG Attention Classification Model Based on Spatiotemporal Feature Fusion
[0041] The preprocessed EEG signals are input into an EEG attention classification model based on spatio-temporal feature fusion (hereinafter referred to as "model" or "classification model") for analysis and classification, so as to accurately identify the attention state of the tester.
[0042] As Figure 2 , Figure 3 shown, the classification model includes a convolutional block, a graph convolutional network, a temporal convolutional network, an attention feature fusion network, and a classifier.
[0043] 3.1 Convolutional Block
[0044] The convolutional block contains two two-dimensional convolutional layers and a depthwise separable convolutional layer, which are used to process the temporal and spatial features of EEG signals.
[0045] First, in the time dimension, independent operations are performed on the EEG signals of each channel. Specifically, a two-dimensional convolutional kernel of size (1, 64) is used to capture the dynamic characteristics spanning multiple time points. Through the sliding window mechanism, the convolutional kernel performs a convolutional operation on the time axis to generate a feature map reflecting the temporal correlation between channels. Subsequently, these temporal feature maps are batch-normalized to accelerate the model training process and stabilize the training effect.
[0046] Next, on the basis of batch normalization, the second two-dimensional convolutional layer is applied, which works in the spatial dimension this time, aiming to extract the relationships between different channels. The convolutional kernel size used here is (32, 1), which can cover the local areas of multiple channels, thereby capturing the interaction information between EEG signal channels and generating new spatial feature maps. These spatial feature maps are batch-normalized again and the ELU (Exponential Linear Unit) activation function is used to more effectively represent the spatial features. Note that the sizes of the convolutional kernels of the two two-dimensional convolutional layers are only exemplary.
[0047] Finally, an average pooling operation is performed on the spatially feature maps that have been batch-normalized. This step helps to further compress the feature representation, reduce the data dimension, and retain important spatial information.
[0048] Perform depthwise separable convolution operations on the pooled spatial feature map. First, perform spatial convolution on each input channel using a depthwise convolution that is independent in the spatial dimension (e.g., a convolution kernel of size 32×1), keeping the number of channels the same as the input. The feature map of each channel is sequentially passed through batch normalization and the ReLU non-linear activation. Next, perform cross-channel feature fusion through a 1×1 pointwise convolution, flexibly adjusting the number of channels using a learnable weight matrix, and apply the ReLU activation function after this convolution layer to enhance the non-linear expression ability. After feature fusion, use an average pooling layer for spatial downsampling to reduce the spatial dimension of the feature map. In practical applications, to prevent overfitting and improve the generalization ability of the model, the Dropout technique can be introduced between the average pooling layer of the depthwise separable convolution operation and the subsequent layer (such as the graph convolution layer), suppressing the over-reliance on specific features during the training process by randomly masking neurons.
[0049] 3.2. Graph Convolutional Network
[0050] Use two graph convolutional network layers to extract the spatial dependence of features.
[0051] Represent the feature corresponding to the feature map with reduced spatial dimension in the previous step as the node feature matrix H (0) , and the feature vector of each node represents the temporal feature of an electrode. Construct an adjacency matrix A based on the physical distance and signal correlation between electrodes, defining the connection structure in the graph. Where A ij represents the connection strength between node i and node j, and usually a Gaussian kernel function or other similarity measurement methods can be used to define the value of A ij . Then perform the first layer of graph convolution operation, and the calculation formula is as follows:
[0052]
[0053] Where: H (1) is the node feature matrix of the first layer; is the adjacency matrix after adding self-connections, and I is the identity matrix; is 's degree matrix, defined as W (0) is the trainable weight matrix, and σ(·) is the activation function ReLU.
[0054] Then input the output H (1) of the first layer into the second layer of graph convolution to further extract deeper spatial features, and the calculation formula is the same as above.
[0055]
[0056] The graph convolution operation aggregates node features through the adjacency matrix and degree matrix, which can capture the spatial dependencies between electrodes, thus better modeling the features of EEG signals.
[0057] 3.3 Temporal Convolutional Network
[0058] The temporal convolutional network is used to capture the dynamic changes over time in EEG signals.
[0059] Taking the spatial features extracted by the graph convolutional network as input, the first-layer dilated causal convolution is performed in the temporal dimension with a dilation rate set to 2. After the convolution operation, batch normalization and the ELU activation function are applied. To enhance the generalization ability of the model, the Dropout technique can be introduced in the first-layer dilated causal convolution. The second-layer dilated causal convolution is similar to the first layer, but the dilation rate is set to 4 to capture dependencies at different time scales. Similarly, batch normalization and the ELU activation function are applied. Subsequently, the input is directly added to the output of the second-layer dilated causal convolution to form a residual connection. This design ensures that the gradient can be effectively propagated in the deep network, alleviating the vanishing gradient problem. Finally, the ELU activation function is applied again to the output after the residual connection to further enhance the non-linear feature expression ability.
[0060] 3.4 Attentional Feature Fusion Network
[0061] The Attentional Feature Fusion (AFF) network is adopted to focus on the most important features.
[0062] First, the spatial feature map extracted by the graph convolutional network and the temporal feature map extracted by the temporal convolutional network are preliminarily fused to form the initial fused feature map F fusion . Then, global average pooling is performed on the initially fused feature map F fusion to obtain the global context information. The formula for global average pooling is as follows:
[0063]
[0064] where is the global feature vector, C fusion is the number of channels of the fused feature, T represents the total number of time steps, t is the time step, and c is the channel.
[0065] The global feature vector z is input into a fully connected layer, and after passing through the non-linear activation function ReLU, the attention weight α is generated. The calculation formula is as follows:
[0066] α = ReLU(Wz + b);
[0067] where α is the attention weight, and W and b are the learnable weight matrix and bias vector.
[0068] Weight the fused feature map F using the calculated attention weight α fusion to highlight important feature channels. The weighted feature map F attention is calculated as follows:
[0069] F attention (t, c) = α c ·F fusion (t, c);
[0070] where F attention retains the temporal information while enhancing the expression of key features.
[0071] Perform weighted summation on the original feature map and the attention-weighted feature map to obtain the final fused feature F final . The calculation formula is as follows;
[0072] F final = F fusion + λ·F attention ;
[0073] where F final is the final fused feature, and λ is a learnable fusion weight used to balance the importance of the original feature and the attention-weighted feature.
[0074] 4. Classification and Output
[0075] The final fused feature is mapped to a vector of the number of classes through a fully connected layer, and then the class probabilities are calculated through the Softmax function, and the classifier outputs the attention classification result.
[0076] Collect the EEG data of the tester, and use the trained model to detect the EEG signal of the tester to classify the attention state of the tester.
[0077] Collect the EEG data of the tester through an EEG acquisition device such as an EEG cap. It should be noted that it is necessary to also use 32 channels for data acquisition, and the sampling rate is set to 1000 Hz.
[0078] Use the trained model to detect the EEG signal of the tester to classify the attention state of the tester.
[0079] The present invention also provides an embodiment of a computer system. This computer system is a broad concept that not only covers traditional computers but also includes external devices related to brain-computer interfaces. The memory is used to store non-transitory computer-readable instructions (such as one or more computer program modules). The processor is responsible for executing these instructions; when the processor runs these non-transitory computer-readable instructions, it can implement one or more steps in the aforementioned EEG attention classification method.
[0080] The memory and the processor can be interconnected through a bus system or other forms of connection mechanisms to ensure the effective transmission of data and control signals. This design allows for efficient instruction execution and support for complex computing tasks, such as the specific implementation of the EEG attention classification method.
[0081] For example, the processor can be a central processing unit (CPU), a graphics processing unit (GPU), or other forms of processing units with data processing capabilities and / or program execution capabilities. For example, the central processing unit (CPU) can be of the X86 or ARM architecture, etc. The processor can be a general-purpose processor or a dedicated processor, and can control other components in the computer to perform the desired functions.
[0082] For example, the memory can include any combination of one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory can include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer program modules can be stored on the computer-readable storage media, and the processor can run one or more computer program modules to implement various functions of the computer.
[0083] The present invention also provides a computer-readable storage medium for storing non-temporary computer-readable instructions, which can implement one or more steps in the above-mentioned EEG attention classification method when executed by a computer. When the EEG attention classification model and method provided by the embodiments of the present invention are implemented in software and sold or used as an independent product, they can be stored in a computer-readable storage medium. The relevant description of the storage medium can refer to the corresponding description of the memory in the computer above, and will not be elaborated here.
Claims
1. An EEG attention classification model based on spatiotemporal feature fusion, characterized in that: Include: A convolution block configured to perform a convolution operation on the EEG signal in time and space dimensions; A graph convolutional network, which is configured to perform graph convolution operations on the feature maps output by the convolutional blocks to capture the spatial dependencies between electrodes; A temporal convolutional network, which is configured to perform dilated causal convolution on the feature map output by the graph convolutional network in the temporal dimension to capture the dynamic changes of EEG signals over time; The attention feature fusion network is configured to perform a preliminary fusion of the spatial feature map extracted by the graph convolution network and the temporal feature map extracted by the temporal convolution network to form an initial fusion feature map F fusion , use the attention weight α to the initial fusion feature map F fusion Weighted, get the weighted feature map F attention , for the initial fusion feature map F fusion And the weighted feature map F attention Perform weighted summation to obtain the final fusion feature F final ; as well as The classifier is configured to pass the final fusion feature F through a fully connected layer. final Map it to a vector of the number of categories, then calculate the category probability through the Softmax function, and finally output the attention classification result.
2. The EEG attention classification model based on spatiotemporal feature fusion according to claim 1 is characterized in that: A convolutional block contains: A first two-dimensional convolutional layer is configured to perform independent convolution operations on the EEG signal of each channel to capture dynamic characteristics across multiple time points and generate an initial time feature map reflecting the temporal correlation between each channel; a first batch normalization layer, configured to perform batch normalization on the initial temporal feature map generated by the first two-dimensional convolutional layer; A second two-dimensional convolutional layer is configured to perform convolution on the initial time feature map processed by batch normalization in the spatial dimension, extract the relationship between different channels, and generate an initial spatial feature map; A second batch normalization layer, configured to perform batch normalization on the initial spatial feature map; An activation function configured to apply a nonlinear transformation to the batch normalized initial spatial feature map; A pooling layer, which is configured to perform an average pooling operation on the initial spatial feature map after the activation function; as well as A depthwise separable convolutional layer is configured to perform a depthwise separable convolution operation on the initial spatial feature map after average pooling.
3. The EEG attention classification model based on spatiotemporal feature fusion according to claim 2 is characterized in that: The depthwise separable convolutional layer contains: Depthwise convolution, which performs spatial convolution on each input channel, keeping the number of channels the same as the input; and Point-wise convolution, which performs cross-channel feature fusion.
4. The EEG attention classification model based on spatiotemporal feature fusion according to claim 3 is characterized in that: After the depthwise convolution performs spatial convolution on each input channel, the feature map of each channel is batch normalized and ReLU nonlinearly activated in turn.
5. The EEG attention classification model based on spatiotemporal feature fusion according to claim 4 is characterized in that: The ReLU activation function is applied after the point-wise convolution.
6. The EEG attention classification model based on spatiotemporal feature fusion according to claim 5 is characterized in that: The average pooling layer is configured after the ReLU activation function, which performs spatial downsampling to reduce the spatial dimension of the feature map input to the graph convolutional network.
7. The EEG attention classification model based on spatiotemporal feature fusion according to claim 1 is characterized in that: The initial fusion feature map F fusion Perform global average pooling to obtain the global feature vector, input the global feature vector into a fully connected layer, and generate the attention weight α after passing through the nonlinear activation function ReLU.
8. A method for classifying EEG attention, characterized in that: The method comprises: collecting EEG signals of subjects; inputting the EEG signals into an EEG attention classification model based on spatiotemporal feature fusion as described in any one of claims 1 to 7, wherein the model outputs an attention classification result.
9. A computer system, characterized in that: include: processor; a memory including one or more computer program modules; Wherein, the one or more computer program modules are stored in the memory and configured to be executed by the processor, and the one or more computer program modules include instructions for implementing the EEG attention classification method described in claim 8.
10. A computer-readable storage medium for storing non-transitory computer-readable instructions, characterized in that: When the non-transitory computer-readable instructions are executed by a computer, the EEG attention classification method of claim 8 can be implemented.
Citation Information
Patent Citations
Attention mechanism-based action recognition method of adaptive graph convolutional network
CN113688765A
Skeleton action recognition method based on space-time adaptive feature fusion graph convolutional network
CN116665300A
Multi-modal feature fusion emotion recognition method based on gating cross-attention mechanism
CN117370828A
Space-time diagram multi-dimensional information fusion electroencephalogram domain generalization decoding method and device and storage medium
CN117521009A
Classification method for motor imagery electroencephalogram signals
CN117860271A
Cited By
Epileptic seizure detection system based on double-branch space-time diagram neural network
CN121647615A