Emotion recognition method based on space-time multi-scale attention convolutional neural network

By introducing a space-time multi-scale attention convolution neural network into emotion recognition, combining the time and space feature extraction modules, learning asymmetric representation between the left and right hemispheres, the problem of ignoring spatial dimension information in the existing technology is solved, and a more efficient emotion recognition effect is achieved.

CN120162652APending Publication Date: 2025-06-17SHIHEZI UNIVERSITY
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510375412.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art ignores the spatial dimension information between different electrode positions in emotion recognition, and the model parameters are large, resulting in poor generalization and migration of the system.

Method used

A method of emotion recognition based on space-time multi-scale attention convolutional neural network is proposed. A lightweight convolutional neural network with dual-stream space-time feature construction layer, mixed attention mechanism layer, high-order fusion layer and classification layer is proposed. Combined with the time and space feature extraction module, asymmetric representation between the left and right hemispheres are learned, and feature extraction is enhanced through attention mechanism.

Benefits of technology

It improves the accuracy and efficiency of emotion recognition, reduces the dependence on manual feature engineering, improves the degree of automation and robustness of the system, and reduces the number of parameters of the model, enhancing generalization ability and migration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162652A_ABST
    Figure CN120162652A_ABST
Patent Text Reader

Abstract

The invention discloses an emotion recognition method based on a space-time multi-scale attention convolutional neural network, and the method comprises the steps: collecting electroencephalogram data of a subject, and carrying out the preprocessing of the electroencephalogram data, so as to obtain electroencephalogram (EEG) data containing a space dimension and a time dimension; constructing a lightweight convolutional neural network comprising a double-flow spatial-temporal feature construction layer, a mixed attention mechanism layer, a high-order fusion layer and a classification layer; wherein the double-flow spatial-temporal feature construction layer comprises a temporal feature extraction module and a parallel spatial feature extraction module; the mixed attention mechanism layer combines a channel attention mechanism, a space attention mechanism and a self-attention mechanism; the high-order fusion layer is used for carrying out re-learning from the learned global convolution kernel to the representation of the local hemisphere convolution kernel; and recognizing the EEG data by using the trained lightweight convolutional neural network to obtain an emotion recognition result of the subject. The precision and efficiency of emotion recognition driven by electroencephalogram signals are improved by constructing a lightweight model with a small number of parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and relates to, but is not limited to, an emotion recognition method based on a spatio-temporal multi-scale attention convolutional neural network. Background Art

[0002] Emotion is a basic factor in human daily life. As a physiological state that responds to external stimuli, it not only affects decision-making, perception, human interaction, and intelligence, but is also closely related to people's health and has a significant impact on decision-making. Emotion recognition has been widely studied in many fields. For example, in the treatment of mental disorders such as generalized anxiety disorder and depression, emotion recognition plays a key role, especially in fields such as cognitive behavioral therapy, emotion regulation therapy, and emotion-focused therapy. In addition, emotion recognition also plays a key role in human-computer interaction, enabling computers or intelligent robots to understand and respond to the emotional states of users, thereby improving the personalized and natural user experience.

[0003] Electroencephalogram (EEG) signals record the electrical activities of the brain, usually including different frequency bands such as delta waves, theta waves, alpha waves, beta waves, and gamma waves. These EEG signals have been widely used in fields such as emotion recognition, fatigue monitoring, and neuropsychological disease research. The neural activities of the brain affect the emotional state by coordinating the central nervous system and the autonomic nervous system. The activities in different regions of the brain are closely related to specific emotional states. For example, high activity in the left prefrontal cortex may be related to positive emotions, while high activity in the right prefrontal cortex may be related to negative emotions. This regional specificity of EEG signals makes emotion recognition possible.

[0004] Traditional methods for emotion recognition using EEG signals are to manually design features based on EEG signal characteristics, such as by analyzing intrinsic mode functions or using wavelet transforms, and then use machine learning-based methods to classify the manually extracted features. Deep learning methods directly learn features from the signals, thus reducing the burden of manual feature extraction.

[0005] The above technologies have the following defects: Based on traditional machine learning methods, (1) they highly rely on the quality of manually extracted features, and in classification tasks, the generalization and transferability of the system are poor; in addition, the task of manually extracting features is cumbersome and time-consuming; (2) when manually extracting features, the current methods mainly focus on extracting temporal dimension features in EEG signals, while ignoring the spatial dimension information between different electrode positions.

[0006] The research methods based on deep learning have the following problems: (1) Currently, the research still mainly extracts features from the time dimension in EEG signals. Similar to traditional machine learning methods, it ignores the spatial dimension information between different electrode positions; (2) Many deep learning methods, such as deep belief networks and stacked autoencoders, are poor at processing two-dimensional data; when conventional convolutional neural networks convolve EEG signals, they have a large number of parameters. Summary of the Invention

[0007] In view of this, embodiments of the present invention provide an emotion recognition method based on a spatio-temporal multi-scale attention convolutional neural network, which at least solves the problems that the existing technology only extracts features from the time dimension of EEG signals, ignores the spatial dimension information between different electrode positions, and the problem of a large number of model parameters.

[0008] The technical solution of the embodiment of the present invention is specifically as follows:

[0009] Embodiments of the present invention provide an emotion recognition method based on a spatio-temporal multi-scale attention convolutional neural network, and the method includes:

[0010] Collect electroencephalogram data of a subject for preprocessing, and perform channel mapping according to the physical positions of the electrodes to obtain EEG data including spatial and time dimensions;

[0011] Construct a lightweight convolutional neural network including a two-stream spatio-temporal feature construction layer, a hybrid attention mechanism layer, a high-order fusion layer, and a classification layer; wherein, the two-stream spatio-temporal feature construction layer includes: a time feature extraction module for learning time-frequency feature representations using multi-scale one-dimensional time convolution kernels, and a parallel spatial feature extraction module for learning asymmetric representations between the left and right hemispheres using hemisphere convolution kernels and local hemisphere convolution kernels whose lengths correspond to the number of channels in the left and right hemispheres; the hybrid attention mechanism layer combines channel attention mechanism, spatial attention mechanism, and self-attention mechanism to enhance the ability of feature extraction and data processing; the high-order fusion layer is used for re-learning from the representations from the learned global convolution kernels to local hemisphere convolution kernels;

[0012] Use the trained lightweight convolutional neural network to identify the EEG data to obtain the emotion recognition result of the subject.

[0013] In some embodiments, the time feature extraction module includes multi-scale one-dimensional time convolution kernels, and the size of the i-th level one-dimensional time convolution kernel can be defined as where f s is the EEG signal sampling rate, i ∈ [1, 2,..., L], L is the level number of the one-dimensional time convolution kernel layer, and the proportionality coefficient δ when i takes 1, 2, 3 iThey are 0.25, 0.5, and 1.0 respectively; the proportionality coefficient corresponding to the one-dimensional temporal convolution kernel of the higher level is less than that corresponding to the one-dimensional temporal convolution kernel of the lower level.

[0014] In some embodiments, the output of the i-th level one-dimensional temporal convolution kernel in the temporal feature extraction module is defined as where X is the input EEG data, Conv1D() is a one-dimensional convolution operation with a convolution kernel size of (1, 1) and a stride of (1, 1), Φ L-ReLU () is the Leaky ReLU activation function, and AP() is the average pooling operation; the outputs of the one-dimensional temporal convolution kernels at each level are concatenated along the time dimension, and a batch normalization operation is added to obtain the output result of the temporal feature extraction module.

[0015] In some embodiments, the spatial feature extraction module has one-dimensional spatial convolution kernels of multiple scales, including: a global convolution kernel for learning global spatial information, and a hemisphere convolution kernel and a local hemisphere convolution kernel for extracting the relationship between the left and right hemispheres through a shared convolution kernel; the size of the one-dimensional spatial convolution kernel can be defined as where c is the total number of channels of the input EEG segment, j takes 1, 2, and 3 respectively to represent the global convolution kernel, the hemisphere convolution kernel, and the local hemisphere convolution kernel, and the corresponding δ j are 0.25, 0.5, and 1.0 respectively; the output of the j-th type of spatial convolution kernel is defined as

[0016]

[0017] where, n is the number of samples, s is the number of each type of one-dimensional spatial convolution kernel, c m is the number of channels after the m-th spatial convolution, f is the feature length after each spatial convolution operation, Z S is the multi-scale spatial representation generated by parallel multi-scale spatial convolution kernels for the input EEG data, Conv1D() is a one-dimensional convolution operation, and its convolution kernel size is (c, 1) for the global convolution kernel, (0.5×c, 1) for the hemisphere convolution kernel, and (0.25×c, 1) for the local hemisphere convolution kernel, Φ L-ReLU is the Leaky ReLU activation function, and AP is the average pooling layer.

[0018] In some embodiments, the method further includes: for the EEG data input to the hemispherical convolution kernel, deleting the electrode data of Fz, Cz, Pz, and Oz located at the midline position, and setting the channel arrangement order as [channel left ,channel right , where channel left represents the channels located in the left hemisphere, and channel right represents the channels located in the right hemisphere; the channel order on each hemisphere is rearranged so that each kernel weight is shared between electrode pairs symmetrically placed on the two hemispheres.

[0019] In some embodiments, the following operations are both adopted in the temporal convolution operation and the spatial convolution operation to reduce the parameters and the amount of computation: first, perform independent depth convolution on the input channels of each EEG data, and then use 1x1 convolution to perform feature combination between channels.

[0020] In some embodiments, in the self-attention mechanism part of the hybrid attention mechanism layer, use [Q, K, V] = z0U qkv in the self-attention mechanism part of the hybrid attention mechanism layer to generate the query vector Q, the key vector K, and the value vector V; in the formula, z0 is the input feature vector, and U qkv is the linear transformation matrix; calculate the self-attention output Attention(Q, K, V) through the following formula:

[0021]

[0022] Scale V = ||V -1 ||2, Scale K = ||K -1 ||2;

[0023] In the formula, D is the dynamic scaling matrix, ⊙ represents the Hadamard product, U copy and U sum are the linear transformation matrices; Scale V is the L2 norm of V, and Scale K is the L2 norm of K.

[0024] In some embodiments, in the hybrid attention mechanism layer, add the improved Convolutional Block Attention Module (CBAM) after the spatial features output by the two-stream spatio-temporal feature construction layer. This improved CBAM module combines the channel attention mechanism and the spatial attention mechanism, takes the feature map F output by the spatial feature extraction module as the input, and sequentially obtains the channel attention map M c(F) and spatial attention map M s (F'), and uses a convolutional layer as the shared network to replace the shared MLP layer in the original CBAM, and outputs the final feature representation; the convolutional layer consists of a convolution including a hidden layer; where M c (F) and M s (F) are calculated respectively by the following formulas:

[0025]

[0026]

[0027] In the formula, AvgPool(F) and MaxPool(F) respectively represent the average pooling operation and the max pooling operation on the feature map F output by the spatial feature extraction module, and AvgPool(F') and MaxPool(F') respectively represent the average pooling operation and the max pooling operation on the feature map F’ output by the channel attention; and are respectively the average pooling feature and the max pooling feature in the channel attention; and are respectively the average pooling feature and the max pooling feature in the spatial attention; W0 and W1 are both learnable weight matrices, f k represents the convolution operation, and σ() is the Sigmoid function.

[0028] In some embodiments, the advanced fusion layer is used to fuse the learning information from the global convolution kernel, the hemisphere convolution kernel and the local hemisphere convolution kernel along the spatial dimension by using a one-dimensional convolutional layer with a kernel size of (3,1), with the time feature as an additional feature input, and the finally learned spatio-temporal feature representation Z of the global hemisphere fusion fusion is generated by the following formula:

[0029] Z fusion = GAP(f bn (AP(Φ L-ReLu (Conv1D(Z,(3,1))))));

[0030] Z = Attention(Q,K,V)+F″+F;

[0031] Among them, Z represents the concatenation of the self-attention output Attention(Q,K,V), the feature map F output by the spatial feature extraction module, and the spatial feature F” output by F processed by CBAM, f bn is the batch normalization function, GAP() represents the global average pooling layer, Φ L-ReLUIt is the LeakyReLU activation function, AP is the average pooling layer, and Conv1D() represents a one-dimensional convolution operation with a convolution kernel size of (3,1);

[0032] Z fusion is input into the fully connected layer and activated by the softmax function. The final output OutPut can be calculated as follows:

[0033] OutPut = Φsoftmax(W'Φdp(ΦReLU(W(Γ(Zfusion)) + b)) + b');

[0034] OutPut = Φ softmax (W′Φ dp (Φ ReLU (W(Γ(Z fusion )) + b)) + b′);

[0035] Among them, Γ is the squeeze operation, W and W' are trainable weight matrices, b and b' are bias terms, and Φ ReLU is the ReLU activation function applied to the output of the linear transformation, and Φ dp is a dropout operation used to prevent overfitting, and φ softmax is the softmax function used to convert the output into a probability distribution.

[0036] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:

[0037] In the embodiments of the present invention, a method for emotion recognition based on a lightweight spatio-temporal multi-scale attention convolutional neural network is proposed. This model can not only extract features in the time dimension of data but also combines the principle of brain asymmetry in neuroscience research. In the design of the convolution kernel, the brain emotion asymmetry is introduced. By using a hemisphere convolution kernel with a length corresponding to the number of channels in the left and right hemispheres and a local hemisphere convolution kernel with a length corresponding to half of the number of channels in the left and right hemispheres, the hemisphere asymmetry patterns related to emotional states are extracted, which helps to improve the emotion recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings, where:

[0039] Figure 1 is a schematic flowchart of the method for emotion recognition based on a spatio-temporal multi-scale attention convolutional neural network provided by the embodiments of the present invention;

[0040] Figure 2 Schematic diagram of the overall architecture of the lightweight convolutional neural network provided by the embodiments of the present invention;

[0041] Figure 3 Schematic diagram of the electrode position processing of three spatial convolution kernels provided by the embodiments of the present invention; among them, part (a) is the global convolution kernel, part (b) is the hemisphere convolution kernel, and part (c) is the local hemisphere convolution kernel;

[0042] Figure 4 Schematic diagram of the principle of the self-attention mechanism provided by the embodiments of the present invention; among them, part (a) is the traditional self-attention mechanism, and part (a) is the improved self-attention mechanism;

[0043] Figure 5 Schematic diagram of the structure of the CBAM module provided by the embodiments of the present invention; among them, part (a) is the overall CBAM structure, part (b) is the channel attention structure, and part (c) is the spatial attention structure;

[0044] Figure 6 Statistical chart of the accuracy rate of arousal and valence for each subject on the DEAP dataset provided by the embodiments of the present invention;

[0045] Figure 7 Statistical chart of the F1 score of arousal and valence for each subject on the DEAP dataset provided by the embodiments of the present invention. Detailed implementation manners

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0047] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0048] It should be noted that the terms "first / second / third" involved in the embodiments of the present invention are only used to distinguish similar objects and do not represent a specific sorting of the objects. Understandably, "first / second / third" can be interchanged with a specific order or sequence under allowable circumstances, so that the embodiments of the present invention described here can be implemented in an order other than that illustrated or described here.

[0049] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms used here (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which the embodiments of the present invention belong. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0050] Figure 1 It is a schematic flowchart of an emotion recognition method based on a spatio-temporal multi-scale attention convolutional neural network provided for the embodiments of the present invention, as Figure 1 shown, the method at least includes the following steps:

[0051] Step S110, collect the electroencephalogram data of the subject for preprocessing, and perform channel mapping according to the physical positions of the electrodes to obtain EEG data including spatial dimension and time dimension.

[0052] Here, the EEG data is regarded as a two-dimensional time series, and its dimensions are respectively the spatial dimension (EEG electrode position distribution) and the time dimension. The time dimension reflects the change of brain activity over time. The spatial dimension shows the activation patterns of different functional regions through different electrode positions on the brain.

[0053] Step S120, construct a lightweight convolutional neural network including a two-stream spatio-temporal feature construction layer, a hybrid attention mechanism layer, a high-order fusion layer, and a classification layer.

[0054] Here, as Figure 2 shown, the two-stream spatio-temporal feature construction layer includes: a time feature extraction module for learning time-frequency feature representations using multi-scale one-dimensional time convolutional kernels, and a parallel spatial feature extraction module for learning asymmetric representations between the left and right hemispheres using hemisphere convolutional kernels and local hemisphere convolutional kernels whose lengths correspond to the number of channels in the left and right hemispheres; the hybrid attention mechanism layer combines a channel attention mechanism, a spatial attention mechanism, and a self-attention mechanism to enhance the ability of feature extraction and data processing; the high-order fusion layer is used for re-learning from the representations from the learned global convolutional kernels to the local hemisphere convolutional kernels.

[0055] The role of the dual-stream spatio-temporal feature construction layer is to extract features from the input data in different temporal and spatial dimensions. To extract more discriminative time-frequency representations, the temporal feature extraction module in the dual-stream spatio-temporal feature construction layer uses multi-scale one-dimensional temporal convolutional kernels to enrich the learned time-frequency feature representations. The spatial feature extraction module utilizes the findings of neuroscience research that the relationship between the activities of the left and right hemispheres of the brain and emotions is asymmetric, and uses a kind of hemispheric convolutional kernel and local hemispheric convolutional kernel to learn the asymmetric representations between the left and right hemispheres. In this way, the final model can capture the changes and patterns in the data at different scales.

[0056] The hybrid attention mechanism layer is a feature enhancement module that combines three mechanisms: self-attention, channel attention, and spatial attention, to improve the model's ability to understand complex data. The self-attention mechanism enables the model to capture long-range dependencies by calculating the similarities between input features, thus better understanding the context information. The channel attention mechanism weights each channel of the feature map, emphasizes important channels, and suppresses irrelevant channels to optimize the feature representation. The spatial attention mechanism models the spatial dimension of the feature map to identify key regions and enhance the model's sensitivity to spatial information. With this design, the model can more effectively identify and utilize the key information in the input data.

[0057] The advanced fusion layer fuses the learned information from the global convolutional kernel, hemispheric convolutional kernel, and local hemispheric convolutional kernel to learn advanced spatial representations, with the temporal features as additional feature inputs. The design of the advanced fusion layer makes the network more compact and suitable for real-time use.

[0058] This lightweight convolutional neural network aims to identify the most significant time-frequency-channel-specific EEG features corresponding to the user's emotional state.

[0059] In step S130, the trained lightweight convolutional neural network is used to identify the EEG data to obtain the emotion recognition result of the subject.

[0060] Here, the EEG data is input into the trained lightweight convolutional neural network, and the emotion recognition result is output through the classification layer. This result is a specific emotional dimension, such as arousal, valence, or dominance, etc.

[0061] An embodiment of the present invention proposes a lightweight convolutional neural network based on a spatio-temporal multi-scale attention mechanism to improve the accuracy and efficiency of emotion recognition driven by electroencephalogram (EEG) signals. By comprehensively extracting the temporal and spatial features of EEG signals and introducing the principle of cerebral emotional asymmetry, the model can capture the hemispheric asymmetry patterns related to emotional states, thereby improving the recognition accuracy. In addition, this method reduces the dependence on manual feature engineering, enhances the automation degree and robustness of the system. By constructing a lightweight model with fewer parameters, the generalization ability and transferability of the model are enhanced, enabling it to perform excellently under different individuals and environmental conditions. Finally, this method will verify its practical application potential in the fields of mental health intervention and human-computer interaction, promoting the realization of intelligent and personalized services.

[0062] In some embodiments, the temporal feature extraction module includes one-dimensional temporal convolutional kernels of multiple scales, and the size of the i-th level one-dimensional temporal convolutional kernel can be defined as where f s is the EEG signal sampling rate, i ∈ [1, 2,..., L], L is the number of levels of the one-dimensional temporal convolutional kernel layers, and the proportionality coefficients δ i are 0.25, 0.5, and 1.0 respectively when i takes 1, 2, and 3; the proportionality coefficient corresponding to the higher-level one-dimensional temporal convolutional kernel is smaller than the proportionality coefficient corresponding to the lower-level one-dimensional temporal convolutional kernel.

[0063] Here, in order to enable the neural network to learn dynamic temporal representations, the length of the one-dimensional temporal convolutional kernel is set to a specific proportion of the EEG signal sampling rate f s These proportionality coefficients are defined as δ i ∈ R, where i is the level of the temporal convolutional kernel layer. If there are L levels of temporal convolutional kernels, then i will vary from 1 to L.

[0064] Emotion-related activations are mainly observed in the Alpha (8 - 12 Hz), Beta (12 - 30 Hz), and Gamma (above 30 Hz) bands. In this work, the present invention expands the temporal receptive field, and the proportionality coefficient transformation δ i is [0.25, 0.50, 1.00], and L = 3 is set, with i = 1 to 3 to learn diverse frequency representations. The present invention assumes that the multi-scale temporal convolutional kernels can learn rich dynamic frequency representations from EEGs, providing more emotion-related information. From a temporal perspective, the multi-scale T convolutional kernels can capture long-term and short-term temporal patterns and learn more diverse representations. The higher-level T convolutional kernels have smaller proportionality coefficients, so the convolutional kernel length is shorter, and vice versa. The long temporal convolutional kernels can learn diverse representations of long-term time and low frequencies, while the short convolutional kernels extract short-term time and high-frequency representations.

[0065] In some embodiments, the output of the i-th level one-dimensional temporal convolution kernel in the temporal feature extraction module is defined as where X is the input EEG data, Conv1D() is a one-dimensional convolution operation with a convolution kernel size of and a stride of (1,1), Φ L-ReLU () is the Leaky ReLU activation function, and AP() is the average pooling operation; the outputs of the one-dimensional temporal convolution kernels at each level are concatenated along the temporal dimension, and a batch normalization operation is added to obtain the output result of the temporal feature extraction module.

[0066] Here, let X represent the EEG input sample, X = [X 0 , X 1 ,..., X n , where X n ∈R c×l where n is the number of EEG samples, c is the number of channels, and l is the length of each sample. The multi-scale temporal representation can be generated by parallel multi-scale temporal convolution kernels for the input EEG samples. Then, after passing through the Leaky ReLU activation function, the feature map is further downsampled by average pooling (AP). The reason for using AP is to reduce the influence of noise and the feature dimension because the EEG signal has a high dimension and a low signal-to-noise ratio. Let the output of the i-th level temporal convolution kernel where n is the number of samples, t is the number of T convolution kernels at each level, c is the number of channels, and f i is the feature length after the i-th level convolution operation.

[0067] Set L = 3, and i takes 1, 2, and 3 respectively. The output result of the temporal feature extraction module is:

[0068]

[0069] In the formula, f bn is the batch normalization operation, concat is the concatenation operation, and dim = 1 indicates along the temporal dimension;

[0070] In this way, the outputs of the one-dimensional temporal convolution kernels at each level are concatenated along the temporal dimension, and batch normalization is added after the temporal feature module. This reduces the problem of internal covariate shift in the neural network.

[0071] In some embodiments, the spatial feature extraction module has multi-scale one-dimensional spatial convolution kernels, including: a global convolution kernel for learning global spatial information, and a hemisphere convolution kernel and a local hemisphere convolution kernel for extracting the relationship between the left and right hemispheres through a shared convolution kernel; the size of the one-dimensional spatial convolution kernel can be defined as where c is the total number of channels of the input EEG segment, and j takes 1, 2, and 3 to represent the global convolution kernel, the hemisphere convolution kernel, and the local hemisphere convolution kernel respectively, and the corresponding δ j are 0.25, 0.5, and 1.0 respectively; the output of the j-th type of spatial convolution kernel is defined as

[0072]

[0073] where n is the number of samples, s is the number of one-dimensional spatial convolution kernels of each type, c m is the number of channels after the m-th spatial convolution, f is the feature length after each spatial convolution operation, Z S is the multi-scale spatial representation generated by the parallel multi-scale spatial convolution kernel for the input EEG data, Conv1D() is a one-dimensional convolution operation, and its convolution kernel size is The stride is (c, 1) for the global convolution kernel, the stride is (0.5×c, 1) for the hemisphere convolution kernel, and the stride is (0.25×c, 1) for the local hemisphere convolution kernel, Φ L-ReLU is the LeakyReLU activation function, and AP is the average pooling layer.

[0074] Here, the size of the one-dimensional spatial convolution kernel is related to the position of the EEG electrode channels. The size of the global spatial convolution kernel is (c, 1), where c is the number of channels. Since the length of the kernel is the same as the channel dimension of the input EEG segment, it can learn global spatial information. In the embodiments of the present invention, the frontal lobe region related to brain emotion asymmetry is combined into the kernel design. The proposed hemisphere spatial convolution kernel and local hemisphere spatial convolution kernel extract the relationship between the left and right hemispheres by sharing the convolution kernel. The size of the hemisphere spatial convolution kernel is (0.5·c, 1), and the stride is also (0.5·c, 1). The size of the local hemisphere spatial convolution kernel is (0.25·c, 1), and the stride is also (0.25·c, 1), where c is the total number of channels. The hemisphere kernel is shared by the two hemispheres and does not overlap, so that asymmetric patterns can be extracted.

[0075] In some embodiments, the method further includes: for the EEG data input to the hemisphere convolution kernel, deleting the electrode data of Fz, Cz, Pz, and Oz located at the midline position, and setting the channel arrangement order as [channel left ,channel right , where channel left represents the channels located in the left hemisphere, and channel right represents the channels located in the right hemisphere; the channel order on each hemisphere is rearranged so that each kernel weight is shared between electrode pairs symmetrically placed on the two hemispheres.

[0076] Here, when extracting spatial features, the frontal lobe regions with brain emotional asymmetry are considered. Because of the research involving brain asymmetry and emotion processing, researchers usually focus on the activity differences between the left and right hemispheres. Electrodes such as Fz, Cz, Pz, and Oz are located in the midline and cannot provide independent information about the left and right hemispheres. Therefore, when comparing the activities of the left and right hemispheres, the electrode position data at these positions in the dataset of the embodiments of the present invention should be excluded, and the processing process is as follows Figure 3 As shown, where part (a) is the global convolution kernel, part (b) is the hemisphere convolution kernel, and part (c) is the local hemisphere convolution kernel. It can be seen from (a) to (b) that before applying the hemisphere convolution kernel, the electrode data of Fz, Cz, Pz, and Oz located in the midline position are deleted.

[0077] In some embodiments, the following operations are adopted in both the temporal convolution operation and the spatial convolution operation to reduce the number of parameters and the amount of computation: first, perform independent depth convolution on the input channels of each EEG data, and then use 1x1 convolution to combine features between channels.

[0078] Here, assume that the input feature map is X, and the output of the depth convolution is X d , and the output of the point convolution (1x1 convolution) is X p , then there is:

[0079] X d = DepthwiseConv(X);

[0080] X p = PointwiseConv(X d );

[0081] In traditional convolution operations, the size of the convolution kernel is K×K×C, where K is the spatial size of the convolution kernel and C is the number of channels of the input feature map. For each output channel, such a convolution kernel is applied. Therefore, both the amount of computation and the number of parameters are proportional to the number of input and output channels, and its amount of computation is K×K×C×N×H×W. The amounts of computation of the depth convolution and the point convolution in DSC are K×K×C×H×W and 1×1×C×N×H×W respectively, and the total amount of computation is K×K×C×H×W + 1×1×C×N×H×W, where H and W are the height and width of the feature map, and N is the number of output channels. Therefore, compared with the standard convolution, the depthwise separable convolution reduces the amount of computation from K×K×C×N×H×W to K×K×C×H×W + 1×1×C×N×H×W.

[0082] In this way, the Depthwise Separable Convolution (DSC) technique is used in the temporal convolution operation and the spatial convolution operation. First, an independent spatial convolution (depth convolution) is performed on each input channel, and then a 1x1 convolution (point convolution) is used to combine features between channels. This significantly reduces the number of parameters and the amount of computation.

[0083] In some embodiments, in the self-attention mechanism part of the hybrid attention mechanism layer, the formula [Q, K, V] = z0U is used in the self-attention mechanism part of the hybrid attention mechanism layer. qkv Generate the query vector Q, the key vector K, and the value V; where z0 is the input feature vector and U qkv is the linear transformation matrix; calculate the self-attention output Attention(Q, K, V) through the following formula:

[0084]

[0085] Scale V = ||V -1 ||2, Scale K = ||K -1 ||2;

[0086] where D is the dynamic scaling matrix, ⊙ represents the Hadamard product, U copy and U sum are the linear transformation matrices; Scale V is the L2 norm of V, and Scale K is the L2 norm of K.

[0087] Here, the principle of the self-attention mechanism is as Figure 4 shown, where part (a) is the traditional self-attention mechanism and part (a) is the improved self-attention mechanism. It can be seen that the traditional self-attention mechanism (Self-Attention) can effectively capture long-range dependencies by calculating the relationships between each element in the input sequence, overcoming the limitations of traditional recurrent neural networks (RNNs) in processing long sequences. The core idea of self-attention is to convert the input vector into three different representations: query (Query), key (Key), and value (Value). By calculating the dot product of the query and the key to generate attention scores and using the Softmax function to convert them into a probability distribution, the values are weighted and summed to obtain the final output.

[0088] In the present invention, the input feature vector is multiplied by a linear transformation matrix to generate queries (Q), keys (K), and values (V). Then, the self-attention output Attention(Q, K, V) is calculated using a formula. Through dynamic scaling and linear transformation, the model can effectively process the input queries, keys, and values, and finally obtain the output of the self-attention mechanism through projection. To improve the stability and performance of the model, the queries (Q), keys (K), and values (V) are normalized by defining the L2 norm. Among them, the L2 norm is used to normalize vectors to ensure the numerical stability during the calculation of attention.

[0089] In some embodiments, in the hybrid attention mechanism layer, an improved convolutional block attention module (CBAM) is added after the spatial features output by the two-stream spatio-temporal feature construction layer. This improved CBAM module combines the channel attention mechanism and the spatial attention mechanism, takes the feature map F output by the spatial feature extraction module as the input, and sequentially obtains the channel attention map M c (F) and the spatial attention map M s (F'), and uses a convolutional layer as a shared network to replace the shared MLP layer in the original CBAM to output the final feature representation; the convolutional layer consists of a convolution including a hidden layer; where M c (F) and M s (F) are calculated through the following formulas respectively:

[0090]

[0091] In the formulas, AvgPool(F) and MaxPool(F) respectively represent the average pooling operation and the max pooling operation on the feature map F output by the spatial feature extraction module, and AvgPool(F') and MaxPool(F') respectively represent the average pooling operation and the max pooling operation on the feature map F’ output by the channel attention output; and are respectively the average pooling feature and the max pooling feature in the channel attention; and are respectively the average pooling feature and the max pooling feature in the spatial attention; W0 and W1 are both learnable weight matrices, and f k represents the convolution operation, and () is the Sigmoid function.

[0092] Here, since the convolution operation extracts information features by mixing cross-channel and spatial information, it is important to add attention representations in the channel dimension and the spatial dimension because not every channel and spatial position is equally important. In the embodiment of the present invention, a CBAM module is added after the spatial features output by the two-stream spatio-temporal feature construction layer, and it takes the feature map F ∈ R output by the spatial feature extraction module C×LAs input, where C represents the number of channels and L represents the spatial dimension of the feature map. The schematic structure of the CBAM module is as shown in Figure 5 . Among them, part (a) is the overall CBAM structure, part (b) is the channel attention structure, and part (c) is the spatial attention structure.

[0093] CBAM first calculates the channel attention map M c ∈R C×1 . In channel attention, average pooling and max pooling are used to aggregate spatial information. Among them, average pooling calculates the average value of the elements within the pooling window, while max pooling calculates the maximum value of the elements within the pooling window. Through average pooling and max pooling, the average pooling feature and the max pooling feature are generated respectively. These two features are passed through a shared network to generate the channel attention map M c . The update process of the final channel attention output feature map can be represented by the following formula:

[0094]

[0095] where, represents element-wise multiplication, and F' represents the output of channel attention.

[0096] Next, CBAM calculates the spatial attention map M s ∈R 1×L , which is generated by aggregating channel information. Spatial attention first aggregates the channel information of the channel attention output feature F' through two pooling operations to generate two maps, representing the cross-channel average pooling feature and the max pooling feature respectively. After these two features are concatenated, a spatial attention map is generated through a convolutional layer. After being weighted by channel attention and spatial attention, the final update process of the feature map is:

[0097]

[0098] where, represents element-wise multiplication, F' represents the output of channel attention, and F” represents the final output after CBAM.

[0099] When the input spatial features pass through the channel attention and spatial attention in CBAM in sequence, the shared MLP layer in the original CBAM is improved to a convolutional layer, which enables the model to reduce a certain amount of parameter quantity. This shared network consists of a convolution containing a hidden layer. To reduce parameter overhead, the output of the hidden layer is R C / r×1, where r is the scaling ratio. In this way, CBAM can more effectively focus on important feature channels and spatial positions, thereby improving the overall performance of the model.

[0100] In some embodiments, the advanced fusion layer is used to fuse the learning information from the global convolution kernel, the hemisphere convolution kernel, and the local hemisphere convolution kernel along the spatial dimension by using a one-dimensional convolution layer with a kernel size of (3,1), with the time feature as an additional feature input, and finally the learned spatio-temporal feature representation Z of the global hemisphere fusion is obtained. fusion Generated by the following formula:

[0101] Z fusion = GAP(f bn (AP(Φ L-ReLu (Conv1D(Z, (3,1))))));

[0102] Z = Attention(Q, K, V) + F + F;

[0103] where Z represents the concatenation of the self-attention output Attention(Q, K, V), the feature map F output by the spatial feature extraction module, and the spatial feature F'' output by F after being processed by CBAM, f bn is the batch normalization function, GAP() represents the global average pooling layer, Φ L-ReLU is the LeakyReLU activation function, AP is the average pooling layer, Conv1D() represents the one-dimensional convolution operation with a kernel size of (3,1);

[0104] Z fusion is input into the fully connected layer and activated by the softmax function. The final output OutPut can be calculated in the following way:

[0105] OutPut = Φsoftmax(W'Φdp(ΦReLU(W(Γ(Zfusion)) + b)) + b')

[0106] OutPut = Φ softmax (W′Φ dp (Φ ReLU (W(Γ(Z fusion )) + b)) + b′);

[0107] where Γ is the squeezing operation, W and W' are trainable weight matrices, b and b' are bias terms, φ ReLU is the ReLU activation function, applied to the output of the linear transformation, Φ dp is a dropout operation for preventing overfitting, Φ softmax is the softmax function for converting the output into a probability distribution.

[0108] Here, for the feature output Z of the previous stage, a one-dimensional convolutional layer with a kernel size of (3, 1) is used to fuse information along the spatial dimension. After LeakyReLU, average pooling, and batch normalization, a global average pooling layer (GAP) is added to overcome overfitting and reduce the model size. Finally, the spatio-temporal feature representation Z of the feature output is learned. fusion , Z fusion will be input into the fully connected layer. The final output layer is activated by the softmax function. Thus, the learning information from the global, hemispherical, and local hemispherical is fused into the high-level spatio-temporal feature representation.

[0109] The following describes the above emotion recognition method based on the spatio-temporal multi-scale attention convolutional neural network in conjunction with a specific embodiment. However, it should be noted that this specific embodiment is only for better explaining the present invention and does not constitute an improper limitation to the present invention.

[0110] In the experimental setup, each trial was segmented into 4 non - overlapping segments, called cropped experiments, and subject - based 10 - fold cross - validation was performed for each subject in the dataset to prevent potential data leakage problems. The reason for conducting cropped experiments is that predictions on shorter segments are more favorable than trial - based predictions evaluated in previous studies to achieve an efficient recognition system. Additionally, for the real - world situation where the test data is unknown to the model, a decoding model with good generalization ability is needed. In each trial, the subject was required to watch or listen to a certain stimulus, which should trigger a certain type of emotion. Since emotion is a continuous cognitive process in the brain, the data segments within a single trial are highly correlated. Therefore, randomly shuffling the segments across subjects before the training - test split of the data might cause adjacent segments to appear in both the training and test data simultaneously, leading to high classification results. However, when the model has never seen highly correlated segments in the real - world situation, the accuracy will decrease. To obtain a more general evaluation, 10 - fold cross - validation was performed with splits across subjects to ensure that adjacent segments within a subject do not appear in both the training and test data simultaneously. In each step of the 10 - fold cross - validation, one fold was selected as the test data, and the remaining 9 folds were used as training data. In these 9 training folds, the data was randomly divided into 80% training data and 20% validation data. During training, the network was trained for 500 epochs on the training data, and the network was evaluated on the validation data in each epoch. Among these 500 epochs, the model with the highest accuracy on the validation data was saved and tested on the test data. The above process was repeated 10 times for each subject in the dataset until each fold was selected as the test fold once. In each fold, the test data remained completely unseen during all stages of training and validation. Finally, when performing 10 - fold cross - validation, for each subject, the highest accuracy and the highest F1 - score of the model in each fold were calculated, and then the average of these 10 folds was obtained to evaluate the overall performance of the model.

[0111] The proportionality coefficients of the T - kernel length for the DEAP dataset are [0.25, 0.50, 1.00]. The sampling rate of the DEAP data is 128 Hz. Therefore, according to the time - kernel lengths, they are 32, 64, and 128 respectively. The maximum number of training epochs for model training is 500. The batch size for the DEAP dataset was set to 64, and the Adam optimizer was used to optimize the training process, with an initial learning rate of 1e -3 . At the same time, the cross - entropy loss was selected as the loss function to guide the training process, and the relevant calculation formula is as follows:

[0112]

[0113] where, is the loss value, y is the true label (0 or 1), is the probability predicted by the model (i.e., the probability of predicting as class 1).

[0114] To evaluate the trained lightweight convolutional neural network model, experiments were conducted on a publicly available benchmark dataset, namely the Database for Emotion Analysis using Physiological Signals (DEAP). DEAP is a multimodal human emotion state dataset that includes electroencephalogram (EEG), facial expressions, and galvanic skin response (GSR). Thirty-two subjects watched music video clips while their EEG, facial expressions, and GSR were recorded.

[0115] During the processing of the DEAP dataset, first, the 3-second pre-trial baseline of each trial was removed. Subsequently, the data was downsampled from the original 512 Hz to 128 Hz to reduce the data volume and potentially reduce the computational load. And an independent component analysis method was used to remove the electrooculogram (EOG) signals to eliminate the artifacts generated by eye movements. To effectively eliminate low-frequency and high-frequency noise, a band-pass filter starting from 4.045 Hz was applied to the original EEG. Finally, the EEG channels were averaged to a common reference point to reduce the influence of the reference electrode and make the signals easier to compare. The emotion dimensions of the DEAP dataset include arousal, valence, and dominance. In the processing of the labels, the class label range for each emotion dimension is from 1 to 9. Therefore, 5 was selected as the threshold to project the 9 discrete values into the low class and high class for each dimension. The present invention only focuses on the two dimensions of arousal and valence.

[0116] Considering that deep neural networks have a relatively large number of trainable parameters, in order to optimally learn the emotional state representation in EEG, a large number of labeled data samples are required, but the number of trials in the selected dataset is relatively small. To address this challenge, the present invention provides a data augmentation strategy by splitting each trial into non-overlapping 4-second segments. These segments are then used to train the constructed lightweight convolutional neural network model to improve the performance of the model.

[0117] One of the evaluation metrics is Accuracy, which is one of the most commonly used evaluation metrics in classification problems. It is the ratio of the number of correctly predicted samples to the total number of samples. For a binary classification problem, Accuracy can be defined as:

[0118]

[0119] where TP is true positive, TN is true negative, FP is false positive, and FN is false negative.

[0120] Accuracy can measure the precision of predictions for class - balanced datasets. After preprocessing the labels mentioned in the preprocessing section, the labels become imbalanced. To better evaluate the performance of the classifier on class - imbalanced datasets, the F1 - score is added. It combines the precision and recall of the classifier and is defined as the harmonic mean of the classifier's precision and recall. The F1 is defined as shown in the following formula.

[0121]

[0122] Among them, TP is the true positive, TN is the true negative, FP is the false positive, and FN is the false negative.

[0123] Finally, the evaluation results are shown in Table 1. Table 1 shows the performance comparison of different methods on the arousal and valence tasks, including indicators such as accuracy, F1 - score, and the number of parameters.

[0124] Table 1 Comparison of evaluation results between the method of the present invention and existing methods

[0125]

[0126] Among the methods for emotion classification tasks based on EEG signals listed in Table 1, the principles and characteristics of each method are different. SVM separates different emotion classes by constructing a hyperplane and is suitable for simple linearly separable tasks; KNN votes to determine the emotion class based on the distance between samples, but has limited processing ability for high - dimensional data; EEGNet uses a convolutional neural network to extract spatio - temporal features of EEG signals and can capture complex local patterns; SCN adapts to the complexity of EEG data by dynamically adjusting the network structure, improving the generalization ability of the model; while DCN extracts high - level pattern features of EEG signals through deep convolution, further enhancing the ability to capture complex emotion patterns.

[0127] From Figure 6 and Figure 7 The statistical results show that the method of the present invention is significantly superior to other methods in terms of accuracy and F1 - score for arousal and valence, reaching 85.39% and 87.12% (arousal) and 71.23% and 72.44% (valence) respectively, and the number of parameters is only 6978, reflecting high performance and model efficiency. This result indicates that the method of the present invention has significant advantages in extracting EEG signal features and emotion classification tasks.

[0128] In contrast, traditional methods (such as SVM and KNN) have low performance due to limited feature extraction capabilities. Compared with EEGNet, although the number of parameters of the method of the present invention has increased, the accuracy rates of arousal and valence have increased by 27.10% and 16.67% respectively, and the F1 scores have increased by 26.52% and 14.83% respectively. The improvement in performance far exceeds the increase in the number of parameters, demonstrating a higher model efficiency. In addition, compared with other deep learning methods (such as SCN and DCN), while the method of the present invention significantly leads in performance, the number of parameters is much lower than the number of parameters of SCN, which is 48,162, and the number of parameters of DCN, which is 151,252, further verifying its balance advantage between performance and complexity.

[0129] Further analysis shows that the method of the present invention significantly reduces the model complexity while ensuring high classification performance by optimizing the network structure and feature extraction strategy. This efficient design makes the method of the present invention not only applicable to the emotion classification task, but also has strong practicability, especially showing great application potential in real-time emotion monitoring scenarios with limited computing resources.

[0130] The emotion recognition method based on the spatio-temporal multi-scale attention convolutional neural network provided by the present invention, on the one hand, introduces brain emotion asymmetry. By using hemisphere convolutional kernels with lengths corresponding to the number of channels in the left and right hemispheres, hemisphere asymmetry patterns related to emotional states are extracted, which helps to improve the emotion recognition accuracy. On the other hand, a multi-scale one-dimensional convolutional layer is adopted, which can parallelly extract features in different time and space dimensions from the input data. In this way, the model can capture the features and transformation patterns of the data at different scales in time and space.

[0131] This method realizes an end-to-end electroencephalogram emotion recognition model without manually extracting features. By constructing a multi-scale time and space convolutional network, the internal connection between the local channels and global channels of the electroencephalogram signal can be effectively learned. There is a significant improvement in the recognition accuracy, and the model operation parameters are effectively reduced.

[0132] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present invention. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present invention, the magnitudes of the serial numbers of the above processes do not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention. The serial numbers of the embodiments of the present invention above are only for description and do not represent the advantages and disadvantages of the embodiments.

[0133] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including such element.

[0134] In several embodiments provided by the present invention, it should be understood that the disclosed methods can be implemented in other ways. The methods disclosed in several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments. The features disclosed in several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.

[0135] As mentioned above, it is only the implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. An emotion recognition method based on spatiotemporal multi-scale attention convolutional neural network, characterized in that: include: Collect the EEG data of the subjects for preprocessing, and perform channel mapping according to the physical location of the electrodes to obtain EEG data containing spatial and temporal dimensions; Constructing a lightweight convolutional neural network including a dual-stream spatiotemporal feature construction layer, a hybrid attention mechanism layer, a high-order fusion layer and a classification layer; wherein the dual-stream spatiotemporal feature construction layer includes: a temporal feature extraction module for learning time-frequency feature representation using multi-scale one-dimensional temporal convolution kernels, and a parallel spatial feature extraction module for learning asymmetric representations between the left and right hemispheres using hemispheric convolution kernels and local hemispheric convolution kernels whose lengths correspond to the number of left and right hemisphere channels; the hybrid attention mechanism layer combines the channel attention mechanism, the spatial attention mechanism and the self-attention mechanism to enhance the capabilities of feature extraction and data processing; the high-order fusion layer is used to relearn from the learned global convolution kernel to the representation of the local hemispheric convolution kernel; The EEG data is recognized using the trained lightweight convolutional neural network to obtain emotion recognition results of the subject.

2. The emotion recognition method according to claim 1, characterized in that: The temporal feature extraction module includes a multi-scale one-dimensional temporal convolution kernel, the size of the i-th level one-dimensional temporal convolution kernel is It can be defined as where f s is the EEG signal sampling rate, i∈[1,2,...,L], L is the number of levels of the one-dimensional temporal convolution kernel layer, and the proportional coefficient δ is 1, 2, or 3. i They are 0.25, 0.50, and 1.00 respectively; the proportional coefficient corresponding to the high-level one-dimensional time convolution kernel is smaller than the proportional coefficient corresponding to the low-level one-dimensional time convolution kernel.

3. The emotion recognition method according to claim 2, characterized in that: The output of the i-th level one-dimensional temporal convolution kernel in the temporal feature extraction module Defined as Where X is the input EEG data, Conv1D() is the convolution kernel size A one-dimensional convolution operation with a stride of (1,1), Φ L-ReLU ( ) is the Leaky ReLU activation function, AP( ) is the average pooling operation; The output of the one-dimensional temporal convolution kernel at each level The output result of the time feature extraction module is obtained by connecting them in series along the time dimension and adding a batch normalization operation.

4. The emotion recognition method according to claim 1, characterized in that: The spatial feature extraction module has a multi-scale one-dimensional spatial convolution kernel, including: a global convolution kernel for learning global spatial information, and a hemispheric convolution kernel and a local hemispheric convolution kernel for extracting the relationship between the left and right hemispheres through a shared convolution kernel; The size of the one-dimensional spatial convolution kernel It can be defined as Where c is the total number of channels of the input EEG fragment, j is 1, 2, and 3, respectively, representing the global convolution kernel, hemispherical convolution kernel, and local hemispherical convolution kernel, and the corresponding δ j are 0.25, 0.5, and 1.0 respectively; the output of the j-th type of spatial convolution kernel Defined as: in, n is the number of samples, s is the number of one-dimensional spatial convolution kernels of each type, and c m is the number of channels after the mth spatial convolution, f is the feature length after each spatial convolution operation, and Z S It is a multi-scale spatial representation of the input EEG data generated by the parallel multi-scale spatial convolution kernel. Conv1D() is a one-dimensional convolution operation with a convolution kernel size of The step size is (c, 1) for the global convolution kernel, the step size is (0.5×c, 1) for the hemispherical convolution kernel, and the step size is (0.25×c, 1) for the local hemispherical convolution kernel. L-ReLU is the LeakyReLU activation function and AP is the average pooling layer.

5. The emotion recognition method according to claim 4, characterized in that: The method further comprises: For the EEG data input into the hemispheric convolution kernel, the Fz, Cz, Pz, and Oz electrode data located at the midline position are deleted, and the channel arrangement order is set to [channel left ,channel right ], where channel left Indicates the channel located in the left hemisphere, channel right represents channels located in the right hemisphere; the order of channels on each hemisphere was rearranged so that each kernel weight is shared between symmetrically placed electrode pairs on both hemispheres.

6. The emotion recognition method according to any one of claims 1 to 5, characterized in that: The following operations are used in both temporal and spatial convolution operations to reduce parameters and computation: First, an independent depth convolution is performed on each input channel of the EEG data, and then a 1x1 convolution is used to combine features between channels.

7. The emotion recognition method according to any one of claims 1 to 5, characterized in that: In the self-attention mechanism part of the hybrid attention mechanism layer, the formula [Q, K, V] = z0U is used. qkv Generate query vector Q, key vector K and value V; where z0 is the input feature vector, U qkv is the linear transformation matrix; The self-attention output Attention(Q,K,V) is calculated by the following formula: Scale V =||V -1 ||2,Scale K =||K -1 ||2; Where D is the dynamic scaling matrix, ⊙ represents the Hadamard product, and U copy and U sum is the linear transformation matrix; Scale V is the L2 norm of V, Scale K is the L2 norm of K.

8. The emotion recognition method according to any one of claims 1 to 5, characterized in that: In the hybrid attention mechanism layer, the improved convolutional block attention module CBAM is added to the spatial features output by the dual-stream spatiotemporal feature construction layer. The improved CBAM module combines the channel attention mechanism and the spatial attention mechanism, takes the feature map F output by the spatial feature extraction module as input, and obtains the channel attention map M in sequence. c (F) and the spatial attention map M s (F'), and replace the shared MLP layer in the original CBAM with the convolutional layer as the shared network to output the final feature representation; The convolutional layer consists of a convolution layer including a hidden layer; where M c (F) and M s (F) are calculated by the following formula: Where AvgPool(F) and MaxPool(F) respectively represent the average pooling operation and the maximum pooling operation on the feature map F output by the spatial feature extraction module, and MaxPool(F') and MaxPool(F') respectively represent the average pooling operation and the maximum pooling operation on the feature map F' output by the channel attention. and They are the average pooling features and the maximum pooling features in channel attention respectively; and They are the average pooling features and the maximum pooling features in spatial attention respectively; W0 and W1 are both learnable weight matrices, f k Represents the convolution operation, σ() is the Sigmoid function.

9. The emotion recognition method according to any one of claims 1 to 5, characterized in that: The advanced fusion layer is used to fuse the learning information from the global convolution kernel, the hemispherical convolution kernel and the local hemispherical convolution kernel along the spatial dimension using a one-dimensional convolution layer with a kernel size of (3,1), and add the temporal feature as an additional feature input. The final learned global hemispherical fused spatiotemporal feature representation Z fusion Generated by: Z=Attention(Q,K,V)+F”+F; Where Z represents the concatenation of the self-attention output Attention(Q,K,V), the feature map F output by the spatial feature extraction module, and the spatial feature F” output by F after CBAM processing, f bn is the batch normalization function, GAP( ) represents the global average pooling layer, Φ L-ReLU is the LeakyReLU activation function, AP is the average pooling layer, Conv1D() represents a one-dimensional convolution operation, and its convolution kernel size is (3,1); Z fusion It is input into the fully connected layer and activated by the softmax function. The final output OutPut can be calculated as follows: Where Γ is the squeeze operation, W and W' are trainable weight matrices, b and b' are bias terms, and Φ ReLU is the ReLU activation function, applied to the output of the linear transformation, Φ dp is a dropout operation used to prevent overfitting, Φ softmax is the softmax function, which is used to convert the output into a probability distribution.

Citation Information

Cited By

  • Electroencephalogram abnormal signal detection method based on step-by-step identification and multi-agent decision

    CN120694664A

  • Method and system for identifying spatial-temporal characteristics of electroencephalogram signals

    CN120705665A

  • A method and system for recognizing spatio-temporal features of electroencephalogram signals

    CN120705665B

  • SAR image ship identification method based on multi-scale feature fusion and sphere space embedding

    CN120783226A

  • Chinese text difficulty classification method and system and storage medium

    CN120951076A