Multi-scale attention emotion recognition method, electronic equipment and storage medium

By adopting a multi-scale attention method in EEG signal emotion recognition, using temporal attention and multi-scale dense spatiotemporal feature extraction module, combined with attention-based feature fusion module, the problem of difficulty in capturing space-time dependence and ignoring spatial information in the prior art is solved, and more efficient and accurate emotional recognition is achieved.

CN120011857APending Publication Date: 2025-05-16SHANGHAI ZERO UNIQUE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510126732.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art has problems in the recognition of EEG signal sentiment, which are difficult to capture long-distance time dependencies, ignore spatial information, high demand for computing resources, and high model complexity.

Method used

The multi-scale attention emotion recognition method is adopted to process EEG data weighted by the time attention module, and combined with the multi-scale dense spatiotemporal feature extraction module, use 2D-dense blocks and transition layers of different scales to extract features, and integrate features of different scales in the attention-based multi-scale feature fusion module.

Benefits of technology

It improves the spatio-temporal pattern capture capability of the model in EEG data, enhances the accuracy and stability of emotion recognition, and reduces the complexity of the model and computing resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011857A_ABST
    Figure CN120011857A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a multi-scale attention emotion recognition method, electronic equipment and a storage medium. The method comprises the following steps: acquiring emotion stimulation electroencephalogram data; determining a first feature of time step average signal intensity and a second feature of maximum signal intensity in a time attention module, and weighting the emotion stimulation electroencephalogram data along a time dimension based on an attention map generated by splicing to obtain attention-enhanced electroencephalogram data; inputting the electroencephalogram data into a multi-scale dense spatial-temporal feature extraction module, and outputting electroencephalogram comprehensive features of different scales for capturing a spatial-temporal mode; performing dimension alignment and fusion on the electroencephalogram comprehensive features of different scales to generate multi-scale fusion features; and determining an emotion recognition result through the multi-scale fusion features. According to the embodiment of the invention, the defects in EEG signal emotion recognition are effectively overcome, the accuracy, stability and adaptability of the model are improved, and a more effective emotion recognition solution is provided for the field of emotion brain-computer interfaces.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of emotional brain-computer interface, and in particular to a method, system, electronic device and storage medium for multi-scale attention emotion recognition and emotion recognition model training. Background Art

[0002] Emotional brain-computer interface includes two tasks: emotion recognition and emotion regulation. In order to accurately recognize emotions, existing technologies usually use the following methods: Methods based on convolutional neural networks (CNNs): Utilize the local feature capture capability of CNNs to process EEG signals, extract features from signals through operations such as convolutional layers and pooling layers to identify emotion-related patterns, such as EEGNet, whose structure is mainly composed of convolutional layers, pooling layers, and fully connected layers. The convolutional layer is used to extract local features of EEG signals, and capture feature information of different scales through convolution kernels of different sizes; the pooling layer is used to reduce data dimensions and reduce the amount of calculation; the fully connected layer maps the extracted features to the corresponding emotion categories. In specific implementation, the EEG signal is input into the network, and features are extracted through the convolutional layer to obtain a series of feature maps; then, the pooling layer downsamples the feature maps to reduce the resolution; then, after multiple layers of convolution and pooling operations, the fully connected layer performs classification prediction and outputs the emotion category.

[0003] Transformer-based methods: The Transformer self-attention mechanism is used to learn long-distance dependencies in EEG signals to classify emotional states, such as ST-Trans, whose structure includes a Transformer encoder, which contains a multi-head self-attention mechanism and a feedforward neural network. The multi-head self-attention mechanism can simultaneously focus on different parts of the input sequence and capture long-distance dependencies; the feedforward neural network is used to further process and transform the features. In the specific implementation, the EEG signal is first converted into a format suitable for Transformer processing and then input into the Transformer encoder. In the encoder, the multi-head self-attention mechanism is used to calculate the degree of correlation between each position and other positions to obtain a weighted feature representation; then, the feature is transformed through the feedforward neural network; finally, the emotion classification result is output through the fully connected layer.

[0004] Combining CNN and Transformer methods: Some models combine the advantages of CNN and Transformer, first using CNN to extract local spatial features, and then using Transformer to capture long-distance temporal dependencies to improve emotion recognition performance, such as CNN-Trans, whose structure consists of CNN modules and Transformer modules. The CNN module is responsible for extracting the spatial features of EEG signals, which is achieved through a combination of convolutional layers and pooling layers; the Transformer module is used to learn long-distance dependencies in the temporal dimension, including multi-head self-attention layers and feedforward layers. In the specific implementation, the EEG signal first enters the CNN module, and after a series of convolution and pooling operations, the spatial features are extracted; then, the obtained feature map is converted into a sequence form and input into the Transformer module; in the Transformer module, the temporal dependencies are captured through the multi-head self-attention mechanism, and then processed through the feedforward layer; finally, the emotion classification results are output through the fully connected layer.

[0005] In the process of implementing the present invention, the inventors found that there are at least the following problems in the related art: CNN-based methods: It is difficult to capture long-distance temporal dependencies and lacks the ability to integrate global information. Specifically, CNN has a limited receptive field and focuses mainly on local features. It is difficult to effectively capture long-distance temporal dependencies between different channels in EEG signals, which may result in the loss of important temporal dynamic information and affect the accuracy of emotion recognition. For example, a small-sized convolution kernel can only cover a shorter time segment when processing EEG signals, and cannot perceive the impact of signal changes at distant time points on the current moment. When processing EEG signals, CNN focuses on local feature extraction, and lacks the ability to integrate global information of the entire signal. It cannot fully utilize all the information in the signal for comprehensive judgment, resulting in limited recognition of complex emotional patterns.

[0006] Transformer-based methods: Ignore spatial information and have high computational resource requirements. Specifically, although Transformer can capture long-distance temporal dependencies well, it may relatively ignore the spatial information in EEG signals during feature extraction because its self-attention mechanism focuses on the relationship between elements in the sequence, while the spatial structural relationship between different channels is insufficiently utilized, affecting the recognition of emotion-related spatial patterns. The multi-head self-attention mechanism in Transformer has a high computational complexity. When processing large-scale EEG data, it requires a lot of computing resources and time, which limits its efficiency and scalability in practical applications.

[0007] Combining CNN and Transformer: The feature fusion effect needs to be optimized and the model complexity is relatively high. Specifically, although the advantages of CNN and Transformer are combined, in actual applications, the feature fusion method of the two may not be flexible and efficient enough, and their respective strengths cannot be fully utilized, resulting in limited performance improvement of the model under complex data conditions and failure to achieve the best emotion recognition effect. Since both CNN and Transformer modules are included, the model structure is relatively complex, which increases the difficulty of model training and the complexity of parameter adjustment, and is prone to problems such as overfitting. In addition, the hardware equipment requirements are relatively high in actual deployment. Summary of the invention

[0008] In order to at least solve the problems of limited emotion recognition ability, poor efficiency and scalability, and high complexity in the existing technology.

[0009] In a first aspect, an embodiment of the present invention provides a multi-scale attention emotion recognition method, comprising: Acquire EEG data of emotional stimulation; Inputting the emotion stimulation EEG data into a time attention module, determining in the time attention module a first feature for reflecting the average signal strength of a time step and a second feature of the maximum signal strength in each time step, using an attention map generated by splicing the first feature and the second feature, and weighting the emotion stimulation EEG data along the time dimension based on the attention map to obtain attention-enhanced EEG data; Inputting the EEG data into a multi-scale dense spatiotemporal feature extraction module, wherein the multi-scale dense spatiotemporal feature extraction module includes a plurality of combination layers of different scales, wherein each combination layer includes a plurality of 2D-dense blocks and transition layers, and outputs EEG comprehensive features of different scales that capture spatiotemporal patterns; Dimensionally aligning and fusing the EEG comprehensive features of different scales to generate multi-scale fusion features; The emotion recognition result is determined by the multi-scale fusion feature.

[0010] In a second aspect, an embodiment of the present invention provides a method for training a multi-scale attention emotion recognition model, comprising: Acquire emotional stimulation EEG data for training, and input the emotional stimulation EEG data into the emotion recognition model, wherein the emotion recognition model includes: a time attention module, a multi-scale dense spatiotemporal feature extraction module, and an attention-based multi-scale feature fusion module; Determining in the temporal attention module a first feature for reflecting the average signal strength of a time step and a second feature of the maximum signal strength in each time step, using an attention map generated by splicing the first feature and the second feature, and weighting the emotion stimulation EEG data along the time dimension based on the attention map to obtain attention-enhanced EEG data; Input the EEG data into a multi-scale dense spatiotemporal feature extraction module, wherein the multi-scale dense spatiotemporal feature extraction module includes a plurality of combination layers of different scales, each combination layer includes a plurality of 2D-dense blocks and a transition layer, in each combination layer, the output of the nth 2D-dense block and the EEG data are used as the input of the n+1th 2D-dense block, and the transition layer after each 2D-dense block is used to reduce the number of feature channels and reduce the temporal resolution, so that the combination layer outputs comprehensive EEG features of different scales from the original input to the learning of each layer; In the attention-based multi-scale feature fusion module, dimensionally aligning and fusing the EEG comprehensive features of different scales to generate multi-scale fusion features; The predicted emotion recognition result is determined by the multi-scale fusion feature, and the emotion recognition model is trained based on the difference between a preset benchmark emotion recognition result and the predicted emotion recognition result until a preset goal is achieved.

[0011] In a third aspect, an embodiment of the present invention provides a multi-scale attention emotion recognition system, comprising: EEG acquisition module, used to acquire EEG data of emotional stimulation; A signal highlighting module, used for inputting the emotion stimulation EEG data into a time attention module, determining in the time attention module a first feature for reflecting the average signal strength of a time step and a second feature of the maximum signal strength in each time step, using an attention map generated by splicing the first feature and the second feature, and weighting the emotion stimulation EEG data along the time dimension based on the attention map to obtain attention-enhanced EEG data; A multi-scale feature capture module, used for inputting the EEG data into a multi-scale dense spatiotemporal feature extraction module, wherein the multi-scale dense spatiotemporal feature extraction module comprises a plurality of combination layers of different scales, wherein each combination layer comprises a plurality of 2D-dense blocks and transition layers, and outputs EEG comprehensive features of different scales capturing spatiotemporal patterns; A fusion module is used to align and fuse the EEG comprehensive features of different scales to generate multi-scale fusion features; The recognition module is used to determine the emotion recognition result through the multi-scale fusion feature.

[0012] In a fourth aspect, an embodiment of the present invention provides a training system for a multi-scale attention emotion recognition model, comprising: An EEG acquisition module, used to acquire emotional stimulation EEG data for training, and input the emotional stimulation EEG data into the emotion recognition model, wherein the emotion recognition model includes: a time attention module, a multi-scale dense spatiotemporal feature extraction module, and an attention-based multi-scale feature fusion module; A signal highlighting module, configured to determine in the time attention module a first feature for reflecting the average signal strength of a time step and a second feature of the maximum signal strength in each time step, and to weight the emotion stimulation EEG data along the time dimension based on the attention map by using the attention map generated by splicing the first feature and the second feature, so as to obtain attention-enhanced EEG data; A multi-scale feature capture module is used to input the EEG data into a multi-scale dense spatiotemporal feature extraction module, wherein the multi-scale dense spatiotemporal feature extraction module includes a plurality of combination layers of different scales, each combination layer includes a plurality of 2D-dense blocks and a transition layer, in each combination layer, the output of the nth 2D-dense block and the EEG data are used as the input of the n+1th 2D-dense block, and the transition layer after each 2D-dense block is used to reduce the number of feature channels and reduce the temporal resolution, so that the combination layer outputs comprehensive EEG features of different scales from the original input to the learning of each layer, A fusion module, used for dimensional alignment and fusion of the EEG comprehensive features of different scales in the attention-based multi-scale feature fusion module to generate multi-scale fusion features; The training module is used to determine the predicted emotion recognition result through the multi-scale fusion feature, and train the emotion recognition model based on the difference between the preset baseline emotion recognition result and the predicted emotion recognition result until a preset goal is achieved.

[0013] In a fifth aspect, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the multi-scale attention emotion recognition method and the emotion recognition model training method of any embodiment of the present invention.

[0014] In a sixth aspect, an embodiment of the present invention provides a storage medium having a computer program stored thereon, characterized in that when the program is executed by a processor, the steps of the multi-scale attention emotion recognition method and the emotion recognition model training method of any embodiment of the present invention are implemented.

[0015] In the seventh aspect, an embodiment of the present invention provides a computer program product, including a computer program / instructions, characterized in that when the computer program / instructions are executed by a processor, the steps of the multi-scale attention emotion recognition method and the emotion recognition model training method of any embodiment of the present invention are implemented.

[0016] The beneficial effects of the embodiments of the present invention are: effectively solving the defects of the prior art in EEG signal emotion recognition, improving the accuracy, stability and adaptability of the model, and providing a more effective emotion recognition solution for the field of emotional brain-computer interface. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 is a flow chart of a multi-scale attention emotion recognition method provided by an embodiment of the present invention; Figure 2 is a schematic diagram of an emotion recognition model structure of a multi-scale attention emotion recognition method provided by an embodiment of the present invention; Figure 3 It is a flowchart of a training method of a multi-scale attention emotion recognition model provided by an embodiment of the present invention; Figure 4 It is a schematic diagram of the performance of different models on Task 1 and Task 2 of a training method of a multi-scale attention emotion recognition model provided by an embodiment of the present invention; Figure 5 It is a schematic diagram of an ablation experiment of MSADM in Task 1 of a training method of a multi-scale attention emotion recognition model provided by an embodiment of the present invention; Figure 6 is a schematic diagram of the structure of a multi-scale attention emotion recognition system provided by an embodiment of the present invention; Figure 7 It is a structural schematic diagram of a training system for a multi-scale attention emotion recognition model provided by an embodiment of the present invention; Figure 8 A schematic structural diagram of an embodiment of an electronic device for multi-scale attention emotion recognition provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0020] like Figure 1 FIG. 1 is a flowchart of a multi-scale attention emotion recognition method provided by an embodiment of the present invention, comprising the following steps: S11: Acquiring EEG data of emotional stimulation; S12: inputting the emotion stimulation EEG data into a time attention module, determining in the time attention module a first feature for reflecting the average signal strength of a time step and a second feature of the maximum signal strength in each time step, using an attention map generated by splicing the first feature and the second feature, and weighting the emotion stimulation EEG data along the time dimension based on the attention map to obtain attention-enhanced EEG data; S13: inputting the EEG data into a multi-scale dense spatiotemporal feature extraction module, wherein the multi-scale dense spatiotemporal feature extraction module comprises a plurality of combination layers of different scales, wherein each combination layer comprises a plurality of 2D-dense blocks and transition layers, and outputting EEG comprehensive features of different scales that capture spatiotemporal patterns; S14: Dimensionally aligning and fusing the EEG comprehensive features of different scales to generate multi-scale fusion features; S15: Determine the emotion recognition result through the multi-scale fusion feature.

[0021] In this embodiment, considering the prior art, if the above defects are to be improved, one may usually try to optimize the convolution kernel design of CNN or increase the number of network layers to enhance the ability to capture long-distance temporal dependencies. However, this often leads to an increase in model parameters, an increase in computational complexity, and an overfitting problem. For Transformer, one may try to add some position encoding or spatial perception modules to its structure to introduce spatial information, but this may destroy the original structural advantages of Transformer, and it is difficult to ensure that the newly introduced modules work effectively with the self-attention mechanism. In the method combining CNN and Transformer, more complex feature fusion methods may be tried, such as weighted summation, cascading, and then passing through additional convolutional layers. However, these methods lack an in-depth understanding of the spatiotemporal characteristics of EEG signals, make it difficult to adaptively fuse features, and easily introduce additional computational overhead and parameters.

[0022] In the field of emotional brain-computer interfaces (aBCIs), it was found that existing deep learning-based emotion recognition methods have many problems in processing EEG signals. Through the analysis and research of a large amount of experimental data, it was found that the CNN-based method has obvious deficiencies in capturing long-distance temporal dependencies, and the Transformer-based method has certain advantages in processing temporal dependencies, but ignores spatial information. The method combining the two is not efficient enough in feature fusion. Based on these problems, we began to think about how to make full use of the spatiotemporal characteristics of EEG signals for emotion recognition.

[0023] This method focuses on enhancing EEG signal representation through three key modules, such as Figure 2 The figure shows the MSADM (Multi-Scale Attention-Based Dense Spatial-Temporal Model) of this method. TAM (Temporal Attention Module) effectively applies temporal attention by combining average pooling and max pooling operations. The multi-scale dense spatiotemporal feature extraction module captures temporal and spatial dependencies by processing EEG signals with 2D-dense blocks at multiple scales. Each has a tailored configuration for optimal feature extraction, which helps to maximize information flow between layers and improve gradient transmission. Finally, the attention-based multi-scale feature fusion module integrates features from different scales using a multi-head attention mechanism to refine the final representation of each classification task. Together, these modules enhance the model's ability to capture complex spatiotemporal patterns in EEG data. This method has the following two tasks, including: Task 1 classifies EEG patterns of positive and negative emotions induced by olfactory stimulation. Task 2 evaluates whether EEG signals from olfactory and non-olfactory stages are distinguishable.

[0024] For step S11, collect emotional stimulation EEG (Electroencephalogram) signal data, use specially designed bandpass filter (0.1 - 70Hz) and notch filter (50Hz) to filter the raw EEG data to remove noise interference and make the data purer. Then, downsample the signal to 200Hz, which not only reduces the amount of data, but also improves the efficiency of subsequent processing.

[0025] Among them, the emotional stimulation EEG data includes: EEG data of olfactory emotional stimulation. These emotional stimulation EEG data are obtained in the following way. According to the experimental design, the EEG data of each subject are segmented. Each experiment is divided into three stages: smelling, self-assessment, and resting. The 15-second EEG record of the smelling stage is selected as the data of Task 1, and the EEG data of the smelling stage and the resting breathing stage are selected as the data of Task 2, and each EEG segment is divided into non-overlapping 1-second samples. Finally, the data is divided into a training set and a test set in a ratio of 8:2, ready to be input into the model. In actual emotion recognition, the subject's EEG can be directly input for emotion recognition. The above preprocessing removes noise interference, reduces the amount of data, and makes the data cleaner, more stable, and more suitable for model processing.

[0026] For step S12, the emotional stimulation EEG data is input into TAM (Temporal Attention Module). For the input EEG sample, TAM performs two operations along the channel dimension to calculate features. In the temporal attention module, a first feature reflecting the average signal strength of the time step and a second feature of the maximum signal strength in each time step are determined.

[0027] In the specific implementation, average pooling and maximum pooling operations can be performed in the time attention module. The average pooling operation averages the signal values ​​of all channels at each time step to obtain a feature value that reflects the average signal strength of the time step; secondly, the maximum pooling operation finds the maximum value of all channels at each time step as the feature value of the time step. Then the feature maps obtained by these two pooling operations are concatenated, and then an attention map is generated through a 1D convolution layer (the convolution kernel size is 3 and is properly padded) and a sigmoid function. The function of this attention map is to assign a weight between 0 and 1 to each time step to indicate the importance of the time step. Finally, this attention map is used to re-weight the input signal along the time dimension, so that the model can pay more attention to those important time features, while suppressing unimportant time information, and obtain EEG data with enhanced attention.

[0028] By capturing significant patterns in the time dimension through average pooling and maximum pooling operations, and then using the generated attention map to reweight the input signal, the model is able to focus on key time features and suppress unimportant time information. This enables the model to better capture emotion-related temporal dynamic changes when processing EEG signals and improve the ability to recognize emotional patterns in time. For example, in the case of rapid emotional changes, TAM can highlight the signal features of the time points related to emotional transitions, thereby improving the accuracy of the model's emotion classification.

[0029] For step S13, the EEG data is input into a multi-scale dense spatiotemporal feature extraction module, which includes a combination layer of 2D-dense blocks and transition layers of multiple different scales (divided into three scales: small, medium and large). At each scale, the 2D-dense block is composed of multiple densely connected layers. Specifically, in each combination layer, the output and input of the nth 2D-dense block are used as the input of the n+1th 2D-dense block, and the transition layer after each 2D-dense block is used to reduce the number of feature channels and reduce the temporal resolution, so that the combination layer outputs comprehensive EEG features of different scales from the original input to the learning of each layer.

[0030] In this implementation, the output of each layer is concatenated with the input and then passed to the next layer. There are two important parameters in the layer, the bottleneck size and the growth rate. Taking a 2D-dense block with several layers as an example, the output of each layer is obtained by first performing a 1D convolution operation with a kernel size of 1 (this operation will change the number of channels according to the set bottleneck size and growth rate), then using the exponential linear unit (ELU) activation function, and then performing a 1D convolution operation with a kernel size of 3 (changing the number of channels again), and finally passing through the ELU activation function and Dropout operation (to prevent overfitting). The output of the entire 2D-dense block is to concatenate the input feature map and the output feature map of each layer in the channel dimension, thus forming a comprehensive feature map that contains rich feature information learned from the original input to each layer, and the dimension of the feature map will change at this time. Each 2D-dense block is followed by a transition layer, which first reduces the number of channels of the feature map through a 1D convolution operation (the convolution kernel size is 1, and the number of output channels becomes half of the number of input channels), and then performs an average pooling operation (kernel size is 2, step size is 2) to reduce the temporal resolution and obtain a new output feature map. After being processed by a set of 2D-dense blocks and transition layers, the new feature map can capture more complex spatiotemporal patterns, and the temporal resolution will be halved. After multiple such operations, the temporal resolution of the final feature map will be reduced to 1, and the dimensions of the final feature map obtained at different scales are the same. Through this multi-scale design, the model can learn the spatiotemporal features of EEG signals from different levels, enhance its adaptability and robustness to complex data, and then capture the comprehensive EEG features of spatiotemporal patterns at different scales.

[0031] By combining 2D-dense blocks and transition layers of different scales, the spatiotemporal features of EEG signals can be learned at different levels. Smaller growth rates and more layers help capture local detail features, while larger growth rates and fewer layers are conducive to learning global features. This multi-scale design enables the model to adapt to different types of EEG signal patterns and enhances the model's robustness to complex data distributions. For example, when processing EEG signals from different subjects or in different emotional states, multi-scale feature extraction can capture feature changes at different scales, thereby improving the generalization ability of the model.

[0032] For step S14, the EEG comprehensive features of different scales are dimensionally aligned and fused using an attention-based multi-scale feature fusion module. The EEG comprehensive features of different scales are spliced ​​along the preset new scale dimension. In order to align the dimensions, the tensor needs to be transposed first. Then, the query, key and value matrices are calculated using a multi-head attention mechanism, and the features are adaptively enriched through a specific calculation method to obtain a new feature representation. Next, the output of the attention mechanism is projected back to the original dimension through a dropout layer, and then a maximum pooling operation is performed along the scale dimension to generate a multi-scale fusion feature. This fusion feature integrates important features at different scales. The features of different scales are adaptively fused through a multi-head attention mechanism, and dynamically adjusted according to the importance of the features, so that the model can make full use of the advantages of features of different scales and generate more representative fusion features. This improves the accuracy of the final classification of the model because the fused features contain richer and more comprehensive information. For example, when distinguishing similar emotional states, multi-scale feature fusion can integrate feature differences at different scales, so that the model can more accurately judge the emotional category.

[0033] For step S15, the multi-scale fusion features are passed through a fully connected layer and an exponential linear unit (ELU) activation function, and then classified emotion recognition is performed to obtain the emotion recognition result.

[0034] Through this implementation, it can be seen that the difference and advantages of this method and the CNN-based method are: unlike the traditional CNN method which is limited to fixed receptive fields and local feature extraction, the multi-scale design and temporal attention mechanism of MSADM can better capture long-distance temporal dependencies and global information. For example, when processing long-term EEG signals, it can pay attention to the dynamic changes between different moments and improve the accuracy of emotion recognition. The difference and advantages of the method based on Transformer: Compared with the Transformer method, MSADM explicitly considers spatial information through the multi-scale feature extraction module, and adopts a more flexible attention mechanism when fusing features, avoiding the Transformer's neglect of spatial information. In practical applications, it can better utilize the spatial relationship between different channels in the EEG signal and improve the model's ability to recognize emotion-related patterns. The difference and advantages of the method combining CNN and Transformer: Compared with the method of simply combining CNN and Transformer, the design of each module of MSADM is more closely combined with the characteristics of EEG signals and the requirements of emotion recognition tasks. The parameter configuration of the multi-scale feature extraction module and the application of the attention mechanism make feature fusion more efficient, the model complexity is better controlled, the overfitting risk is reduced, and the generalization ability of the model is improved. Therefore, this method can effectively solve the defects of the existing technology in EEG signal emotion recognition, improve the accuracy, stability and adaptability of the model, and provide a more effective emotion recognition solution for the field of emotional brain-computer interface.

[0035] like Figure 3 FIG. 1 is a flow chart of a method for training a multi-scale attention emotion recognition model provided by an embodiment of the present invention, comprising the following steps: S21: Acquire emotional stimulation EEG data for training, and input the emotional stimulation EEG data into the emotion recognition model, wherein the emotion recognition model includes: a time attention module, a multi-scale dense spatiotemporal feature extraction module, and an attention-based multi-scale feature fusion module; S22: determining in the time attention module a first feature for reflecting the average signal strength of a time step and a second feature of the maximum signal strength in each time step, using an attention map generated by splicing the first feature and the second feature, and weighting the emotion stimulation EEG data along the time dimension based on the attention map to obtain attention-enhanced EEG data; S23: Input the EEG data into a multi-scale dense spatiotemporal feature extraction module, wherein the multi-scale dense spatiotemporal feature extraction module includes a plurality of combination layers of different scales, each combination layer includes a plurality of 2D-dense blocks and a transition layer, in each combination layer, the output of the nth 2D-dense block and the EEG data are used as the input of the n+1th 2D-dense block, and the transition layer after each 2D-dense block is used to reduce the number of feature channels and reduce the temporal resolution, so that the combination layer outputs comprehensive EEG features of different scales from the original input to the learning of each layer; S24: dimensionally aligning and fusing the EEG comprehensive features of different scales in the attention-based multi-scale feature fusion module to generate multi-scale fusion features; S25: Determine a predicted emotion recognition result through the multi-scale fusion feature, and train the emotion recognition model based on a difference between a preset baseline emotion recognition result and the predicted emotion recognition result until a preset goal is achieved.

[0036] In this embodiment, model training is performed based on a multi-scale attention emotion recognition method. The repeated recognition steps are not described here. The structure and training steps of each module in the model training are described in detail as follows.

[0037] For step S21, after obtaining the emotion stimulation EEG data for training, the method further includes: filtering and downsampling the emotion stimulation EEG data to obtain clean emotion stimulation EEG data; the acquisition of emotion stimulation EEG data for training includes: testing the subjects in three stages: odor stimulation, self-assessment, and rest, to obtain EEG data of olfactory emotion stimulation and baseline emotion recognition results. This ensures the accuracy of subsequent model training and prediction, because high-quality data is the basis for effective model learning and accurate classification. For example, in experiments, if filtering is not performed, noise may mask weak emotion-related features in the EEG signal, making it difficult for the model to accurately identify emotional states.

[0038] For step S22, regarding the temporal attention module, TAM aims to enhance the representation ability of temporal features by focusing on the elements with the most information in the temporal dimension. This method introduces a temporal attention mechanism for EEG signals. Given an EEG sample , where C represents the number of channels and T represents the time step. TAM applies attention to adaptively reweight features. First, implicit and salient temporal patterns are captured by utilizing average pooling and max pooling operations. This method computes the average pooled features along the channel dimension. and the maximum merged features : The two feature maps are concatenated to form a combined feature map, which is then passed through a 1D convolutional layer with a kernel size of 3 and appropriate padding to generate the attention map. : Where σ(·) represents the sigmoid function, which ensures that the attention value is within [0,1]. Then, the attention map A is used to reweight the input signal X along the time dimension. The output of TAM Defined as: This operation enhances relevant temporal features while suppressing less informative features, allowing the model to focus on key components in the EEG signal. Subsequently, convolution and pooling operations are applied to better fit the 2D-dense block and reduce the temporal resolution.

[0039] For step S23, regarding the multi-scale dense spatiotemporal feature extraction module, unlike the three-dimensional matrix representation of images, EEG signals are represented as two-dimensional matrices. Therefore, this method specifically designs 2D-dense blocks and transition layers for EEG signals to learn dense embedding. Given the optimized EEG signal , the model can effectively capture local and global temporal dependencies across channels through N groups of 2D-dense blocks and transition layers.

[0040] 2D-dense blocks consist of multiple densely connected layers, where the output of each layer is connected to its input and passed to the next layer. Each layer contains two important parameters: bottleneck size and growth rate. The bottleneck size bn is used to reduce the number of input feature maps, while the growth rate gr controls the amount of information added by each layer. The higher the gr, the richer the feature representation, while the lower the gr, the more compact the network.

[0041] For a block with L layers, the output of layer 1 is represented as , is given by: Among them, Conv1D1 has a kernel size of 1 and an input channel Conv1D2 is a 1D convolution with kernel size 3, input channels bn×gr, and output channels gr. is the exponential linear unit (ELU) activation function. Let is the input of the block, and the output of the entire 2D-dense block is the concatenation of all intermediate feature maps. , where L = C / gr. After the 2D-dense block, the feature map belong .

[0042] Each 2D-dense block is paired with a transition layer to reduce the number of channels and the temporal resolution of the feature map. This layer applies 1D convolution followed by average pooling, outputting The calculation is as follows: Among them, the kernel size of Conv1D is 1 and the output channels are half of the input channels. AvgPool with a kernel size of 2 and a stride of 2 reduces the temporal resolution by half.

[0043] After passing through a set of 2D-dense blocks and transition layers, the new feature maps capture more complex spatiotemporal patterns in the EEG signal while keeping the same number of channels and halving the temporal resolution. The final temporal resolution of is reduced to 1 by processing through N 2D-dense blocks and transition layers.

[0044] This method does not simply set gr to four times that of bn. The purpose of configuring bn is to reduce the number of parameters, while a higher gr will increase the number of feature maps faster, resulting in higher computational cost. These two goals are inherently conflicting. To address this issue, this method introduces a multi-scale dense module, where bn and gr are set differently at each scale, but their product remains unchanged. This ensures that the intermediate embeddings of networks with different depths and growth rates are mapped to the same dimensional space.

[0045] Specifically, this method classifies the growth rate gr into three levels: small, medium, and large, from low to high. For each gr, the bottleneck size bn and the number of layers L within the 2D-dense block are adjusted accordingly. The values ​​of gr, bn, and L for the three scales are set as follows: {8, 32, 8}, {16, 16, 4}, and {32, 8, 2}, respectively. A larger gr allows the model to use more convolutional filters, helps to learn richer feature representations, and search for features more widely in the space, which is crucial for capturing complex patterns and details. On the other hand, a larger L corresponds to a deeper network, enabling the model to express more complex functions and decision boundaries. Therefore, deeper layers capture high-level features and gradually build complex feature representations. With our parameter configuration, the final feature map size obtained at all three scales is consistent, i.e. By combining the results of three scales, the limitations of a single scale are effectively alleviated, and the adaptability and robustness of the model in dealing with different data distributions are enhanced.

[0046] For step S24, regarding the multi-scale feature fusion module, the attention-based cross-scale feature fusion module aims to integrate features of different scales through the attention mechanism, adaptively enrich features, and finally generate the final output through the classification layer. The first step of this module is to connect the input embedding along the new scale dimension. Since each embedding has a single channel dimension, the concatenation is performed after permuting the tensors to align their dimensions correctly. By using a multi-head attention mechanism, this module can capture various relationships between elements in the input sequence and compute the query Q, key K, and value V matrices. Embedding The calculation is as follows: Where dk is the dimension of the key vector. The output of the attention mechanism is then projected to the original dimension C with a dropout layer, followed by maximum pooling on the scale dimension to produce multi-scale fused features .

[0047] For step S25, finally, the loss is calculated by binary cross entropy loss: in, and y are the predicted emotion recognition results and the benchmark emotion recognition results, respectively.

[0048] It can be seen from this implementation that by applying the attention mechanism in the time dimension and combining multi-scale dense feature extraction, the model can focus on key time features while capturing rich spatiotemporal features at different scales. Aiming at the two-dimensional matrix representation and spatiotemporal characteristics of EEG signals, a 2D-dense block, transition layer and multi-scale feature fusion method are specially designed to effectively process EEG signals and improve emotion recognition performance.

[0049] Experimental description of this method. This method designs two tasks (task 1 and task 2 in the above steps) to explore the relationship between olfactory stimulation and EEG patterns. In order to ensure the rationality of subsequent experiments, as well as the quality and reliability of the data set, appropriate stimulus materials are required. Therefore, this method recruited 50 subjects for a pilot experiment to refine our selection of stimuli. After sniffing each odor, the subjects rated the odor according to their emotional response and intensity. Positive and negative scores indicate whether the odor will cause pleasant or unpleasant emotions. The higher the absolute value, the stronger the reaction. Based on the distribution and central tendency of the data, this method selects the two odors with the lowest scores as negative stimuli and the two tastes with the highest scores as positive stimuli.

[0050] In the formal experiment, this method recruited 32 subjects (16 males; mean age: 24.19). All subjects were right-handed, with normal or corrected-to-normal vision and normal sense of smell. All subjects were fully informed of the experimental procedures before the experiment and signed informed consent. To ensure data quality, the experiment was conducted in a controlled laboratory environment to minimize noise and other environmental interference. EEG signals were collected using an ESI neuroscan system with a 62-channel electrode cap (1000 Hz) according to the international 10-20 system. Each subject participated in three experiments with an interval of more than 24 hours to achieve data independence. Each session consisted of 6 rounds, each of which consisted of 4 trials, and each round of the experiment randomly presented four different olfactory stimuli. Each trial was divided into three stages. The first stage was to smell the odor, which was released continuously for 15 seconds by an air pump. The second stage was self-assessment, in which the subjects recorded their emotional reactions as positive or negative. The third stage was to rest while breathing fresh air. Before and after each odor release, corresponding prompts were given to remind the subjects.

[0051] Following the above procedure, the present method obtained three continuous complete EEG recordings for each subject. For each trial, 15 seconds of EEG recordings from the olfactory phase and the corresponding self-labeled emotions were selected as the raw data and labels for Task 1. For Task 2, in addition to the EEG recordings from the first olfactory phase, EEG data corresponding to the third resting breathing phase were selected. The labels for this task were naturally derived based on whether the EEG data was recorded during the olfactory or non-olfactory phase. Each EEG segment was divided into non-overlapping 1-second samples. The data were divided into training and test sets in a ratio of 8:2. The training and test sets and the divisions were consistent across all methods. A bandpass filter (0.1-70Hz) and a notch filter (50Hz) were applied, and the signal was downsampled to 200Hz.

[0052] This method sets the number of 2D-dense blocks to 5, the batch size to 64, and the dropout rate to 0.5. The learning rate is set to 1e-3, the weight decay is 1e-5, and the Adam optimizer is applied.

[0053] All methods are CNN-based or Transformer-based advanced algorithms designed for the task of raw EEG signal classification. Each method follows the same hyperparameter settings. For each subject, the accuracy and area under the receiver operating characteristic curve (AUROC) are calculated. The mean and standard deviation are calculated for all subjects.

[0054] Compared with the baseline: Figure 4As shown, our MSADM model consistently outperforms other models on evaluation metrics and is a very robust choice in terms of both effectiveness and stability. The results for ST-Trans show that its approach to capturing spatiotemporal features has specific limitations. The performance of pre-trained BIOT shows that while pre-training may not significantly improve accuracy, it does improve the stability of predictions. FFCL and ContraWR achieve reasonable accuracy but lack sufficient discriminative power. While SPaRCNet and CNNTrans perform well in specific tasks, their overall results are still slightly lower than our model and reveal limitations in cross-task generalization. MSADM has consistently shown robustness and effectiveness in both tasks, outperforming other models in cross-task generalization, especially when dealing with more complex evaluation thresholds.

[0055] This method conducts ablation experiments on the more challenging task 1. Figure 5 As shown in Figure 3, the results show that with the gradual integration of multi-scale, TAM, and attention-based multi-scale feature fusion modules, the performance of the model continues to improve in all indicators. The performance of small, medium, and large modules is relatively similar, indicating that each scale provides limited performance enhancement when used independently, with minimal changes across scales. However, applying multi-scale dense feature extraction can significantly improve performance, demonstrating the effectiveness of multi-scale modules in leveraging different scales to improve overall performance. Combining TAM can further improve performance, indicating that TAM enhances task-related feature channels by assigning different weights to each channel, thereby reducing the impact of less relevant channels and emphasizing key features. Finally, the addition of the fusion module brings the highest performance, demonstrating its ability to fully integrate the extracted features and significantly improve the model's ability to distinguish the final predictions.

[0056] Overall, this method captures spatiotemporal dependencies in EEG signals through temporal attention, multi-scale feature extraction, and attention-based fusion. Combining olfactory stimulation provides a novel and stable emotion induction paradigm. MSADM surpasses existing methods and facilitates the development of personalized aBCIs, while opening up new avenues for olfactory-based emotion induction in various applications.

[0057] like Figure 6 The figure shows a schematic diagram of the structure of a multi-scale attention emotion recognition system provided by an embodiment of the present invention. The system can execute the multi-scale attention emotion recognition method described in any of the above embodiments and be configured in a terminal.

[0058] A multi-scale attention emotion recognition system 10 provided in this embodiment includes: an EEG acquisition module 11, a signal highlighting module 12, a multi-scale feature capturing module 13, a fusion module 14 and a recognition module 15.

[0059] Among them, the EEG acquisition module 11 is used to acquire emotion stimulation EEG data; the signal highlighting module 12 is used to input the emotion stimulation EEG data into the time attention module, determine the first feature used to reflect the average signal strength of the time step, and the second feature of the maximum signal strength in each time step in the time attention module, use the first feature and the second feature to generate an attention map, and weight the emotion stimulation EEG data along the time dimension based on the attention map to obtain attention-enhanced EEG data; the multi-scale feature capture module 13 is used to input the EEG data into the multi-scale dense spatiotemporal feature extraction module, wherein the multi-scale dense spatiotemporal feature extraction module includes a plurality of combination layers of different scales, wherein each combination layer includes a plurality of 2D-dense blocks and transition layers, and outputs EEG comprehensive features of different scales that capture spatiotemporal patterns; the fusion module 14 is used to dimensionally align and fuse the EEG comprehensive features of different scales to generate multi-scale fusion features; the recognition module 15 is used to determine the emotion recognition result through the multi-scale fusion features.

[0060] The embodiment of the present invention further provides a non-volatile computer storage medium, the computer storage medium stores computer executable instructions, and the computer executable instructions can execute the multi-scale attention emotion recognition method in any of the above method embodiments; As an implementation mode, the non-volatile computer storage medium of the present invention stores computer executable instructions, and the computer executable instructions are configured as follows: Acquire EEG data of emotional stimulation; Inputting the emotion stimulation EEG data into a time attention module, determining in the time attention module a first feature for reflecting the average signal strength of a time step and a second feature of the maximum signal strength in each time step, using an attention map generated by splicing the first feature and the second feature, and weighting the emotion stimulation EEG data along the time dimension based on the attention map to obtain attention-enhanced EEG data; Inputting the EEG data into a multi-scale dense spatiotemporal feature extraction module, wherein the multi-scale dense spatiotemporal feature extraction module includes a plurality of combination layers of different scales, wherein each combination layer includes a plurality of 2D-dense blocks and transition layers, and outputs EEG comprehensive features of different scales that capture spatiotemporal patterns; Dimensionally aligning and fusing the EEG comprehensive features of different scales to generate multi-scale fusion features; The emotion recognition result is determined by the multi-scale fusion feature.

[0061] like Figure 7The figure shows a structural diagram of a training system for a multi-scale attention emotion recognition model provided by an embodiment of the present invention. The system can execute the training method for the multi-scale attention emotion recognition model described in any of the above embodiments and be configured in a terminal.

[0062] A training system 20 for a multi-scale attention emotion recognition model provided in this embodiment includes: an EEG acquisition module 21, a signal highlighting module 22, a multi-scale feature capture module 23, a fusion module 24 and a training module 25.

[0063] Among them, the EEG acquisition module 21 is used to obtain emotional stimulation EEG data for training, and input the emotional stimulation EEG data into the emotion recognition model, wherein the emotion recognition model includes: a time attention module, a multi-scale dense spatiotemporal feature extraction module and an attention-based multi-scale feature fusion module; the signal highlighting module 22 is used to determine a first feature used to reflect the average signal strength of the time step and a second feature with the maximum signal strength in each time step in the time attention module, and use the attention map generated by splicing the first feature and the second feature to weight the emotional stimulation EEG data along the time dimension based on the attention map to obtain attention-enhanced EEG data; the multi-scale feature capture module 23 is used to input the EEG data into the multi-scale dense spatiotemporal feature extraction module, wherein the multi-scale dense spatiotemporal feature extraction module includes a plurality of combination layers of different scales, each combination layer includes a plurality of 2D-dense blocks and a transition layer, in each combination layer, the output of the nth 2D-dense block and the EEG data are used as the input of the n+1th 2D-dense block, and in each 2D-dense The transition layer after the block is used to reduce the number of feature channels and reduce the temporal resolution, so that the combination layer outputs EEG comprehensive features of different scales learned from the original input to each layer; the fusion module 24 is used to align and fuse the EEG comprehensive features of different scales in the attention-based multi-scale feature fusion module to generate multi-scale fusion features; the training module 25 is used to determine the emotion recognition results through the multi-scale fusion features.

[0064] The embodiment of the present invention further provides a non-volatile computer storage medium, the computer storage medium stores computer executable instructions, and the computer executable instructions can execute the training method of the multi-scale attention emotion recognition model in any of the above method embodiments; As an implementation mode, the non-volatile computer storage medium of the present invention stores computer executable instructions, and the computer executable instructions are configured as follows: Acquire emotional stimulation EEG data for training, and input the emotional stimulation EEG data into the emotion recognition model, wherein the emotion recognition model includes: a time attention module, a multi-scale dense spatiotemporal feature extraction module, and an attention-based multi-scale feature fusion module; Determining in the temporal attention module a first feature for reflecting the average signal strength of a time step and a second feature of the maximum signal strength in each time step, using an attention map generated by splicing the first feature and the second feature, and weighting the emotion stimulation EEG data along the time dimension based on the attention map to obtain attention-enhanced EEG data; Input the EEG data into a multi-scale dense spatiotemporal feature extraction module, wherein the multi-scale dense spatiotemporal feature extraction module includes a plurality of combination layers of different scales, each combination layer includes a plurality of 2D-dense blocks and a transition layer, in each combination layer, the output of the nth 2D-dense block and the EEG data are used as the input of the n+1th 2D-dense block, and the transition layer after each 2D-dense block is used to reduce the number of feature channels and reduce the temporal resolution, so that the combination layer outputs comprehensive EEG features of different scales from the original input to the learning of each layer; In the attention-based multi-scale feature fusion module, dimensionally aligning and fusing the EEG comprehensive features of different scales to generate multi-scale fusion features; The predicted emotion recognition result is determined by the multi-scale fusion feature, and the emotion recognition model is trained based on the difference between a preset benchmark emotion recognition result and the predicted emotion recognition result until a preset goal is achieved.

[0065] As a non-volatile computer-readable storage medium, it can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions / modules corresponding to the method in the embodiment of the present invention. One or more program instructions are stored in the non-volatile computer-readable storage medium, and when executed by the processor, the multi-scale attention emotion recognition method and the multi-scale attention emotion recognition model training method in any of the above method embodiments are executed.

[0066] Figure 8 FIG. 1 is a schematic diagram of the hardware structure of an electronic device of a multi-scale attention emotion recognition method provided by another embodiment of the present application, such as Figure 8 As shown, the device includes: One or more processors 810 and memory 820, Figure 8 A processor 810 is taken as an example. The device of the multi-scale attention emotion recognition method may also include: an input device 830 and an output device 840.

[0067] The processor 810, the memory 820, the input device 830 and the output device 840 may be connected via a bus or other means. Figure 8 The example of connecting through bus is taken in the following.

[0068] The memory 820, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions / modules corresponding to the multi-scale attention emotion recognition method in the embodiment of the present application. The processor 810 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions and modules stored in the memory 820, that is, the multi-scale attention emotion recognition method and the training method of the multi-scale attention emotion recognition model in the above method embodiment are implemented.

[0069] The memory 820 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data, etc. In addition, the memory 820 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 820 may optionally include a memory remotely arranged relative to the processor 810, and these remote memories may be connected to the mobile device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0070] The input device 830 can receive input digital or character information. The output device 840 can include a display device such as a display screen.

[0071] The one or more modules are stored in the memory 820, and when executed by the one or more processors 810, perform the multi-scale attention emotion recognition method and the training method of the multi-scale attention emotion recognition model in any of the above method embodiments.

[0072] The above-mentioned product can execute the method provided in the embodiment of the present application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of the present application.

[0073] The non-volatile computer-readable storage medium may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the device, etc. In addition, the non-volatile computer-readable storage medium may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the non-volatile computer-readable storage medium may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0074] An embodiment of the present invention also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the steps of the multi-scale attention emotion recognition method of any embodiment of the present invention.

[0075] The electronic device of the embodiment of the present application exists in various forms, including but not limited to: (1) Mobile communication equipment: This type of equipment is characterized by having mobile communication functions and its main purpose is to provide voice and data communications. This type of terminal includes: smart phones, multimedia phones, functional phones, and low-end phones.

[0076] (2) Ultra-mobile personal computer devices: These devices fall into the category of personal computers, have computing and processing capabilities, and generally also have mobile Internet access features. These terminals include: PDAs, MIDs, and UMPC devices, such as tablet computers.

[0077] (3) Portable entertainment devices: These devices can display and play multimedia content. They include audio and video players, handheld game consoles, e-books, smart toys, and portable car navigation devices.

[0078] (4) Other electronic devices with data processing functions.

[0079] In this article, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include" and "comprise" include not only those elements, but also other elements not explicitly listed, or also include elements inherent to such processes, methods, articles or equipment. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the existence of other identical elements in the process, method, article or equipment that includes the elements.

[0080] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0081] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-scale attention emotion recognition method, comprising: Acquire EEG data of emotional stimulation; Inputting the emotion stimulation EEG data into a time attention module, determining in the time attention module a first feature for reflecting the average signal strength of a time step and a second feature of the maximum signal strength in each time step, using an attention map generated by splicing the first feature and the second feature, and weighting the emotion stimulation EEG data along the time dimension based on the attention map to obtain attention-enhanced EEG data; Inputting the EEG data into a multi-scale dense spatiotemporal feature extraction module, wherein the multi-scale dense spatiotemporal feature extraction module includes a plurality of combination layers of different scales, wherein each combination layer includes a plurality of 2D-dense blocks and transition layers, and outputs EEG comprehensive features of different scales that capture spatiotemporal patterns; Dimensionally aligning and fusing the EEG comprehensive features of different scales to generate multi-scale fusion features; The emotion recognition result is determined by the multi-scale fusion feature.

2. The method according to claim 1, wherein: The step of dimensionally aligning and fusing the EEG comprehensive features of different scales to generate multi-scale fusion features includes: splicing the EEG comprehensive features of different scales, performing tensor transposition processing after the splicing, and performing calculation of a multi-head attention mechanism after the tensor transposition processing to obtain a feature representation representing the emotional stimulus; The feature representation is projected back to the original dimension through a random dropout layer, and then a maximum pooling process is performed to generate a multi-scale fusion feature that integrates feature differences at different scales.

3. The method according to claim 1, wherein: The EEG data is input into a multi-scale dense spatiotemporal feature extraction module, and the output of the EEG comprehensive features of different scales capturing the spatiotemporal pattern includes: In each combination layer, the output of the nth 2D-dense block and the input of the nth 2D-dense block are used as the input of the n+1th 2D-dense block. The transition layer after each 2D-dense block is used to reduce the number of feature channels and reduce the temporal resolution, so that the combination layer outputs comprehensive EEG features of different scales learned from the original input to each layer.

4. The method according to claim 1, wherein: The emotion stimulation EEG data includes: EEG data of olfactory emotion stimulation.

5. A training method for a multi-scale attention emotion recognition model, comprising: Acquire emotional stimulation EEG data for training, and input the emotional stimulation EEG data into the emotion recognition model, wherein the emotion recognition model includes: a time attention module, a multi-scale dense spatiotemporal feature extraction module, and an attention-based multi-scale feature fusion module; Determining in the temporal attention module a first feature for reflecting the average signal strength of a time step and a second feature of the maximum signal strength in each time step, using an attention map generated by splicing the first feature and the second feature, and weighting the emotion stimulation EEG data along the time dimension based on the attention map to obtain attention-enhanced EEG data; Input the EEG data into a multi-scale dense spatiotemporal feature extraction module, wherein the multi-scale dense spatiotemporal feature extraction module includes a plurality of combination layers of different scales, each combination layer includes a plurality of 2D-dense blocks and a transition layer, in each combination layer, the output of the nth 2D-dense block and the EEG data are used as the input of the n+1th 2D-dense block, and the transition layer after each 2D-dense block is used to reduce the number of feature channels and reduce the temporal resolution, so that the combination layer outputs comprehensive EEG features of different scales from the original input to the learning of each layer; In the attention-based multi-scale feature fusion module, dimensionally aligning and fusing the EEG comprehensive features of different scales to generate multi-scale fusion features; The predicted emotion recognition result is determined by the multi-scale fusion feature, and the emotion recognition model is trained based on the difference between a preset benchmark emotion recognition result and the predicted emotion recognition result until a preset goal is achieved.

6. The method according to claim 5, wherein: After acquiring the emotion stimulation EEG data for training, the method further includes: Filtering and down-sampling the emotion-stimulated EEG data to obtain clean emotion-stimulated EEG data; The acquisition of emotional stimulation EEG data for training includes: conducting three-stage tests on the subject, namely, odor stimulation, self-assessment, and rest, to obtain EEG data of olfactory emotional stimulation and baseline emotion recognition results.

7. A multi-scale attention emotion recognition system, comprising: EEG acquisition module, used to acquire EEG data of emotional stimulation; A signal highlighting module, used for inputting the emotion stimulation EEG data into a time attention module, determining in the time attention module a first feature for reflecting the average signal strength of a time step and a second feature of the maximum signal strength in each time step, using an attention map generated by splicing the first feature and the second feature, and weighting the emotion stimulation EEG data along the time dimension based on the attention map to obtain attention-enhanced EEG data; A multi-scale feature capture module, used for inputting the EEG data into a multi-scale dense spatiotemporal feature extraction module, wherein the multi-scale dense spatiotemporal feature extraction module comprises a plurality of combination layers of different scales, wherein each combination layer comprises a plurality of 2D-dense blocks and transition layers, and outputs EEG comprehensive features of different scales capturing spatiotemporal patterns; A fusion module is used to align and fuse the EEG comprehensive features of different scales to generate multi-scale fusion features; The recognition module is used to determine the emotion recognition result through the multi-scale fusion feature.

8. A training system for a multi-scale attention emotion recognition model, comprising: An EEG acquisition module, used to acquire emotional stimulation EEG data for training, and input the emotional stimulation EEG data into the emotion recognition model, wherein the emotion recognition model includes: a time attention module, a multi-scale dense spatiotemporal feature extraction module, and an attention-based multi-scale feature fusion module; A signal highlighting module, configured to determine in the time attention module a first feature for reflecting the average signal strength of a time step and a second feature of the maximum signal strength in each time step, and to weight the emotion stimulation EEG data along the time dimension based on the attention map by using the attention map generated by splicing the first feature and the second feature, so as to obtain attention-enhanced EEG data; A multi-scale feature capture module, used for inputting the EEG data into a multi-scale dense spatiotemporal feature extraction module, wherein the multi-scale dense spatiotemporal feature extraction module includes a plurality of combination layers of different scales, each combination layer includes a plurality of 2D-dense blocks and a transition layer, in each combination layer, the output of the nth 2D-dense block and the EEG data are used as the input of the n+1th 2D-dense block, and the transition layer after each 2D-dense block is used to reduce the number of feature channels and reduce the temporal resolution, so that the combination layer outputs comprehensive EEG features of different scales from the original input to the learning of each layer; A fusion module, used for dimensional alignment and fusion of the EEG comprehensive features of different scales in the attention-based multi-scale feature fusion module to generate multi-scale fusion features; The training module is used to determine the predicted emotion recognition result through the multi-scale fusion feature, and train the emotion recognition model based on the difference between the preset baseline emotion recognition result and the predicted emotion recognition result until a preset goal is achieved.

9. A storage medium having a computer program product stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.

10. A computer program product having instructions embedded on a storage medium, wherein the instructions implement the steps of the method according to any one of claims 1 to 6.

11. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method described in any one of claims 1 to 6.