Multi-scale emotion recognition method based on DBBCapsNe model

By introducing the DBBCapsNe model in EEG emotion recognition, combining diverse branch modules, deep separable convolution, SEBlock and capsule network, the problem of insufficient feature extraction and model generalization capabilities in the existing technology is solved, and multi-scale emotion recognition with high accuracy and strong generalization capabilities is achieved.

CN120144920APending Publication Date: 2025-06-13CHONGQING UNIV OF TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510229729.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing EEG emotion recognition method based on capsule networks has shortcomings in the comprehensiveness of feature extraction and model generalization ability, and it is difficult to fully capture the deep-seated information required for complex tasks, which affects the final performance of the model.

Method used

A multi-scale emotion recognition method based on the DBBCapsNe model is proposed, combining the diversified branch module (DBB), depth separable convolution, SEBlock channel attention mechanism and capsule network, multi-scale temporal features are extracted through parallel multi-branch structure and depth separable convolution, and spatial information and multi-level features are integrated through SEBlock and capsule network.

Benefits of technology

It significantly improves the accuracy of emotion recognition, enhances the generalization ability of the model, reduces the computational complexity, and realizes end-to-end multi-scale feature extraction and emotion classification to achieve the current state-of-the-art performance level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144920A_ABST
    Figure CN120144920A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale emotion recognition method based on a DBBCapsNe model, and relates to the technical field of electroencephalogram signal processing and emotion recognition. According to the method, the diversified branch modules, the depth separable convolution, the SEBlock channel attention mechanism and the capsule network are combined, and efficient feature extraction and emotion classification of the electroencephalogram signals are achieved. According to the method, through an end-to-end model architecture, the whole network is divided into a time feature extractor and a capsule network, the time feature extractor module effectively extracts time domain features through a diversified branch module and depth separable convolution, and SEBlock is integrated, so that weight connection between channels is enhanced, the feature expression ability is further improved, and the time domain feature extraction efficiency is improved. And inputting the extracted feature information into the capsule network to extract spatial domain features and effectively integrate local and global relationships.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electroencephalogram signal processing and emotion recognition, and specifically to a multi-scale emotion recognition method based on the DBBCapsNe model. Background Technique

[0002] The electroencephalogram emotion recognition technology judges an individual's emotional state by analyzing the electrical signals (electroencephalogram waves) generated by the brain. This technology is of great significance in the diagnosis of medical diseases, such as helping to diagnose emotional disorders (such as depression, post-traumatic stress disorder PTSD, etc.). The sources of emotion signals can be divided into physiological signals (such as electroencephalogram EEG, electrooculogram EOG, electrocardiogram ECG, electromyogram EMG) and non-physiological signals (such as facial expressions, language, body postures). However, non-physiological signals (such as speech, facial expressions, and postures) are easily covered up or disguised by individuals intentionally or unintentionally, resulting in inaccurate emotion recognition based on these signals. In contrast, electroencephalogram, as a central nervous physiological signal, has the advantages of objectivity and difficulty in being covered up, and is particularly suitable for emotion recognition. In addition, due to its non-invasive, easy-to-use, and low-cost characteristics, EEG has received more and more attention in the research and application of emotion recognition.

[0003] Emotion recognition models are mainly divided into discrete emotion models and dimensional emotion models. The discrete emotion model defines emotions as finite independent categories, such as the six basic emotions (happiness, sadness, fear, anger, surprise, and disgust) proposed by Paul Ekman, but it is difficult to cover complex or mixed emotions. The dimensional emotion model describes emotions through a multi-dimensional space, such as Russell's Circumplex Model, with valence and arousal as the core dimensions, and can accurately describe complex emotional states, such as anxiety (low valence, high arousal) and satisfaction (high valence, low arousal). The dimensional model has significant advantages in emotion intensity measurement, cross-cultural comparison, and data characteristic matching with physiological signals, and has become the core tool for emotion research.

[0004] In recent years, with the development of deep learning technology, researchers have proposed a variety of EEG-based emotion recognition methods, including convolutional neural networks (CNNs), recurrent neural networks (RNNs), graph neural networks (GNNs), and hybrid neural networks. These methods have made certain progress in extracting spatial and temporal features of EEG signals. For example, Asif et al. used the short-time Fourier transform to extract features and combined a classification model composed of a CNN-LSTM hybrid layer to achieve an accuracy of 95.65% on the DEAP dataset. Yao et al. constructed a new multi-feature fusion model for cross-channel spatial information and cross-frequency band information. By combining the power of different sub-bands and the electrode spatial images to construct a three-dimensional multi-spectral image for training, the information in the EEG signals was fully explored. However, traditional CNNs have limitations in characterizing feature coupling. The fixed receptive field limits the ability to capture global features, and the pooling operation may lead to information loss. In addition, existing architectures generally ignore the inherent spatial topological characteristics of EEG signals, affecting the final performance of the model.

[0005] The emergence of Capsule Network provides a new technical path for modeling the spatial features of EEG signals. Through its unique vector neuron structure and dynamic routing mechanism, Capsule Network can effectively capture the spatial topological relationship and feature pose parameters of EEG signals, making up for the deficiencies of traditional convolutional neural networks in spatial information modeling. However, existing Capsule Network-based EEG emotion recognition methods still have deficiencies in the comprehensiveness of feature extraction and the model generalization ability, making it difficult to fully capture the deep information required for complex tasks and affecting the final performance of the model.

[0006] Therefore, a new solution to the above problems is needed. Summary of the Invention

[0007] The purpose of the present invention is to provide a multi-scale emotion recognition method based on the DBBCapsNe model, aiming to improve the recognition accuracy of emotion states in EEG signals and solve the technical problems proposed in the background technology.

[0008] To achieve the above purpose, the present invention provides the following technical solution: A multi-scale emotion recognition method based on the DBBCapsNe model, at least including the following steps:

[0009] S1: Select the DEAP and DREAMER public datasets;

[0010] S2: Perform data preprocessing to standardize the signals into a format suitable for model input. The preprocessing includes, but is not limited to, downsampling and band-pass filtering of EEG signals;

[0011] S3: Construct the DBBCapsNet model, where the DBBCapsNet model includes a temporal feature extractor and a capsule network. The temporal feature extractor includes a Diversified Branch Block (DBB) module, depthwise separable convolutions, and a Squeeze-and-Excitation (SE) Block. The Diversified Branch Block is the DBB module, and the SE Block is the channel attention mechanism;

[0012] S4: Use 10-fold cross-validation to train the DBBCapsNet model and evaluate its performance;

[0013] S5: Apply the trained DBBCapsNet model for application, and then output the emotion classification results.

[0014] Furthermore, the DBB module consists of four parallel branches, which respectively use different convolutional kernels and pooling operations to extract multi-scale temporal features, and stabilize the training through batch normalization. The specific structure of the DBB module is as follows:

[0015] The first branch: Use a 3x3 convolutional kernel for feature extraction;

[0016] The second branch: Use a 1x1 convolution to adjust the number of channels;

[0017] The third branch: First use a 1x1 convolution to reduce the number of channels, and then apply average pooling;

[0018] The fourth branch: First use a 1x1 convolution to reduce the number of channels, and then use a 6x6 convolutional kernel to extract larger-scale local features;

[0019] After each convolution and pooling operation, the batch normalization layer is used to stabilize the training of the network, further improving the robustness and generalization ability of the model;

[0020] Among them, the outputs of each branch are merged through a stacking operation (concatenate) to enhance the richness of feature expression. Then, the DBB module extracts features from different scales and perspectives through a multi-branch structure, reducing the computational amount and preventing overfitting;

[0021] In the DBB module, the input is set as: the input dimension is I ∈ R C×H×W , where C is the number of channels, H is the height, and W is the width;

[0022] The branch outputs are set as: B 1 , B 2 , B 3 and B 4 respectively represent the output dimensions of each branch, where B i ∈ R D ×H′×W′ , where H' and W' are the new heights and widths obtained after processing H and W;

[0023] Perform a stacking operation, and the output obtained after passing through one layer of the DBB network is O. Refer to the following formula:

[0024] O = concatenate(B 1 , B 2 , B 3 , B 4 )

[0025] Among them, concatenate represents the stacking operation, and O ∈ R 4D×H′×W′ .

[0026] Furthermore, on the basis of the DBB module, in order to further extract more features and minimize the computational amount as much as possible, depthwise separable convolution is adopted as the second content of the temporal feature extractor;

[0027] Different from traditional convolution, depthwise separable convolution splits the standard convolution into two stages: depthwise convolution and pointwise convolution;

[0028] Depthwise convolution performs independent convolution operations on the input channels, and pointwise convolution uses 1x1 convolution to combine all channels to obtain the final output;

[0029] Since only one convolution kernel convolves one channel in depthwise convolution, the number of channels generated in this process is exactly the same as the number of input channels. Assuming the input is I ds , the depthwise convolution operation is f de ;

[0030] Since depthwise convolution does not effectively utilize the feature information of different channels at the same spatial position, pointwise convolution is needed to combine these feature maps into new feature maps;

[0031] Pointwise convolution is similar to traditional convolution operation, and its convolution kernel size is 1x1xC, where C is the number of channels in the previous layer. Therefore, pointwise convolution will perform weighted combination in the depth direction on the feature maps of depthwise convolution;

[0032] The specific operation of the depthwise separable convolution is as follows:

[0033] O ds = f pt (f de (I ds ))

[0034] Among them, f de is the depthwise convolution operation, and f pt is the pointwise convolution operation.

[0035] Furthermore, the SEBlock dynamically adjusts the weights of feature channels through a "squeeze-excitation" structure to enhance the response of key features;

[0036] Specifically, the SEBlock optimizes the weight allocation of features through the following process:

[0037] Squeeze: First, the input feature map is compressed through global average pooling to aggregate the global information of the channels, converting the features of each channel into a global statistic;

[0038] In this way, the network can capture the global context information of the entire feature map rather than simply relying on local information. Assuming the given input is The channel description vector after the compression operation is:

[0039]

[0040] where C is the number of channels, H is the height, W is the width, and X i,j,c represents the input feature Figure X at the pixel value corresponding to position (i, j) and channel k, and the vector z contains the global information of each channel;

[0041] Excitation: After the squeeze step, a fully connected layer (usually including ReLU activation and Sigmoid activation functions) is used to adaptively adjust the weights of each channel. In this way, the DBBCapsNet model can self-learn the importance of each channel, thereby adjusting the activation intensity of each feature channel, strengthening the feature representation of important channels, and suppressing redundant or unimportant channel information;

[0042] The output s of the excitation process is obtained through the Sigmoid activation function σ, which generates a weight coefficient s for each channel c , representing the importance of the channel, as shown in the following formula:

[0043] s = σ(W 2 ReLU(W 1 z))

[0044] where W 1 , W 2 are the dimensionality reduction matrix and the dimensionality increase matrix respectively, r is the compression ratio, and σ is the sigmoid activation function;

[0045] Scale: The learned weights are applied to the original input feature map in a pointwise multiplication manner, and the final output is The output feature Figure X 's each element is calculated through pointwise multiplication:

[0046] X i ' ,j,k =s c ·X i,j,k

[0047] In the scaling step, the learned weights s are multiplied channel by channel. c Apply to original feature Figure X In this way, the network can adaptively amplify or suppress the corresponding features according to the importance of each channel.

[0048] Furthermore, the capsule network includes a primary capsule layer and an emotion capsule layer;

[0049] The main capsule layer converts the feature map into an 8D capsule vector;

[0050] The emotion capsule layer iteratively generates high-level capsules through a dynamic routing mechanism and calculates the L2 norm output classification probability; in the dynamic routing algorithm, the capsule vector is normalized using the squash activation function, and the feature representation is optimized by iteratively updating the coupling coefficient;

[0051] Margin loss is used as the total loss function in the capsule network to optimize the feature extractor and classifier parameters.

[0052] Furthermore, the main capsule layer serves as the initial layer of the capsule network, and takes the output of the feature extraction layer as input, so that the original multi-level feature map is converted into an original capsule;

[0053] In the main capsule operation, the original 256 channel map is converted into 32 8D capsule vectors, that is, each capsule vector contains 8 units;

[0054] Furthermore, the emotion capsule layer is the core part of the capsule network, the input and output of the emotion capsule layer are both capsule vectors, the input of the emotion capsule layer accepts the vector input of the main capsule layer, and the output of the emotion capsule layer searches for the advanced capsule vector through the dynamic routing mechanism algorithm.

[0055] Furthermore, the dynamic routing is the core mechanism of the capsule network, which is used to dynamically allocate weights between capsules and determine how to transfer learning from low-level capsules to high-level capsules; finally, due to different tasks, the number of capsule vectors is set to 2 or 4, corresponding to HV / LV and HA / LA, respectively, where HV / LV stands for high / low valence and HA / LA stands for high / low arousal;

[0056] In the dynamic routing algorithm, the formula for normalizing the capsule vector using the squash activation function is as follows:

[0057]

[0058] Among them, s j represents the j-th capsule of the emotional capsule layer;

[0059] After being converted into a capsule vector, the squash activation function is used to normalize it to obtain the original capsule u i ; The original capsule is used as the input of the dynamic routing mechanism, and it needs to be converted into the capsule of the emotional capsule layer through the dynamic routing mechanism;

[0060] During the dynamic routing process:

[0061] The weight matrix W from the low-level capsule i to the high-level capsule j is used ij to learn the input features and represent them as higher-level emotional features, as shown in the following formula:

[0062]

[0063] Among them, is the prediction vector from the low-level capsule i and is used as the input of the high-level capsule j;

[0064] Calculate the input of each high-level capsule s j :

[0065]

[0066] Among them, s j represents the j-th capsule of the emotional capsule layer; c ij is the coupling coefficient;

[0067] The coupling coefficient c ij , is obtained by performing softmax on b ij , and it is stipulated that ∑ j c ij = 1:

[0068] c ij = softmax(b ij )

[0069] Among them, b ij is initialized to 0 to ensure that the coupling coefficients of all paths are the same initially. Then, according to the definition, the length of a capsule indicates the likelihood that it belongs to a specific emotional capsule. Therefore, the capsule length must be between 0 and 1;

[0070] The dynamic routing optimizes the feature representation by iteratively updating the coupling coefficient b ij :

[0071]

[0072] Among them, squash(S j ) is the normalized vector output by the high-level capsule j;

[0073] Through the above process, if and squash(S j ) have a high similarity, then b ij will be updated to a higher value, so as to update c ij . After the iteration is completed, it is determined that the output of the capsule neurons on the path is the capsule of the correct prediction.

[0074] Furthermore, the calculation formula of the margin loss is as follows:

[0075] L class = T k max(0, m + -‖v k ‖) 2 + λ(1 - T k )max(0, ‖v k ‖ - m - ) 2

[0076] The optimization goal is to minimize the total loss L class , so as to optimize the feature extractor and classifier parameters θ f and θ c :

[0077]

[0078] Among them, T k is the sentiment label. If there is a k-th type of sentiment, then T k = 1, otherwise T k = 0; m + and m _ are used to punish false positives and false negatives, and are set to 0.85 and 0.15 respectively; ‖v k ‖ is the L2 norm of the output capsule vector of the emotion capsule layer.

[0079] Furthermore, the application using the trained DBBCapsNet model at least includes the following steps:

[0080] The preprocessed EEG signal first passes through the DBB module. The DBB module extracts local and global features in parallel on different branches through various convolution and pooling operations, so as to effectively capture the diversity information in the signal;

[0081] Next, the extracted features are input into depthwise separable convolutions to further extract feature information in the time domain and enhance the temporal dependence of the signal;

[0082] After feature extraction, an SEBlock is introduced. Through a "squeeze-and-excitation" structure, it automatically adjusts the feature responses of each channel, strengthens important features, suppresses redundant features, and improves the effectiveness of feature representation;

[0083] Finally, the feature maps processed by the SEBlock are converted into capsule vectors and input into a capsule network. In the capsule network, through a dynamic routing mechanism, spatial information and multi-level feature representations are gradually integrated, and finally accurate classification results are generated.

[0084] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0085] 1. Improve recognition accuracy: By combining a parallel multi-branch structure (Diverse Branch Block, DBB) and depthwise separable convolutions, the present invention can more comprehensively extract the temporal and spatial features of EEG signals, thereby significantly improving the accuracy of emotion recognition.

[0086] 2. Enhance the generalization ability of the model: By introducing the SEBlock channel attention mechanism, the model can automatically adjust feature responses, strengthen the attention to important features, thereby improving the generalization ability of the model and enabling it to perform well in different datasets and tasks.

[0087] 3. Reduce computational complexity: The use of depthwise separable convolutions reduces the computational amount of the model, while improving the training speed and inference efficiency, making the application of the model on large-scale datasets more efficient.

[0088] 4. Wide applicability: The method of the present invention not only performs well on public datasets such as DEAP and DREAMER, but also has good scalability and can be applied to other EEG emotion recognition datasets, having wide applicability.

[0089] 5. End-to-end multi-scale feature extraction: The present invention proposes an end-to-end multi-scale emotion recognition model that can directly extract spatio-temporal features from the original signal for classification, avoiding information loss in the preprocessing stage. Through the parallel multi-branch DBB module, the model can extract features of different scales and more comprehensively capture the feature information of EEG signals.

[0090] 6. Comprehensive spatial feature extraction: By combining with a capsule network (Capsule Network), the present invention can comprehensively extract the spatial feature information of EEG, and use the dynamic routing mechanism to capture the hierarchical relationship between features layer by layer, effectively integrating local and global features, and further improving the accuracy of emotion recognition.

[0091] 7. Excellent performance: The experimental results on two public datasets, DEAP and DREAMER, show that the method of the present invention achieves average accuracies of 98.98%, 98.56% and 98.13% (DEAP dataset) and 94.92%, 95.31% and 93.12% (DREAMER dataset) respectively in valence, arousal and four-classification tasks, significantly outperforming existing methods and achieving the current state-of-the-art performance level. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0093] Figure 1 Schematic diagram of DBBCapsnet of the present invention;

[0094] Figure 2 Schematic diagram of Diverse Branch Block of the present invention;

[0095] Figure 3 Schematic diagram of depthwise separable convolution - attention mechanism of the present invention;

[0096] Figure 4 Schematic diagram of dynamic routing mechanism of the present invention;

[0097] Figure 5 Bar chart comparing DEAP dataset of the present invention with classical models;

[0098] Figure 6 Bar chart comparing DREAMER dataset of the present invention with classical models;

[0099] Figure 7 Bar chart comparing DEAP dataset of the present invention with the latest literature;

[0100] Figure 8 Bar chart comparing DREAMER dataset of the present invention with the latest literature;

[0101] Figure 9 Schematic diagram of confusion matrix of the present invention;

[0102] Figure 10 Schematic diagram of ablation performance comparison of each subject of the present invention;

[0103] Figure 11This is a visualization diagram of the output features at different stages of the present invention. Detailed implementation manners

[0104] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0105] The core of the present invention is to achieve efficient feature extraction and emotion classification of electroencephalogram (EEG) signals through a diverse branch block (DBB), depthwise separable convolution, SEBlock channel attention mechanism, and capsule network. The model architecture is as Figure 1 shown, and the detailed parameters of each layer structure of the network are shown in Table 1.

[0106] Table 1 Detailed parameters of the DBBCapsnet structure

[0107]

[0108] Among them, C represents the number of EEG channels, T represents the number of time samples, F O and F i represent the number of filters, F p represents the number of filters in the primary capsule layer, K 1 represents the size of the convolution kernel in the time feature extractor, K p represents the size of the convolution kernel in the primary capsule layer, Nc represents the number of capsules in the primary capsule layer, Dc represents the dimension size of each capsule, c represents the number of final output classifications, and lc represents the dimension of the final output capsule.

[0109] The time feature extractor includes three modules: the DBB module extracts multi-scale time features in parallel through four branches, uses different convolution kernels (1x1, 3x3, 6x6) and combines pooling operations to enhance feature diversity; depthwise separable convolution further extracts time features and reduces the computational amount; the SEBlock module dynamically adjusts the channel weights through a "squeeze-excitation" structure to strengthen key features. The capsule network converts the extracted features into capsule vectors, integrates local and global features through a dynamic routing mechanism, and outputs the emotion classification result.

[0110] Specifically as follows:

[0111] A multi-scale emotion recognition method based on the DBBCapsNe model includes at least the following steps:

[0112] S1: Select the DEAP and DREAMER public datasets;

[0113] S2: Perform preprocessing of the data, standardize the signals into a format suitable for model input. The preprocessing includes, but is not limited to, downsampling and band-pass filtering of the electroencephalogram (EEG) signals. In this example, the EEG signals are sampled to 128 Hz; band-pass filtering with 4 - 45 Hz is used; the signals are segmented using a 1-second sliding window, the baseline average value is calculated and z-score standardization is performed; different settings can be adopted in different embodiments.

[0114] S3: Construct the DBBCapsNet model. The DBBCapsNet model includes a time feature extractor and a capsule network. The time feature extractor includes a Diverse Branch Block (DBB), Depthwise Separable Convolution, and an SEBlock. The Diverse Branch Block is the DBB module, and the SEBlock is the channel attention mechanism.

[0115] S4: Use the 10-fold cross-validation method to train the DBBCapsNet model and evaluate its performance.

[0116] S5: Apply the trained DBBCapsNet model and then output the emotion classification result.

[0117] Refer to Figure 2 , the DBB module is composed of four parallel branches, which respectively use different convolutional kernels and pooling operations to extract multi-scale time features, and the training is stabilized through batch normalization. The specific structure of the DBB module is as follows:

[0118] The first branch: Use a 3x3 convolutional kernel for feature extraction.

[0119] The second branch: Use a 1x1 convolution to adjust the number of channels.

[0120] The third branch: First use a 1x1 convolution to reduce the number of channels, and then apply average pooling.

[0121] The fourth branch: First use a 1x1 convolution to reduce the number of channels, and then use a 6x6 convolutional kernel to extract a larger range of local features.

[0122] After each convolution and pooling operation, the training of the network is stabilized through a batch normalization layer, further improving the robustness and generalization ability of the model.

[0123] Among them, the outputs of each branch are merged through a concatenate operation to enhance the richness of feature expression. Furthermore, the DBB module extracts features from different scales and perspectives through a multi-branch structure, reducing the computational amount and preventing overfitting.

[0124] In the DBB module, the input is set as: the dimension of the input is \(I\in\mathbb{R}\) C×H×W , where \(C\) is the number of channels, \(H\) is the height, and \(W\) is the width;

[0125] The settings of the branch outputs are: \(B\) 1 , \(B\) 2 , \(B\) 3 and \(B\) 4 represent the output dimensions of each branch respectively, where \(B\) i \(\in\mathbb{R}\) D ×H′×W′ , where \(H'\) and \(W'\) are the new height and width obtained by processing \(H\) and \(W\);

[0126] Performing a stacking operation, the output obtained after passing through one layer of the DBB network is \(O\), as shown in the following formula:

[0127] \(O = \text{concatenate}(B\) 1 , \(B\) 2 , \(B\) 3 , \(B\) 4 )

[0128] where \(\text{concatenate}\) represents the stacking operation and \(O\in\mathbb{R}\) 4D×H′×W′ .

[0129] Based on the DBB module, in order to further extract more features and minimize the computational cost as much as possible, depthwise separable convolution is adopted as the second part of the temporal feature extractor;

[0130] Different from traditional convolution, depthwise separable convolution splits the standard convolution into two stages: depthwise convolution and pointwise convolution;

[0131] Depthwise convolution performs independent convolution operations on the input channels, and pointwise convolution uses 1x1 convolution to combine all channels to obtain the final output;

[0132] Since only one convolution kernel convolves one channel in depthwise convolution, the number of channels generated in this process is exactly the same as the number of input channels. Assuming the input is \(I\) ds , the depthwise convolution operation is \(f\) de ;

[0133] Since depthwise convolution does not effectively utilize the feature information of different channels at the same spatial position, pointwise convolution is needed to combine these feature maps into new feature maps;

[0134] Pointwise convolution is similar to traditional convolution operations. Its convolution kernel size is 1x1xC, where \(C\) is the number of channels in the previous layer. Therefore, pointwise convolution will perform weighted combination in the depth direction on the feature maps of depthwise convolution;

[0135] The specific operation of depthwise separable convolution is as follows:

[0136] O ds = f pt (f de (I ds ))

[0137] Among them, f de is the depthwise convolution operation, and f pt is the pointwise convolution operation.

[0138] Refer to Figure 3 , the SEBlock dynamically adjusts the feature channel weights through a "squeeze-excitation" structure to enhance the key feature responses;

[0139] Specifically, the SEBlock optimizes the weight allocation of features through the following process:

[0140] Squeeze: First, the input feature map is compressed through global average pooling to gather the global information of the channels, and the features of each channel are transformed into a global statistic;

[0141] In this way, the network can capture the global context information of the entire feature map instead of simply relying on local information. Assuming the given input is The channel description vector after the compression operation is:

[0142]

[0143] Among them, C is the number of channels, H is the height, W is the width, and X i,j,k represents the input feature Figure X at the pixel value corresponding to position (i, j) and channel k, and the vector z contains the global information of each channel;

[0144] Excitation: After the squeeze step, the weights of each channel are adaptively adjusted using a fully connected layer (usually including ReLU activation and Sigmoid activation functions). In this way, the DBBCapsNet model can self-learn the importance of each channel, thereby adjusting the activation intensity of each feature channel, strengthening the feature representation of important channels, and suppressing redundant or unimportant channel information;

[0145] The output s of the excitation process is obtained through the Sigmoid activation function σ, which generates a weight coefficient s c for each channel, representing the importance of the channel. Refer to the following formula:

[0146] s = σ(W 2 ReLU(W 1 z))

[0147] Among them W 1 ,W 2 are the dimensionality reduction matrix and the dimensionality increase matrix respectively, r is the compression ratio, and σ is the sigmoid activation function;

[0148] Scaling: The learned weights are applied to the original input feature map in a pointwise multiplication manner, and the final output is output feature Figure X Each element of ' is calculated by pointwise multiplication:

[0149] X i ′ ,j,k = s c ·X i,j,k

[0150] In the scaling step, the learned weights s c are applied to the original features Figure X to adjust the feature representation of each channel. In this way, the network can adaptively amplify or suppress the corresponding features according to the importance of each channel;

[0151] By introducing the SEBlock, the DBBCapsNe model can not only automatically focus on and strengthen the key feature channels, but also dynamically adjust the feature expression, making the network more efficient and accurate in complex EEG emotion recognition tasks. The introduction of the SEBlock further improves the sensitivity of the network to the emotion information in EEG data and enhances its recognition ability in different emotional states. Through the SEBlock, the DBBCapsNe model can automatically focus on and strengthen the key feature channels, dynamically adjust the feature expression, and improve the efficiency and accuracy of the model in complex EEG emotion recognition tasks.

[0152] The capsule network includes a Primary Capsule Layer and an Emotion Capsule Layer;

[0153] The Primary Capsule Layer converts the feature map into 8D capsule vectors;

[0154] The Emotion Capsule Layer iteratively generates high-level capsules through a dynamic routing mechanism and calculates the L2 norm to output classification probabilities; In the dynamic routing algorithm, the squash activation function is used to normalize the capsule vectors, and the feature representation is optimized by iteratively updating the coupling coefficients;

[0155] The margin loss is used as the total loss function in the capsule network to optimize the parameters of the feature extractor and the classifier.

[0156] The main capsule layer is the initial layer of the capsule network, which takes the output of the feature extraction layer as input, so that the original multi-level feature map is converted into the original capsule;

[0157] In the main capsule operation, the original 256 channel map is converted into 32 8D capsule vectors, that is, each capsule vector contains 8 units;

[0158] Taking the DEAP dataset as an example, a convolution operation is performed with a 9x9 convolution kernel, a step size of 2, and an output feature map of shape (32, 8, 256); these feature maps are then split into sub-feature maps with 32 channels, so that the main capsule layer finally obtains a total of 8192 initial capsules.

[0159] The emotion capsule layer is the core part of the capsule network. The input and output of the emotion capsule layer are both capsule vectors. The input of the emotion capsule layer accepts the vector input of the main capsule layer, and the output of the emotion capsule layer searches for the advanced capsule vector through the dynamic routing mechanism algorithm.

[0160] See also Figure 4 Dynamic routing is the core mechanism of the capsule network, which is used to dynamically allocate weights between capsules and determine how to transfer learning from low-level capsules to high-level capsules; finally, due to different tasks, the number of capsule vectors is set to 2 or 4, corresponding to HV / LV and HA / LA, respectively. HV / LV means high / low valence, and HA / LA means high / low arousal.

[0161] This process is the classification process. The L2 norm of the output capsule vector of the emotion capsule layer is calculated to obtain the probability of each emotion category. The emotion recognition result is output according to the category with the highest probability, such as high / low valence (HV / LV) or high / low arousal (HA / LA).

[0162] In the dynamic routing algorithm, the formula for normalizing the capsule vector using the squash activation function is as follows:

[0163]

[0164] Among them, s j Represents the jth capsule of the sentiment capsule layer;

[0165] After conversion to a capsule vector, the squash activation function is used to normalize it to get the original capsule u i ; The original capsule is used as the input of the dynamic routing mechanism, which needs to be converted into capsules of the emotion capsule layer through the dynamic routing mechanism;

[0166] During dynamic routing:

[0167] Use the weight matrix W from low-level capsule i to high-level capsule j ijThe features of the learning input are represented as higher-level sentiment features, as shown in the following formula:

[0168]

[0169] where is the prediction vector from the lower-level capsule i and is used as the input for the higher-level capsule j;

[0170] Calculate the input for each higher-level capsule s j :

[0171]

[0172] where s j represents the j-th capsule in the sentiment capsule layer; c ij is the coupling coefficient;

[0173] The coupling coefficient c ij , is obtained by performing softmax on b ij , and it is stipulated that ∑ j c ij = 1:

[0174] c ij = softmax(b ij )

[0175] where the initial value of b ij is 0 to ensure that the coupling coefficients of all paths are the same initially. Then, by definition, the length of a capsule indicates the likelihood that it belongs to a specific sentiment capsule. Therefore, the capsule length must be between 0 and 1;

[0176] Dynamic routing optimizes the feature representation by iteratively updating the coupling coefficient b ij :

[0177]

[0178] where squash(S j ) is the normalized vector output by the higher-level capsule j;

[0179] Through the above process, if and squash(S j ) are highly similar, then b ij will be updated to a higher value, thereby updating c ij . After the iteration is completed, it is determined that the output of the capsule neurons on the path is the capsule with the correct prediction.

[0180] The formula for the margin loss is as follows:

[0181] L class = Tk max(0, m + - ‖v k ‖) 2 + λ(1 - T k )max(0, ‖v k ‖ - m - ) 2

[0182] The optimization objective is to minimize the total loss L class , so as to optimize the parameters θ f of the feature extractor and the classifier c :

[0183]

[0184] where T k is the sentiment label. If there is a sentiment of the k-th class, then T k = 1; otherwise, T k = 0; m + and m_ are used to penalize false positives and false negatives, and are set to 0.85 and 0.15 respectively; ‖v k ‖ is the L2 norm of the output capsule vector of the emotion capsule layer.

[0185] Applying the trained DBBCapsNet model includes at least the following steps:

[0186] The preprocessed EEG signals first pass through the DBB module. The DBB module extracts local and global features in parallel on different branches through various convolutional and pooling operations, so as to effectively capture the diversity information in the signals;

[0187] Next, the extracted features are input into the depthwise separable convolution to further extract the feature information in the time domain and enhance the time dependence of the signals;

[0188] After feature extraction, the SEBlock is introduced. Through the "squeeze - excitation" structure, it automatically adjusts the feature responses of each channel, strengthens the important features, suppresses the redundant features, and improves the effectiveness of feature representation;

[0189] Finally, the feature map processed by the SEBlock is converted into capsule vectors and input into the capsule network. In the capsule network, through the dynamic routing mechanism, the spatial information and multi - level feature representations are gradually integrated, and finally an accurate classification result is generated. This process effectively utilizes the complementarity of different - level features in deep learning and enhances the model's perception ability for the sentiment classification task.

[0190] Based on the two publicly available datasets DEAP and DREAMER used in this case, the following verification is carried out;

[0191] Experimental verification:

[0192] The present invention has been experimentally verified on two publicly available datasets, DEAP and DREAMER. The experimental results show that on the DEAP dataset, the average accuracies of valence, arousal, and four-classification tasks are 98.98%, 98.56%, and 98.13% respectively. On the DREAMER dataset, the average accuracies of valence, arousal, and four-classification tasks are 94.92%, 95.31%, and 93.12% respectively. Compared with the prior art, the present invention significantly improves the accuracy of emotion recognition, demonstrating its superiority.

[0193] Experimental setup:

[0194] This study was implemented using the open-source software library TensorFlow on a workstation equipped with an NVIDIA RTX 3080Ti GPU, and experimentally verified using two publicly available electroencephalogram (EEG) emotion datasets, DEAP and DREAMER. The DEAP dataset contains EEG data of 32 subjects with a sampling rate of 128 Hz. Each subject has 40 emotional segments, each segment is 63 seconds long, including 3 seconds of baseline time. The DREAMER dataset contains EEG data of 23 subjects with a sampling rate of 128 Hz. Each subject has 18 movie stimulus segments, and the segment duration ranges from 65 to 393 seconds.

[0195] The DEAP dataset and the DREAMER dataset are divided according to trials, channels, and data dimensions as follows:

[0196] Table 2 DEAP and DREAMER datasets

[0197]

[0198] In the data preprocessing stage, the data is segmented using a 1-second sliding window, and the baseline signal fluctuations and noise interferences are eliminated by subtracting the baseline average value. The specific operation is as follows:

[0199] Calculate the baseline average value:

[0200]

[0201] where n is the number of seconds of the baseline duration. Then the segmented signals are respectively subtracted from the average baseline signal:

[0202]

[0203] To avoid excessive differences in the magnitudes of feature values and eliminate the distribution differences between different channels, the calibrated data is standardized using z-score:

[0204]

[0205] Among them, X'mean represents the current sample mean, and X'sigma represents the current sample standard deviation. The final data is obtained through the above operations.

[0206] In the experimental setup, the DEAP dataset is classified into high and low valence and arousal categories with a threshold of 5, and the DREAMER dataset is classified with a threshold of 3. The data of each subject is trained and tested using 10-fold cross-validation. During model training, the ADAM optimizer is used to minimize the loss function. The initial learning rate for the DEAP dataset is set to 1×10 -4 , and the number of training epochs is 50; the initial learning rate for the DREAMER dataset is set to 2×10 -5 , and the number of training epochs is 60. The batch size is uniformly set to 50.

[0207] Experimental results:

[0208] (1) Compared with classical algorithms

[0209] The DBBCapsnet will be used to verify and compare the DEAP and DREAMER datasets with classical methods. These include DT, MLP, 3DCNN, DGCNN, Capsnet, EESNN. Both the DT and MLP models conduct experiments by obtaining DE features. Extracting DE features from four-band signals is more helpful for improving the accuracy of emotion recognition. Similarly, 3DCNN constructs 3D inputs by extracting DE features and combining electrode spatial positions. DGCNN extracts signal features from multiple frequency bands and can better reflect the functional relationships between channels through the dynamic adjacency matrix training for adaptive update. EESNN is based on SNN and CNN, converts EEG signals into 2D frame sequences, and uses LIF nodes to better simulate the biological neuron mechanism and better process frequency domain, time, and spatial information. ACRNN is an attention-based convolutional recurrent neural network. Aiming to mine the temporal information of EEG signals, extended self-attention is introduced into the RNN, and the importance is encoded using the intrinsic similarity of EEG signals. The emergence of Capsnet provides a new direction for extracting spatial features, encoding the spatial information of EEG signals using the dynamic routing method. To verify the effectiveness of the DBBCapsnet proposed in the present invention, the model proposed in the present invention is compared with the above-mentioned 7 most classical methods.

[0210] Table 3 shows the average accuracy and standard deviation of 32 subjects in the valence, arousal, and four-classification tasks in the DEAP dataset. The results show that compared with the other seven methods, the method of the present invention performs best in both the valence and arousal dimensions, reaching accuracies of 98.98% and 98.56% respectively. At the same time, the present invention also uses multi-dimensional combination. By integrating the two dimensions of valence-arousal, compared with single-dimensional recognition, it can provide a new perspective for emotion modeling, so as to explore more features and laws of EEG signals, enabling the model to be applicable to more complex interaction scenarios. In the present invention, the two-dimensional four-classification model reaches an accuracy of 98.13%. Compared with the traditional models DT and MLP, the model proposed by the present invention has an average increase of 19.3% and 17.14% respectively in the accuracy of the two dimensions. Compared with 3DCNN and DGCNN based on spatio-temporal features, the method of the present invention has an average increase of 7.98% and 6.47% respectively in the valence and arousal dimensions. Compared with the neuromorphic-based model EESNN, the model of the present invention has an accuracy increase of 4.42% and 3.25%. Compared with Capsnet, the method of the present invention achieves higher accuracy in the two-dimensional classification task, with an average increase of 2.62% and 2.82%, which also shows that the model of the present invention has stronger ability to capture emotion features when processing EEG data. Please refer to Figure 5 。

[0211] Table 3 Comparison with classical models in the DEAP dataset

[0212]

[0213]

[0214] When further verifying the generality and generalization ability of the method of the present invention, the present invention selected another widely used EEG dataset DREAMER for experimental analysis. On the DREAMER dataset, the method of the present invention also demonstrated excellent performance. Table 4 shows the two-dimensional average accuracy and standard deviation of the method of the present invention and the other five methods on 23 subjects in the DREAMER dataset. The method of the present invention achieved accuracies of 94.92% and 95.31% in the two dimensions of valence and arousal. Compared with the traditional DT and MLP, the method of the present invention increased by 15.24% and 15.59% on average in the two dimensions. Similarly, compared with the 3DCNN and DGCNN based on spatio-temporal features, the method of the present invention increased by 7.86% and 8.43% on average in the valence and arousal dimensions respectively. Compared with Capsnet, the method of the present invention achieved higher accuracies in the two-dimensional classification task, with an average increase of 1.21% and 1.28%, and can be referred to Figure 6 。

[0215] Table 4 Comparison of the DREAMER dataset with other models

[0216]

[0217] (2) Comparison with the latest literature

[0218] To comprehensively evaluate the performance of the model of the present invention, the present invention systematically compared it with the latest models proposed in recent years. Through the comparison, the present invention not only verified the superiority of the model, but also explored its applicability in different tasks and datasets. As shown in the following table, the input formats and final effects of the model proposed by the present invention and each model are shown. It can be divided into three categories according to different input features. One is to use DE features as the input. Compared with GLFANet, the accuracy of the method of the present invention increased by 4.45% in valence and 3.92% in arousal. For the GANSER model with a 2D spatial matrix as the input, the method of the present invention also increased by 5.12% and 4.05% in the two dimensions. For the currently published capsule network-based models, most of them use the original signal as the input of the model. Compared with them, the average recognition rates of the model of the present invention in the two classification tasks increased by 0.9% and 0.49% respectively.

[0219] Based on the DREAMER dataset, the present invention also conducts a comprehensive comparison. Different from the DE features or spatial matrix representations used by GLFANet and GANSER, the method of the present invention is directly based on the original signal, effectively avoiding information loss that may be introduced during the preprocessing process and fully exploiting the original characteristics of EEG signals. Compared with CRAM (93.03%, 92.27%) and MLF-CapsNet (94.59%, 95.26%) which are also based on the original signal, the model of the present invention has further improvements in both tasks, especially showing stronger emotion recognition ability in the Arousal dimension. The method of the present invention is 3.54% higher than MLF-CapsNet in the four-classification task, which also fully demonstrates the superiority of jointly modeling the Valence and Arousal dimensions. Compared with the latest ICaps-ResLSTM (94.71%, 94.97%), the model of the present invention has improved by 0.21% and 0.34% respectively in Valence and Arousal classification, showing higher robustness and generalization ability.

[0220] Table 5 Comparison of performance between DEAP dataset and the latest literature

[0221]

[0222] Table 6 Comparison of performance between DREAMER dataset and the latest literature

[0223]

[0224] Refer to Figure 7 and Figure 8 , the present invention can intuitively understand the recognition performance in each category. The first two in each row are the confusion matrices of single binary classification for valence and Arousal, and the last is the confusion matrix of four-classification combination of the two dimensions. In Figure 9 , the rows represent the true class labels and the columns represent the predicted class labels. The darker the color of the diagonal, the better the classification performance. In the DEAP dataset, the diagonal matrix blocks are dark and have similar colors, indicating that the model shows excellent performance in both binary classification tasks and four-classification tasks. While in the DREAMER dataset, the color in the lower right of the binary classification task for valence and the four-classification task is slightly lighter than that in the upper left, which indicates that the model is more difficult to recognize HV (high valence) than LV for the DREAMER dataset, and there may be missed detections in this category, especially 13.83% of the actual positive examples are missed in the HVLA category.

[0225] Ablation experiment

[0226] To further verify the effectiveness of each component or improvement point in the model, the present invention conducted ablation experiments. By gradually removing key modules or specific designs in the model, the present invention analyzed the contribution of each component to the overall performance. This method can help the present invention quantify the actual impact of each part and evaluate the rationality of the model structure design. The present invention first conducted experiments on the DEAP dataset, as shown in the following table.

[0227] Table 7 DEAP Ablation Experiment

[0228]

[0229] Table 8 DREAMER Ablation Experiment

[0230]

[0231] As shown in Table 8, the model proposed by the present invention achieved the highest accuracy rate in the classification task. To verify the effectiveness of the DBB module, the present invention replaced the original module with a common convolution having the same number of parameters instead of directly removing it. The experimental results show that the DBB module with parallel multi-branches can extract feature information more efficiently, thus significantly improving the accuracy rate of emotion recognition. Especially in the experiment on the DREAMER dataset, the ablation experiment results of the DBB module show that this module improved the accuracy rate by 4.13% and 3.06% respectively in two dimensions. In addition, CapsNet, another key module of the model of the present invention, also significantly enhanced the classification performance of the model by virtue of its strong spatial feature extraction ability. Refer to Figure 10 , Figure 10 a and Figure 10 b are for the DEAP dataset, Figure 10 c and Figure 10 d are for the DREAMER dataset.

[0232] Feature Visualization

[0233] To more intuitively display the performance of the overall model, the present invention uses the t-SNE dimensionality reduction technique to visualize the initial data, the feature map after passing through the DBB module, and the feature map of the final capsule network layer. As Figure 11 shown, Figure 11 a is for the fifth subject in the DEAP dataset, Figure 11 b is for the twelfth subject in the DREAMER dataset;

[0234] The data of the fifth subject in the DEAP dataset and the twelfth subject in the DREAMER dataset were selected for illustration. The display of the initial feature map after dimensionality reduction still appears rather chaotic, and the classification results do not have obvious separations. However, for the feature map processed by the DBB module, a certain degree of inter-class aggregation can already be observed, indicating that the features are distinguishable among certain categories. Finally, the feature map obtained through the capsule network shows a more obvious classification effect, and the boundaries between the features are clearer, verifying the effectiveness of the model in feature learning and classification tasks.

[0235] Summary based on the above content:

[0236] The present invention combines a Diverse Branch Block (DBB), Depthwise Separable Convolution, SEBlock channel attention mechanism, and Capsule Network to achieve efficient feature extraction and emotion classification of electroencephalogram (EEG) signals. This method uses an end-to-end model architecture (DbbCapsnet), and the entire network is divided into a temporal feature extractor and a capsule network. In the temporal feature extractor module, the Diverse Branch Block and Depthwise Separable Convolution are used to effectively extract temporal domain features, and the SEBlock is integrated to enhance the weight connection between channels, further improving the feature expression ability. The extracted feature information is input into the capsule network to extract spatial domain features and effectively integrate local and global relationships. This study adopted an in-subject cross-validation training strategy to evaluate the effectiveness of the classification results. Experimental results on two public datasets, DEAP and DREAMER, show that the average accuracies of the present invention in valence, arousal, and four-classification tasks are 98.98%, 98.56%, 98.13%; 94.92%, 95.31%, 93.12% respectively, significantly superior to the prior art. Ablation experiments and feature visualization were used to verify each core module, and the experimental results show that the method of the present invention can effectively extract spatio-temporal feature information and demonstrate significant superiority in emotion recognition tasks.

[0237] The ablation experiments further verified the key roles of each module. Removing any one module (such as DBB, capsule network, or SEBlock) would lead to a performance decline, emphasizing their synergistic effects in capturing the complex spatio-temporal features of EEG signals. These modules are not only independently effective but also significantly improve the overall performance of the model through mutual cooperation.

[0238] Future research will focus on cross-subject emotion recognition tasks, exploring ways to improve the generalization ability of the model across different subjects in order to build a more general EEG emotion recognition solution. Meanwhile, plans are underway to optimize the model architecture to enhance its adaptability and interpretability for complex emotion patterns, providing new ideas and methods for further research in the field of EEG emotion recognition.

[0239] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Accordingly, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims concerned.

Claims

1. A multi-scale emotion recognition method based on DBBCapsNe model, characterized by: At least the following steps are included: S1: Select DEAP and DREAMER public datasets; S2: performing data preprocessing to standardize the signal into a format suitable for model input, wherein the preprocessing includes but is not limited to downsampling and bandpass filtering of the EEG signal; S3: Construct a DBBCapsNet model, wherein the DBBCapsNet model includes a temporal feature extractor and a capsule network, wherein the temporal feature extractor includes a diversified branch module, a depthwise separable convolution, and a SEBlock, wherein the diversified branch module is a DBB module, and the SEBlock is a channel attention mechanism; S4: Use 10-fold cross validation to train the DBBCapsNet model and evaluate the performance; S5: Use the trained DBBCapsNet model for application and output the emotion classification results.

2. The multi-scale emotion recognition method based on the DBBCapsNe model according to claim 1, characterized in that: The DBB module consists of four parallel branches, which use different convolution kernels and pooling operations to extract multi-scale time features and stabilize the training through batch normalization. The specific structure of the DBB module is as follows: The first branch: use 3x3 convolution kernel for feature extraction; The second branch: Use 1x1 convolution to adjust the number of channels; The third branch: first use 1x1 convolution to reduce the number of channels, and then apply average pooling; The fourth branch: first use 1x1 convolution to reduce the number of channels, and then use 6x6 convolution kernel to extract local features in a larger range; After each convolution and pooling operation, a batch normalization layer is used to stabilize the network training, further improving the robustness and generalization ability of the model; The output of each branch is combined through a stacking operation to enhance the richness of feature expression. The DBB module then extracts features from different scales and perspectives through a multi-branch structure to reduce the amount of calculation and prevent overfitting. In the DBB module, the input is set as follows: the input dimension is I∈R C×H×W , where C is the number of channels, H is the height, and W is the width; The branch output is set as follows: B1, B2, B3 and B4 represent the output dimensions of each branch, where B i ∈R D×H′×W′ , where H′ and W′ are the new height and width obtained after processing H and W; After stacking, the output obtained after one layer of DBB network is O, see the following formula: O = concatenate(B1,B2,B3,B4) Among them, concatenate represents the stacking operation, O∈R 4D×H′×W′ .

3. The multi-scale emotion recognition method based on the DBBCapsNe model according to claim 2 is characterized in that: Based on the DBB module, in order to further extract more features and minimize the amount of calculation, the depth-wise separable convolution is used as the second content of the temporal feature extractor; Depthwise separable convolution splits the standard convolution into two stages: depthwise convolution and pointwise convolution; The depth convolution performs independent convolution operations on the input channels, and the point-by-point convolution uses 1x1 convolution to combine all channels to obtain the final output; Since a channel is convolved by only one convolution kernel, the number of channels generated by this process is exactly the same as the number of input channels. Assuming the input is I ds , the depthwise convolution operation is f de ; Since deep convolution does not effectively utilize the feature information of different channels at the same spatial position, point-by-point convolution is required to combine these feature maps into a new feature map; Point-by-point convolution is similar to traditional convolution operation. Its convolution kernel size is 1x1xC, where C is the number of channels in the previous layer. Therefore, point-by-point convolution performs a weighted combination in the depth direction on the feature map of the depth convolution. The specific operation of the depthwise separable convolution is as follows: O ds =f pt (f de (I ds )) Among them, f de is the depthwise convolution operation, f pt It is a point-by-point convolution operation.

4. The multi-scale emotion recognition method based on the DBBCapsNe model according to claim 3 is characterized in that: The SEBlock dynamically adjusts the feature channel weights through a "compression-excitation" structure to enhance key feature responses; Specifically, SEBlock optimizes the weight distribution of features through the following process: Compression: First, the input feature map is compressed through global average pooling to aggregate the global information of the channel and convert the features of each channel into a global statistic; In this way, the network can capture the global context information of the entire feature map instead of relying solely on local information. Assuming that the given input is Channel description vector after compression operation for: Where C is the number of channels, H is the height, W is the width, and X is the width. i,j,k Represents the pixel value corresponding to position (i, j) and channel k in the input feature map X. The vector z contains the global information of each channel. Motivation: After the compression step, the weights of each channel are adaptively adjusted using the fully connected layer. In this way, the DBBCapsNet model can self-learn the importance of each channel, thereby adjusting the activation strength of each feature channel, strengthening the feature representation of important channels, and suppressing redundant or unimportant channel information; The output s of the excitation process is obtained through the Sigmoid activation function σ, which generates a weight coefficient s for each channel c , indicating the importance of the channel, see the following formula: s = σ(W2ReLU(W1z)) in W1, W2 are the dimension reduction matrix and dimension increase matrix respectively, r is the compression ratio, σ is the sigmoid activation function; Scaling: The learned weights are applied to the original input feature map by point-by-point multiplication, and the final output is Each element of the output feature map X′ is calculated by point-wise multiplication: X′ i,j,k =s c ·X i,j,k In the scaling step, the learned weights s are multiplied channel by channel. c Applied to the original feature map X, the feature representation of each channel is adjusted so that the network can adaptively amplify or suppress the corresponding features according to the importance of each channel.

5. The multi-scale emotion recognition method based on DBBCapsNe model according to claim 1, characterized in that: The capsule network includes a main capsule layer and an emotion capsule layer; The main capsule layer converts the feature map into an 8D capsule vector; The emotion capsule layer iteratively generates high-level capsules through a dynamic routing mechanism and calculates the L2 norm output classification probability; in the dynamic routing algorithm, the capsule vector is normalized using the squash activation function, and the feature representation is optimized by iteratively updating the coupling coefficient; Margin loss is used as the total loss function in the capsule network to optimize the feature extractor and classifier parameters.

6. The multi-scale emotion recognition method based on DBBCapsNe model according to claim 5, characterized in that: The main capsule layer serves as the initial layer of the capsule network, and takes the output of the feature extraction layer as input, so that the original multi-level feature map is converted into an original capsule; In the main capsule operation, the original 256 channel map is converted into 32 8D capsule vectors, that is, each capsule vector contains 8 units.

7. The multi-scale emotion recognition method based on DBBCapsNe model according to claim 6, characterized in that: The emotion capsule layer is the core part of the capsule network. The input and output of the emotion capsule layer are both capsule vectors. The input of the emotion capsule layer accepts the vector input of the main capsule layer, and the output of the emotion capsule layer searches for the advanced capsule vector through the dynamic routing mechanism algorithm.

8. The multi-scale emotion recognition method based on DBBCapsNe model according to claim 7, characterized in that: The dynamic routing is the core mechanism of the capsule network, which is used to dynamically allocate weights between capsules and determine how to transfer learning from low-level capsules to high-level capsules. Finally, due to different tasks, the number of capsule vectors is set to 2 or 4, corresponding to HV / LV and HA / LA, respectively. HV / LV stands for high / low valence, and HA / LA stands for high / low arousal. In the dynamic routing algorithm, the formula for normalizing the capsule vector using the squash activation function is as follows: Among them, s j Represents the jth capsule of the sentiment capsule layer; After conversion to a capsule vector, the squash activation function is used to normalize it to get the original capsule u i ; The original capsule is used as the input of the dynamic routing mechanism, which needs to be converted into capsules of the emotion capsule layer through the dynamic routing mechanism; During dynamic routing: Use the weight matrix W from low-level capsule i to high-level capsule j ij Learn the input features to represent them as higher-level sentiment features, see the following formula: in, is the prediction vector from the lower-level capsule i, used as the input to the higher-level capsule j; Calculate each advanced capsule j Input: Among them, s j represents the jth capsule of the emotion capsule layer; c ij is the coupling coefficient; Coupling coefficient c ij , through b ij Perform softmax to obtain, and stipulate ∑ j c ij =1: c ij =softmax(b ij ) Among them, b ij is initialized to 0 to ensure that the coupling coefficients of all paths are the same at the beginning. Then, by definition, the length of a capsule indicates the probability that it belongs to a specific emotion capsule, so the capsule length must be between 0 and 1; Dynamic routing iteratively updates the coupling coefficient b ij To optimize the feature representation: Among them, squash (S j ) is the normalized vector output by high-level capsule j; Through the above process, if and squash(S j ) have a high similarity, then b ij will be updated to a higher value, thereby updating c ij , after the iteration is completed, it is determined that the output of the capsule neuron on the path is the correctly predicted capsule.

9. The multi-scale emotion recognition method based on DBBCapsNe model according to claim 8, characterized in that: The marginal loss is calculated as follows: L class =T k max(0,m + -‖v k ‖) 2 +λ(1-T k )max(0,‖v k ‖-m - ) 2 The optimization goal is to minimize the total loss L class , thereby optimizing the feature extractor and classifier parameters θ f and θ c : Among them, T k is the emotion label. If there is a k-th emotion, then T k =1, otherwise T k =0;m + and m _ Used to penalize false positives and false negatives, set them to 0.85 and 0.15 respectively; ‖v k ‖ is the L2 norm of the output capsule vector of the emotion capsule layer.

10. The multi-scale emotion recognition method based on DBBCapsNe model according to claim 8, characterized in that: The application of the trained DBBCapsNet model includes at least the following steps: The preprocessed EEG signal first passes through the DBB module, which extracts local and global features in parallel on different branches through various convolution and pooling operations, thereby effectively capturing the diverse information in the signal; Next, the extracted features are input into the depthwise separable convolution to further extract feature information in the time domain and enhance the temporal dependency of the signal; After feature extraction, SEBlock is introduced to automatically adjust the feature response of each channel through the "compression-excitation" structure, strengthen important features, suppress redundant features, and improve the effectiveness of feature representation; Finally, the feature map processed by SEBlock is converted into a capsule vector and input into the capsule network. In the capsule network, the spatial information and multi-level feature representation are gradually integrated through a dynamic routing mechanism to finally generate accurate classification results.

Citation Information

Cited By

  • Deep learning-based pet dog emotion recognition method and system

    CN120708251A

  • Emergency data fusion and decision-making method under cross-modal dynamic routing mechanism

    CN120995384A

  • Method and system for predicting urban park visitor flow based on optimized CNN-LSTM model

    CN121168766A

  • Student mental health assessment method fusing multi-source biological behavior data

    CN121528541A