Multimodal emotion recognition method based on DLKA and EEGNet

Through the DLKA-EEGNet model combined with the time-frequency module and the variable large-core attention mechanism, the multimodal data fusion problem is solved, and the accuracy and robustness of emotion recognition are improved. Especially when fusing EEG and EOG signals, the computational complexity is reduced and the emotion recognition ability is enhanced.

CN120323974BActive Publication Date: 2025-08-22HUNAN UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510820278.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-08-22
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate multimodal data while ensuring the robustness of the model and deeply explore the original features of the EEG signal, affecting the accuracy of emotion recognition.

Method used

The DLKA-EEGNet model is adopted to extract the EEG signal characteristics through the time-frequency module, the space module and the variable large nuclear attention mechanism module, and multi-modal characteristics are fused with the EOG signal. The variable large nuclear attention mechanism module is used to adjust the receptive field adaptively, reduce the parameter amount, and enhance the global information mining of the deep convolutional network.

Benefits of technology

It improves the robustness and accuracy of the emotion recognition model, effectively integrates EEG and EOG feature data, reduces the computational complexity, makes full use of signal complementary information, and improves emotion recognition ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120323974B_ABST
    Figure CN120323974B_ABST
Patent Text Reader

Abstract

This invention discloses a multimodal emotion recognition method based on DLKA and EEGNet, comprising the following steps: extracting and preprocessing EEG signals; constructing a DLKA-EEGNet model and extracting EEG features using the DLKA-EEGNet model; extracting features from EOG signals and fusing them with EEG signals for multimodal feature fusion; and inputting the fused features into a deep learning model for emotion recognition. By introducing a model combining DLKA and EEGNet, this invention proposes a novel multimodal emotion recognition method. The DLKA module employs a transformable large-kernel convolution strategy, which can adaptively adjust the receptive field while reducing the number of parameters and computational complexity. The DLKA module also enhances the deep convolutional network's global information mining of EEG signals, thereby improving the robustness and accuracy of the emotion recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of emotion recognition, and in particular to a multimodal emotion recognition method based on DLKA and EEGNet. Background Art

[0002] Multimodal emotion recognition holds broad application prospects in areas such as intelligent healthcare, intelligent human-computer interaction, and mental health monitoring. Traditional emotion recognition methods typically rely on single-modal data, such as facial expressions, speech, or biosignals (e.g., electroencephalogram (EEG)). However, effectively integrating multimodal data while ensuring model robustness and deeply mining the raw features of EEG signals remains a key challenge in the field. EEG signals are inherently high-dimensional and complex, and effectively extracting emotional information from them has a significant impact on the accuracy of emotion recognition. Summary of the Invention

[0003] In order to solve the above technical problems, the present invention provides a multimodal emotion recognition method based on DLKA and EEGNet with simple algorithm and high recognition accuracy.

[0004] The technical solution of the present invention to solve the above technical problems is: a multimodal emotion recognition method based on DLKA and EEGNet, comprising the following steps:

[0005] S1: Extract features and preprocess EEG signals;

[0006] S2: Build the DLKA-EEGNet model and use it to extract EEG features;

[0007] S3: Extract features from the electrooculogram signal, i.e., the EOG signal, and perform multimodal feature fusion of the EOG signal and the EEG signal;

[0008] S4: Input the fused features into the deep learning model for emotion recognition.

[0009] In the multimodal emotion recognition method based on DLKA and EEGNet, in step S1, the EEG signal is recorded and acquired by 62 electrode channels. The subjects watch 24 emotion-triggering videos at different experimental stages. The length of each emotion-triggering video is different. While maintaining a sampling rate of 200 Hz, a 1-75 Hz bandpass filter is used to eliminate noise. The EEG data size of each emotion-triggering video ranges from 62×9601 to 62×51801.

[0010] In the above-mentioned multimodal emotion recognition method based on DLKA and EEGNet, in step S2, the DLKA-EEGNet model includes a time-frequency module, a spatial module and a variable large-core attention mechanism module, the variable large-core attention mechanism module is the DLKA module, the EEG signal is input into the time-frequency module to obtain feature F1, feature F1 is convolutionally processed by the spatial module to obtain feature F2, and feature F2 is input into the DLKA module for deep feature mining to obtain feature F3.

[0011] In the multimodal emotion recognition method based on DLKA and EEGNet, in step S2, the time-frequency module is used to mine the time-frequency information in the EEG signal. The model learns the frequency filtering features by setting a time convolution operation and uses a one-dimensional convolution kernel for all EEG channels, thereby learning the time sampling characteristics of each channel. The feature generated by the time-frequency convolution process is feature F1.

[0012] In the multimodal emotion recognition method based on DLKA and EEGNet, in step S2, the spatial module extracts the global spatial features of the EEG signal through deep convolution and separable convolution operations, and obtains feature F2 after convolution processing of feature F1.

[0013] In the multimodal emotion recognition method based on DLKA and EEGNet, in step S2, the DLKA module uses two variable convolution kernels: Deform-3 Conv2D and Deform-5 Conv2D, corresponding to convolution kernel sizes of 3×3 and 5×5, respectively. A large convolution kernel is constructed by using depthwise convolution, depthwise dilated convolution and 1×1 point convolution; a channel number of , the input dimension is , the kernel size is The depth convolution layer and the depth expansion convolution layer, represents the height of the input feature map, Represents the convolution kernel size. The kernel functions of the depth convolution layer and the depth expansion convolution layer are as follows:

[0014]

[0015]

[0016] in, represents the effective window size of the dilated convolution, represents the depth of the input data in 3D convolution, Represents the dilated convolution rate; when the dilated convolution rate is greater than 1, the elements in the kernel function are no longer continuous, but are scanned by sliding by skipping a set number of intervals.

[0017] In the above multimodal emotion recognition method based on DLKA and EEGNet, in step S2, the floating-point operation amount of the neural network in the DLKA module is and parameter amount The calculation formula is:

[0018]

[0019]

[0020] in, Indicates the height of the input feature map.

[0021] In the multimodal emotion recognition method based on DLKA and EEGNet, in step S2, the DLKA module minimizes Determining the optimal expansion rate , that is, about Taking the derivative and setting it to zero gives :

[0022] .

[0023] In the multimodal emotion recognition method based on DLKA and EEGNet, in step S2, in the DLKA module, the depth convolution is added with the depth layer , then the new parameter and the new model complexity for:

[0024]

[0025] .

[0026] In the above-mentioned multimodal emotion recognition method based on DLKA and EEGNet, in step S3, the EOG signal includes five types of feature data: pupil diameter, dispersion, gaze duration, saccade, and event statistics. The length of each emotion-triggered video is different, and the EOG data is also collected under different time period tasks. Therefore, the EOG data format of each subject is: the number of feature records × the total number of samples is 31 × 2505; in order to ensure the fusion of the EOG signal and the EEG signal, the EOG signal is first dimensionally mapped, and the EOG feature data and the EEG signal are kept in the same tensor dimension through data reconstruction. The EOG signal after the same dimension transformation is spliced ​​with the EEG feature signal to obtain a complete multimodal emotion feature, which is then input into the subsequent deep learning model.

[0027] The beneficial effects of the present invention are:

[0028] 1. This paper proposes a novel multimodal emotion recognition method by combining the variable large kernel attention (DLKA) mechanism with EEGNet. The DLKA module employs a variable large kernel convolution strategy, which adaptively adjusts the receptive field while reducing the number of parameters and computational complexity. The DLKA module also enhances the deep convolutional network's ability to mine global information from EEG signals, thereby improving the robustness and accuracy of the emotion recognition model.

[0029] 2. The present invention can effectively fuse EEG and EOG feature data, adopt an EOG dimension alignment strategy based on feature mapping, and avoid the information loss problem existing in traditional methods; through multimodal feature fusion, the complementary information of EEG and EOG signals can be fully utilized, thereby improving the emotion recognition ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 Flowchart of the present invention.

[0031] Figure 2 This is a framework diagram of the DLKA-EEGNet model of the present invention.

[0032] Figure 3 This is a framework diagram of the DLKA model of the present invention.

[0033] Figure 4 This is a comparison diagram of the present invention and other models.

[0034] Figure 5 This is the confusion matrix diagram of the first representative subject under EEG single modality.

[0035] Figure 6 This is the confusion matrix diagram of the first representative subject after EEG and EOG multimodal fusion.

[0036] Figure 7 This is the confusion matrix diagram of the second representative subject under EEG single modality.

[0037] Figure 8 This is the confusion matrix diagram of the second representative subject after EEG and EOG multimodal fusion.

[0038] Figure 9 This is the confusion matrix diagram of the third representative subject under EEG single modality.

[0039] Figure 10 This is the confusion matrix diagram of the third representative subject after EEG and EOG multimodal fusion.

[0040] Figure 11 This is the confusion matrix diagram of the fourth representative subject under EEG single modality.

[0041] Figure 12 This is the confusion matrix diagram of the fourth representative subject after EEG and EOG multimodal fusion. DETAILED DESCRIPTION

[0042] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0043] like Figure 1 As shown in FIG, a multimodal emotion recognition method based on DLKA and EEGNet includes the following steps:

[0044] S1: Extract features and preprocess the electroencephalogram signal, i.e., the EEG signal.

[0045] EEG signals were recorded using 62 electrode channels. Participants viewed 24 emotion-triggering videos of varying length at different experimental stages. A 1-75 Hz bandpass filter was used to remove noise while maintaining a 200 Hz sampling rate. The EEG data for each emotion-trigger video ranged from 62 × 9601 to 62 × 51801 (number of channels × number of sampling points per channel). The average number of sampling points per channel was calculated to be 28131. Therefore, the sample size of all raw EEG signals was 62 × 28131. This processing yielded filtered EEG signals with clear time-frequency characteristics, which served as input for subsequent feature extraction.

[0046] S2: Construct the DLKA-EEGNet model and use the DLKA-EEGNet model to extract EEG features.

[0047] like Figure 2 As shown in the figure, the DLKA-EEGNet model includes a time-frequency module, a spatial module and a variable large-core attention mechanism module. The variable large-core attention mechanism module is the DLKA module. The EEG signal is sent to the time-frequency module to obtain feature F1. Feature F1 is convolved by the spatial module to obtain feature F2. Feature F2 is input into the DLKA module for deep feature mining to obtain feature F3.

[0048] The time-frequency module uses 8 time-frequency filters with a convolution kernel size of 1×64 to extract the time-frequency features of the original EEG signal. The step size is set to 1 to ensure that the time-frequency features of the signal are fully learned. The time-frequency module is used to mine the time-frequency information in the EEG signal. The model learns the frequency filtering features by setting the time convolution operation and uses a one-dimensional convolution kernel for all EEG channels to learn the time sampling characteristics of each channel. Through the time-frequency convolution process, the feature generated is feature F1.

[0049] The spatial module extracts global spatial features from the EEG signal through depthwise and separable convolution operations. Feature F1 is fed into the deep convolution module for spatial filtering. Feature extraction is performed using a convolutional layer with a depth of 2. Using a convolution kernel of 62×1, 16 spatial filter features are obtained. At this stage, the spatial features of the EEG signal are extracted and ready for subsequent processing. Feature dimensionality reduction is performed using a 1×4 convolution kernel and an average pooling layer with a stride of 4, resulting in a 16×1×50 feature vector, or feature F2.

[0050] like Figure 3 As shown in the figure, the DLKA module uses two variable convolution kernels: Deform-3 Conv2D and Deform-5Conv2D, corresponding to convolution kernel sizes of 3×3 and 5×5 respectively. In the convolution layer of each variable convolution kernel, a convolution layer is set up specifically responsible for calculating the offset. These offsets will be dynamically adjusted according to the size of the convolution kernel to better extract the local and global features of the data.

[0051] In the DLKA module, the large convolution kernel provides a receptive field similar to the self-attention mechanism. By using depthwise convolution, depthwise dilated convolution, and 1×1 point convolution, the large convolution kernel can be constructed with fewer parameters. , the input dimension is , the kernel size is The depth convolution layer and the depth expansion convolution layer, represents the height of the input feature map, Represents the convolution kernel size. The kernel functions of the depth convolution layer and the depth expansion convolution layer are as follows:

[0052]

[0053]

[0054] in, Represents the effective window size of the dilated convolution, which is equivalent to the actual coverage area of ​​the large kernel. Indicates the depth of the input data in 3D convolution (in 2D convolution =1, can be ignored), Represents the dilated convolution rate; when the dilated convolution rate is greater than 1, the elements in the kernel function are no longer continuous, but are scanned by sliding by skipping a set number of intervals, which can increase the size of the receptive field while keeping the number of parameters of the convolution operation relatively small.

[0055] Floating-point operations of neural networks and parameter amount The calculation formula is:

[0056]

[0057]

[0058] in, Indicates the height of the input feature map.

[0059] The number of floating-point operations will increase linearly with the size of the input dimension, and the number of parameters will increase quadratically with the number of channels and kernel size. Generally speaking, the number of floating-point operations and parameters are relatively small, so they will not become limiting factors. In order to reduce the number of parameters for a fixed kernel size, we can minimize Determining the optimal expansion rate , that is, about Taking the derivative and setting it to zero gives :

[0060] .

[0061] Adding depth layers to depthwise convolution , then the new parameter and the new model complexity for:

[0062]

[0063] .

[0064] S3: Extract features from the electrooculogram signal, i.e., the EOG signal, and perform multimodal feature fusion on the EOG signal and the EEG signal.

[0065] To effectively fuse EOG and EEG signals, the present invention uses smoothed EOG signals. EOG signals include five types of feature data: pupil diameter, dispersion, fixation duration, saccades, and event statistics. Each emotion-triggered video has a different length, and EOG data is collected during different time periods. Therefore, the EOG data format for each subject is: number of feature records × total number of samples = 31 × 2505. To ensure fusion of the EOG and EEG signals, the EOG signals are first dimensionally mapped. Data reconstruction is then performed to maintain the same tensor dimensions as the EEG signal. The EOG signals, after undergoing the same dimensionality transformation, are then concatenated with the EEG feature signals to obtain a complete multimodal emotion feature, which is then input into the subsequent deep learning model.

[0066] The present invention deeply mines the characteristic information of EEG and EOG signals. Through the refined feature extraction of time-frequency module and spatial module and the deep feature mining of DLKA module, it not only optimizes the computational efficiency of the model, but also effectively reduces the computational overhead. It also makes full use of the complementary information of EEG and EOG signals, enhances the performance of emotion recognition tasks, and provides a new idea for the field of emotion recognition.

[0067] S4: Input the fused features into the deep learning model for emotion recognition.

[0068] The present invention can be used in scenarios such as mental health monitoring, intelligent interaction, and driving safety. For example, in intelligent healthcare, it can assist in the detection of mood disorders such as depression and anxiety by collecting patients' EEG and EOG signals in real time, analyzing their emotional state, and providing doctors with more objective data support. In intelligent human-computer interaction, it can enhance the system's ability to perceive the user's emotional state, enabling more natural human-computer interaction. For example, intelligent customer service systems can adjust their response strategies based on the user's emotional state, improving the user experience. In the field of intelligent driving, it can be used to monitor driver emotions and fatigue, and combined with the intervention mechanism of the vehicle system, reduce the risk of traffic accidents. In brain science research and neurorehabilitation training, it can be used to analyze brain activity patterns under different emotional states and assist patients with cognitive impairment in rehabilitation training.

[0069] Experimentation and evaluation

[0070] 1. Experimental Environment and Implementation: The entire experiment was conducted on the SEED-IV dataset. This dataset employed a paradigm in which physiological modal signals generated by 15 subjects (7 males and 8 females) were collected in response to movie clips during three different task periods. In each task period, each subject viewed 24 movie clips. These 24 clips contained four emotion-evoking types: neutral, sad, fearful, and happy, labeled 0, 1, 2, and 3, respectively. Each emotion-evoking type contained six movie clips. Because each movie clip had a different duration, the number of sampling points varied. To maintain data consistency across all subjects, the number of sampling points was controlled to be the same for each movie clip. The SEED-IV dataset, serving as the multimodal emotion dataset for this experiment, contains both EEG and EOG physiological modal signals. The experimental environment is Python 3.9, the CPU processor is 13th Gen Intel (R) Core (TM) i9-13900HX, and the GPU processor is NVIDIA GTX 4060. The proposed DLKA-EEGNet model is implemented using the PyTorch framework, and the data sample size for each input is 64.

[0071] 2. Comparative Analysis of Experimental Results: To validate the advantages of this invention, we conducted an experimental comparison of the DLKA-EEGNet model with four comparison models (deepConvNet, ImageNet, EEGNet, and FBCNet) on the SEED-IV dataset. Table 1 shows the recognition accuracy of the different models on 15 subjects.

[0072]

[0073] As shown in Table 1, the proposed DLKA-EEGNet model achieved an average accuracy of 98.65% in the emotion recognition task across 15 subjects, significantly outperforming the other comparison models. The highest accuracy was 99.86% (Subject 7) and the lowest was 95.79% (Subject 2). Comparing the EEGNet and DLKA-EEGNet models reveals that the proposed method enhances EEG signal feature extraction capabilities by adding a time-frequency feature extraction module and separable convolutional layers to the deep convolutional layers. The introduction of the DLKA module further strengthens the model's global feature extraction capabilities.

[0074] Figure 4 The average accuracy and standard deviation of each model are described in . By analyzing the standard deviation of each model, it can be seen that the standard deviation of DLKA-EEGNet is 0.61, indicating that the model performs more stably on multiple subjects and has better robustness.

[0075] In order to verify the role of EOG signals in this experiment, Figure 5-Figure 12 The confusion matrices for four representative subjects are presented in turn, using EEG unimodality and EEG and EOG multimodality fusion. The labels are neutral, sad, fearful, and happy. Each row represents the actual category, and each column represents the predicted value. The predicted values ​​in the confusion matrix show that DLKA-EEGNet demonstrates good recognition accuracy in both EEG unimodality and EEG and EOG multimodality fusion emotion recognition tasks, demonstrating that the proposed model fully utilizes the characteristics of the original EEG data. Furthermore, a comparison of the confusion matrices for unimodal EEG and multimodal EEG reveals that the addition of the EOG modality reduces the recognition error for sadness and fear, indicating that EOG signals are helpful in distinguishing these two emotions.

[0076] Experimental results demonstrate that the DLKA-EEGNet model outperforms existing deep learning models in multimodal emotion recognition tasks. By incorporating time-frequency feature extraction, deep convolution, and the DLKA module, the DLKA-EEGNet model not only improves emotion recognition accuracy but also demonstrates enhanced robustness. Furthermore, the inclusion of EOG signals plays a significant role in emotion recognition, particularly in reducing classification errors for sadness and fear. This further demonstrates the potential of multimodal fusion in emotion recognition and provides new research insights for the field of affective computing.

Claims

1. A multimodal emotion recognition method based on DLKA and EEGNet, characterized in that: The following steps are involved: S1: Extract features and preprocess EEG signals; S2: Build the DLKA-EEGNet model and use it to extract EEG features; The DLKA-EEGNet model includes a time-frequency module, a spatial module, and a variable large-core attention mechanism module. The variable large-core attention mechanism module is the DLKA module. The EEG signal is fed into the time-frequency module to obtain feature F1. Feature F1 is convolved with the spatial module to obtain feature F2. Feature F2 is then fed into the DLKA module for deep feature mining to obtain feature F3. The time-frequency module is used to mine the time-frequency information in the EEG signal. The model learns the frequency filtering characteristics by setting the time convolution operation and uses a one-dimensional convolution kernel for all EEG channels to learn the time sampling characteristics of each channel. The feature generated by the time-frequency convolution process is feature F1; The spatial module extracts the global spatial features of the EEG signal through deep convolution and separable convolution operations, and obtains feature F2 after convolution processing of feature F1; S3: Extract features from the electrooculogram signal, i.e., the EOG signal, and perform multimodal feature fusion of the EOG signal and the EEG signal; S4: Input the fused features into the deep learning model for emotion recognition.

2. The multimodal emotion recognition method based on DLKA and EEGNet according to claim 1, characterized in that In step S1, EEG signals were recorded and acquired using 62 electrode channels. The subjects watched 24 emotion-triggering videos at different experimental stages. Each emotion-triggering video had a different length. A 1-75 Hz bandpass filter was used to eliminate noise while maintaining a sampling rate of 200 Hz. The EEG data size for each emotion-triggering video ranged from 62 × 9601 to 62 × 51801.

3. The multimodal emotion recognition method based on DLKA and EEGNet according to claim 1, characterized in that In step S2, the DLKA module uses two variable convolution kernels: Deform-3 Conv2D and Deform-5 Conv2D, which correspond to convolution kernel sizes of 3×3 and 5×5, respectively. A large convolution kernel is constructed by using depthwise convolution, depthwise dilated convolution and 1×1 point convolution; a channel number of , the input dimension is , the kernel size is The depth convolution layer and the depth expansion convolution layer, represents the height of the input feature map, Represents the convolution kernel size. The kernel functions of the depth convolution layer and the depth expansion convolution layer are as follows: ; ; in, represents the effective window size of the dilated convolution, represents the depth of the input data in 3D convolution, Represents the dilated convolution rate; when the dilated convolution rate is greater than 1, the elements in the kernel function are no longer continuous, but are scanned by sliding by skipping a set number of intervals.

4. The multimodal emotion recognition method based on DLKA and EEGNet according to claim 3, characterized in that In step S2, in the DLKA module, the floating point operation amount of the neural network, i.e., the model complexity and parameter amount The calculation formula is: ; ; in, Indicates the height of the input feature map.

5. The multimodal emotion recognition method based on DLKA and EEGNet according to claim 4, characterized in that In step S2, the DLKA module minimizes Determining the optimal expansion rate , that is, about Taking the derivative and setting it to zero gives : 。 6. The multimodal emotion recognition method based on DLKA and EEGNet according to claim 5, characterized in that In step S2, in the DLKA module, add the depth convolution , then the new parameter and the new model complexity for: ; 。 7. The multimodal emotion recognition method based on DLKA and EEGNet according to claim 6, characterized in that In step S3, the EOG signal includes five types of feature data: pupil diameter, dispersion, gaze duration, saccade, and event statistics. The length of each emotion-triggered video is different, and the EOG data is also collected under different time period tasks. Therefore, the EOG data format of each subject is: number of feature records × total number of samples is 31 × 2505; in order to ensure the fusion of EOG signals and EEG signals, the EOG signal is first dimensionally mapped, and the EOG feature data and the EEG signal are kept in the same tensor dimension through data reconstruction. The EOG signal after the same dimension transformation is spliced ​​with the EEG feature signal to obtain a complete multimodal emotion feature, which is then input into the subsequent deep learning model.

Citation Information

Patent Citations

  • Double-current self-adaptive convolution circulation mixed electroencephalogram emotion recognition method combined with attention mechanism

    CN118797410A