Multi-modal emotion recognition method based on DLKA and EEGNet

Through the multimodal emotion recognition method combined with DLKA and EEGNet models, the problem of multimodal data fusion is solved, and high-precision emotion recognition is achieved, especially in the feature extraction of EEG signals and the complementary information utilization of EOG signals, which improves the robustness and accuracy of emotion recognition.

CN120323974AActive Publication Date: 2025-07-18HUNAN UNIV OF SCI & TECH

Patent Information

Application Number
CN202510820278.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate multimodal data, especially the high dimension and complexity of EEG signals, while ensuring the robustness of the model, affecting the accuracy of emotion recognition.

Method used

DLKA and EEGNet models are adopted to extract EEG signal characteristics through time-frequency modules, space modules and variable large core attention mechanism modules, and multi-modal features are fused with EOG signals. The variable convolution kernel is used to reduce the parameter amount and enhance the global information mining of deep convolution networks.

Benefits of technology

It improves the robustness and accuracy of the emotion recognition model, effectively integrates EEG and EOG feature data, reduces the computational complexity, makes full use of the complementary information of the signal, and improves the emotion recognition ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120323974A_ABST
    Figure CN120323974A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal emotion recognition method based on DLKA and EEGNet. The multi-modal emotion recognition method comprises the following steps: performing feature extraction and preprocessing on EEG signals; a DLKA-EEGNet model is constructed, and the DLKA-EEGNet model is adopted to extract electroencephalogram characteristics; carrying out feature extraction on the EOG signal, and carrying out multi-modal feature fusion on the EOG signal and the EEG signal; and inputting the fused features into a deep learning model for emotion recognition. According to the method, a model combining DLKA and EEGNet is introduced, a brand-new multi-modal emotion recognition method is provided, a DLKA module adopts a convertible large kernel convolution strategy, a receptive field can be adjusted in a self-adaptive mode, meanwhile, the parameter quantity is reduced, the calculation complexity is reduced, the DLKA module reinforces global information mining of a deep convolutional network on EEG signals, and the recognition efficiency is improved. Therefore, the robustness and the accuracy of the emotion recognition model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of emotion recognition, and particularly to a multimodal emotion recognition method based on DLKA and EEGNet. Background Art

[0002] Multimodal emotion recognition has broad application prospects in fields such as intelligent healthcare, intelligent human-computer interaction, and mental health monitoring. Traditional emotion recognition methods usually rely on single-modal data, such as facial expressions, speech, or biological signals (such as electroencephalogram EEG). However, how to effectively fuse multimodal data while ensuring the robustness of the model and deeply mine the original features of EEG signals remains a key challenge in the current field of emotion recognition. The EEG signal itself has high dimensionality and complexity, and how to effectively extract the emotional information therein has an important impact on the accuracy of emotion recognition. Summary of the Invention

[0003] To solve the above technical problems, the present invention provides a multimodal emotion recognition method based on DLKA and EEGNet with a simple algorithm and high recognition accuracy.

[0004] The technical solution of the present invention to solve the above technical problems is: A multimodal emotion recognition method based on DLKA and EEGNet, comprising the following steps:

[0005] S1: Extract features from the electroencephalogram signal, i.e., the EEG signal, and perform preprocessing;

[0006] S2: Construct a DLKA-EEGNet model, and use the DLKA-EEGNet model to extract electroencephalogram features;

[0007] S3: Extract features from the electrooculogram signal, i.e., the EOG signal, and perform multimodal feature fusion of the EOG signal and the EEG signal;

[0008] S4: Input the fused features into a deep learning model for emotion recognition.

[0009] In the above multimodal emotion recognition method based on DLKA and EEGNet, in step S1, the EEG signal is recorded and obtained by 62 electrode channels. The subject watches 24 emotion-triggering videos in different experimental stages, and the length of each emotion-triggering video is different. On the premise of maintaining a sampling rate of 200Hz, a 1-75Hz band-pass filter is used to eliminate noise; the size range of the EEG data of each emotion-triggering video is from 62×9601 to 62×51801.

[0010] The above-mentioned multi-modal emotion recognition method based on DLKA and EEGNet. In step S2, the DLKA-EEGNet model includes a time-frequency module, a spatial module, and a variable large kernel attention mechanism module. The variable large kernel attention mechanism module is the DLKA module. The EEG signal is sent into the time-frequency module to obtain feature F1. Feature F1 is processed through convolution in the spatial module to obtain feature F2. Feature F2 is input into the DLKA module for in-depth feature mining to obtain feature F3.

[0011] The above-mentioned multi-modal emotion recognition method based on DLKA and EEGNet. In step S2, the time-frequency module is used to mine the time-frequency information in the EEG signal. The model learns frequency filtering features by setting time convolution operations and uses one-dimensional convolutional kernels for all EEG channels, thereby learning the time sampling characteristics of each channel. Through the time-frequency convolution process, the generated feature is feature F1.

[0012] The above-mentioned multi-modal emotion recognition method based on DLKA and EEGNet. In step S2, the spatial module extracts the global spatial features of the EEG signal through depth convolution and separable convolution operations. After feature F1 is processed through convolution, feature F2 is obtained.

[0013] The above-mentioned multi-modal emotion recognition method based on DLKA and EEGNet. In step S2, the DLKA module uses two variable convolutional kernels: Deform-3 Conv2D and Deform-5 Conv2D, corresponding to convolutional kernel sizes of 3×3 and 5×5 respectively. A large convolutional kernel is constructed by using depth convolution, depth dilated convolution, and 1×1 point convolution; a depth convolutional layer and a depth dilated convolutional layer with a number of channels of and an input dimension of and a kernel size of are constructed. represents the height of the input feature map, represents the convolutional kernel size. The kernel functions of the depth convolutional layer and the depth dilated convolutional layer are as follows:

[0014]

[0015]

[0016] Among them, represents the effective window size of the dilated convolution, represents the depth of the input data in 3D convolution, represents the dilation rate of the dilated convolution; when the dilation rate is greater than 1, the elements in the kernel function are no longer continuous, but slide and scan by skipping a set number of intervals.

[0017] In the above multi-modal emotion recognition method based on DLKA and EEGNet, in step S2, in the DLKA module, the floating-point operation amount of the neural network and the number of parameters The calculation formulas are as follows:

[0018]

[0019]

[0020] Wherein, represents the height of the input feature map.

[0021] In the above multi-modal emotion recognition method based on DLKA and EEGNet, in step S2, the DLKA module determines the optimal dilation rate by minimizing , that is, taking the derivative of with respect to and setting the derivative to zero to obtain : :

[0022] .

[0023] In the above multi-modal emotion recognition method based on DLKA and EEGNet, in step S2, in the DLKA module, adding the depth layer number in the depth convolution, then the new number of parameters and the new model complexity are:

[0024]

[0025] .

[0026] In the above multi-modal emotion recognition method based on DLKA and EEGNet, in step S3, the EOG signal includes five types of feature data: pupil diameter, chromatic dispersion, fixation duration, saccade, and event statistics. The length of each emotion-triggering video is different, and the EOG data is also collected under different time period tasks. Therefore, the EOG data format of each subject is: the number of feature records × the total number of samples is 31×2505; in order to ensure the fusion of the EOG signal and the EEG signal, first perform dimension mapping on the EOG signal, and make the EOG feature data and the EEG signal maintain the same tensor dimension through data reconstruction. Concatenate the EOG signal after the same-dimension transformation with the EEG feature signal to obtain the complete multi-modal emotion feature, and then input it into the subsequent deep learning model.

[0027] The beneficial effects of the present invention are as follows:

[0028] 1. The present invention proposes a novel multi-modal emotion recognition method by introducing a model that combines the deformable large kernel attention mechanism (DLKA) with EEGNet. The DLKA module adopts a deformable large kernel convolution strategy, which can adaptively adjust the receptive field while reducing the number of parameters and computational complexity. The DLKA module also enhances the global information mining of EEG signals by the deep convolutional network, thereby improving the robustness and accuracy of the emotion recognition model.

[0029] 2. The present invention can effectively fuse EEG and EOG feature data, adopt an EOG dimension alignment strategy based on feature mapping to avoid the information loss problem existing in traditional methods; through multi-modal feature fusion, the complementary information of EEG and EOG signals is fully utilized to improve the emotion recognition ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 is a flowchart of the present invention.

[0031] Figure 2 is a framework diagram of the DLKA-EEGNet model of the present invention.

[0032] Figure 3 is a framework diagram of the DLKA model of the present invention.

[0033] Figure 4 is a comparison result diagram of the present invention with other models.

[0034] Figure 5 is a confusion matrix diagram of the first representative subject in the EEG unimodal mode.

[0035] Figure 6 is a confusion matrix diagram of the first representative subject after multi-modal fusion of EEG and EOG.

[0036] Figure 7 is a confusion matrix diagram of the second representative subject in the EEG unimodal mode.

[0037] Figure 8 is a confusion matrix diagram of the second representative subject after multi-modal fusion of EEG and EOG.

[0038] Figure 9 is a confusion matrix diagram of the third representative subject in the EEG unimodal mode.

[0039] Figure 10 is a confusion matrix diagram of the third representative subject after multi-modal fusion of EEG and EOG.

[0040] Figure 11 is a confusion matrix diagram of the fourth representative subject in the EEG unimodal mode.

[0041] Figure 12 It is the confusion matrix diagram after the multimodal fusion of EEG and EOG for the fourth representative subject. Specific implementation manners

[0042] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0043] As Figure 1 shown, a multimodal emotion recognition method based on DLKA and EEGNet includes the following steps:

[0044] S1: Extract features from electroencephalogram signals, that is, EEG signals, and perform preprocessing.

[0045] The EEG signals are recorded and obtained by 62 electrode channels. The subject watches 24 emotion-triggering videos in different experimental stages. The length of each emotion-triggering video is different. On the premise of maintaining a sampling rate of 200Hz, a 1-75Hz band-pass filter is used to eliminate noise; the size range of the EEG data of each emotion-triggering video is 62×9601 to 62×51801 (number of channels × number of sampling points in a single channel). After calculation, the average number of sampling points in each channel is 28131. Therefore, the sample size of all original EEG signals is 62×28131 dimensions. Through processing, EEG signals with clear time-frequency features after filtering can be obtained as the input for subsequent feature extraction.

[0046] S2: Construct a DLKA-EEGNet model, and use the DLKA-EEGNet model to extract electroencephalogram features.

[0047] As Figure 2 shown, the DLKA-EEGNet model includes a time-frequency module, a spatial module, and a variable large kernel attention mechanism module. The variable large kernel attention mechanism module is the DLKA module. The EEG signal is sent into the time-frequency module to obtain feature F1. Feature F1 is processed by convolution in the spatial module to obtain feature F2. Feature F2 is input into the DLKA module for in-depth feature mining to obtain feature F3.

[0048] The time-frequency module uses 8 time-frequency filters with a convolution kernel size of 1×64 to extract time-frequency features from the original EEG signal, and the stride is set to 1 to ensure that the time-frequency features of the signal are fully learned; the time-frequency module is used to mine the time-frequency information in the EEG signal. The model learns frequency filtering features by setting time convolution operations and uses one-dimensional convolution kernels for all EEG channels to learn the time sampling characteristics of each channel. Through the time-frequency convolution process, the generated feature is feature F1.

[0049] The spatial module extracts the global spatial features of EEG signals through depth convolution and separable convolution operations. The feature F1 is input into the depth convolution module for spatial filtering. A convolutional layer with a depth of 2 is used for feature extraction, and 16 spatial filter features are obtained through a convolutional kernel of size 62×1. At this stage, the spatial features of the EEG signals are extracted and ready for subsequent processing. A convolutional kernel of 1×4 and an average pooling layer with a stride of 4 are used for feature dimensionality reduction, and finally a feature vector of 16×1×50 is obtained, namely feature F2.

[0050] As Figure 3 shown, the DLKA module uses two variable convolutional kernels: Deform-3 Conv2D and Deform-5Conv2D, corresponding to convolutional kernel sizes of 3×3 and 5×5 respectively. In the convolutional layer of each variable convolutional kernel, a convolutional layer responsible for calculating the offsets is set, and these offsets are dynamically adjusted according to the size of the convolutional kernel, so as to better extract the local and global features of the data.

[0051] In the DLKA module, the large convolutional kernel provides a receptive field similar to the self-attention mechanism. By using depth convolution, depth dilated convolution and 1×1 point convolution, a large convolutional kernel can be constructed with fewer parameters; construct a depth convolutional layer and a depth dilated convolutional layer with a channel number of , an input dimension of , and a kernel size of . represents the height of the input feature map, represents the convolutional kernel size. The kernel functions of the depth convolutional layer and the depth dilated convolutional layer are as follows:

[0052]

[0053]

[0054] Among them, represents the effective window size of the dilated convolution, and its role is to equivalently cover the actual area of the large kernel. represents the depth of the input data in 3D convolution (in 2D convolution = 1, which can be ignored). represents the dilation rate of the dilated convolution; when the dilation rate is greater than 1, the elements in the kernel function are no longer continuous, but slide and scan by skipping a set number of intervals, which can increase the size of the receptive field while keeping the number of parameters of the convolution operation relatively small.

[0055] The floating-point operation amount and the number of parameters of the neural network are calculated as follows:

[0056]

[0057]

[0058] Among them, represents the height of the input feature map.

[0059] Among them, the floating-point operation amount will increase linearly with the size of the input dimension, and the number of parameters will increase quadratically with the number of channels and the kernel size. Generally speaking, the floating-point operation amount and the number of parameters are relatively small, so they will not be limiting factors. To reduce the number of parameters of the fixed kernel size, by minimizing determine the optimal dilation rate , that is, for with respect to take the derivative and set the derivative to zero to obtain :

[0060] .

[0061] Adding the depth layer number in the depth convolution, then the new number of parameters and the new model complexity are:

[0062]

[0063] .

[0064] S3: Extract features from the electrooculogram signal, that is, the EOG signal, and perform multi-modal feature fusion on the EOG signal and the EEG signal.

[0065] In order to effectively fuse the EOG signal and the EEG signal, the present invention uses the smoothed EOG signal. The EOG signal includes five types of feature data: pupil diameter, dispersion, fixation duration, saccade, event statistics. The length of each emotion-triggering video is different, and the EOG data is also collected under different time period tasks. Therefore, the EOG data format of each subject: the number of feature records × the total number of samples is 31×2505; in order to ensure the fusion of the EOG signal and the EEG signal, first perform dimensional mapping on the EOG signal, and make the EOG feature data and the EEG signal maintain the same tensor dimension through data reconstruction. Concatenate the EOG signal after the same-dimensional transformation with the EEG feature signal to obtain the complete multi-modal emotion feature, and then input it into the subsequent deep learning model.

[0066] The present invention deeply explores the characteristic information of EEG and EOG signals. Through the refined feature extraction of the time-frequency module and the spatial module, as well as the deep feature mining of the DLKA module, it not only optimizes the computational efficiency of the model, effectively reduces the computational overhead, but also makes full use of the complementary information of EEG and EOG signals, enhances the performance of the emotion recognition task, and provides a new idea for the field of emotion recognition.

[0067] S4: Input the fused features into a deep learning model for emotion recognition.

[0068] The present invention can be used in scenarios such as mental health monitoring, intelligent interaction, and driving safety. For example, in intelligent healthcare, it can assist in the detection of mood disorders such as depression and anxiety. By collecting the EEG and EOG signals of patients in real time and analyzing their emotional states, it provides more objective data support for doctors. In intelligent human-computer interaction, it can improve the system's ability to perceive the user's emotional state and achieve more natural human-computer interaction. For example, an intelligent customer service system can adjust its response strategy according to the user's emotional state to improve the user experience. In the field of intelligent driving, it can be used for driver emotion and fatigue monitoring. Combined with the intervention mechanism of the in-vehicle system, it reduces the risk of traffic accidents. In brain science research and neurorehabilitation training, it can be used to analyze the brain's activity patterns in different emotional states and assist cognitive impairment patients in rehabilitation training.

[0069] Experiment and Evaluation

[0070] 1. Experimental Environment and Implementation: The entire experiment was tested on the SEED-IV dataset. The experimental paradigm of this data: Physiological modality signals generated by 15 subjects (7 males and 8 females) according to movie clips in three different period tasks were collected. In each period task, each subject would watch 24 movie clips. These 24 movie clips included four emotion induction types: neutral, sad, fearful, and happy, which were labeled with tags 0, 1, 2, and 3 respectively. Among them, each emotion induction type had 6 movie clips. Since the durations of each movie clip were different, the number of sampling points for each movie clip was also different. To maintain the consistency of the data information of all subjects, the subjects were controlled to have the same number of samples when watching the same movie clip. The SEED-IV dataset, as the multimodal emotion dataset for this experiment, included two types of physiological modality signals: EEG and EOG. The experimental environment was Python 3.9, CPU processor 13th Gen Intel (R) Core (TM) i9-13900HX, GPU processor NVIDIA GTX 4060. The proposed DLKA-EEGNet model was implemented using the PyTorch framework, and the data sample size input each time was 64.

[0071] 2. Comparative Analysis of Experimental Results: To verify the advantages of the present invention, on the SEED-IV dataset, the DLKA-EEGNet model was experimentally compared with four comparative models (deepConvNet, ImageNet, EEGNet, FBCNet). Table 1 shows the recognition accuracies of different models for 15 subjects.

[0072]

[0073] As can be seen from Table 1, in the emotion recognition task for 15 subjects, the proposed DLKA-EEGNet model achieved an average accuracy of 98.65%, significantly outperforming other comparative models. The highest accuracy was 99.86% (Subject 7), and the lowest was 95.79% (Subject 2). By comparing the EEGNet and DLKA-EEGNet models, it can be concluded that the present invention adds a time-frequency feature extraction module and a separable convolutional layer on the basis of the deep convolutional layer, enhancing the feature extraction ability for EEG signals. The introduction of the DLKA module further enhances the global feature extraction ability of the model.

[0074] Figure 4 The average accuracy and its standard deviation of each model are described. By analyzing the standard deviation of each model, it can be seen that the standard deviation of DLKA-EEGNet is 0.61, indicating that the model performs more stably on multiple subjects and has good robustness.

[0075] To verify the role of EOG signals in this experiment, Figures 5 - 12 The confusion matrices of 4 representative subjects in the EEG single modality and the fusion of EEG and EOG multimodality are shown in sequence. The labels are neutral, sad, fear, and happy. Each row represents the actual category, and each column represents the predicted value. From the predicted values in the confusion matrix, it can be seen that DLKA-EEGNet shows good recognition accuracy in both the EEG single modality and the emotion recognition task of the fusion of EEG and EOG multimodality, indicating that the proposed model makes full use of the features of the original EEG data. In addition, according to the comparison of the two types of confusion matrices of single modality EEG and multimodality, it can be seen that after adding the EOG modality, the recognition errors of the sad and fear emotions are reduced, indicating that the EOG signal helps to distinguish these two types of emotions.

[0076] The experimental results show that the DLKA-EEGNet model outperforms existing deep learning models in multi-modal emotion recognition tasks. By introducing time-frequency feature extraction, deep convolution, and the DLKA module, the DLKA-EEGNet not only improves the accuracy of emotion recognition but also demonstrates stronger robustness. At the same time, the introduction of EOG signals plays an important role in emotion recognition, especially in reducing the classification errors of sad and fearful emotions, which further proves the potential of multi-modal fusion in emotion recognition and provides new research ideas for the field of affective computing.

Claims

1. A multi-modal emotion recognition method based on DLKA and EEGNet, characterized in that, It includes the following steps: S1: Extract features from the electroencephalogram signal, i.e., the EEG signal, and perform preprocessing; S2: Construct the DLKA-EEGNet model, and use the DLKA-EEGNet model to extract electroencephalogram features; S3: Extract features from the electrooculogram signal, i.e., the EOG signal, and perform multi-modal feature fusion of the EOG signal and the EEG signal; S4: Input the fused features into a deep learning model for emotion recognition.

2. The multimodal emotion recognition method based on DLKA and EEGNet according to claim 1, characterized in that, In the step S1, the EEG signal is recorded and obtained by 62 electrode channels. The subject watches 24 emotion-triggering videos in different experimental stages. The length of each emotion-triggering video is different. On the premise of maintaining a sampling rate of 200Hz, a 1-75Hz band-pass filter is used to eliminate noise; the size range of the EEG data for each emotion-triggering video is from 62×9601 to 62×51801.

3. The multimodal emotion recognition method based on DLKA and EEGNet according to claim 2, characterized in that, In the step S2, the DLKA-EEGNet model includes a time-frequency module, a spatial module, and a variable large kernel attention mechanism module. The variable large kernel attention mechanism module is the DLKA module. The EEG signal is sent into the time-frequency module to obtain feature F1. Feature F1 is processed by convolution in the spatial module to obtain feature F2. Feature F2 is input into the DLKA module for deep feature mining to obtain feature F3.

4. The multimodal emotion recognition method based on DLKA and EEGNet according to claim 3, characterized in that In the step S2, the time-frequency module is used to mine the time-frequency information in the EEG signal. The model learns frequency filtering features by setting time convolution operations and uses one-dimensional convolution kernels for all EEG channels, so as to learn the time sampling characteristics of each channel. Through the time-frequency convolution process, the generated feature is feature F1.

5. The multimodal emotion recognition method based on DLKA and EEGNet according to claim 3, characterized in that In the step S2, the spatial module extracts the global spatial features of the EEG signal through depth convolution and separable convolution operations. After feature F1 is processed by convolution, feature F2 is obtained.

6. The multimodal emotion recognition method based on DLKA and EEGNet according to claim 3, wherein, In the step S2, the DLKA module uses two variable convolutional kernels: Deform-3 Conv2D and Deform-5 Conv2D, corresponding to convolutional kernel sizes of 3×3 and 5×5 respectively, and constructs large convolutional kernels by using depth convolution, depth dilated convolution and 1×1 point convolution; construct a depth convolutional layer and a depth dilated convolutional layer with the number of channels being , input dimension being , and kernel size being . The depth convolutional layer and the depth dilated convolutional layer have the following kernel functions: represents the height of the input feature map, represents the kernel size, and the kernel functions of the depth convolutional layer and the depth dilated convolutional layer are as follows: ; ; Among them, represents the effective window size of the dilated convolution, represents the depth of the input data in the 3D convolution, represents the dilation rate of the dilated convolution; when the dilation rate is greater than 1, the elements in the kernel function are no longer continuous, but slide and scan by skipping a set number of intervals.

7. The multimodal emotion recognition method based on DLKA and EEGNet according to claim 6, characterized in that In the step S2, in the DLKA module, the floating-point operation amount of the neural network, that is, the model complexity and the number of parameters The calculation formula is as follows: ; ; Among them, represents the height of the input feature map.

8. The multimodal emotion recognition method based on DLKA and EEGNet according to claim 7, wherein, In the step S2, the DLKA module minimizes to determine the optimal inflation rate , that is, with respect to take the derivative and set the derivative to zero to obtain : 。 9. The multimodal emotion recognition method based on DLKA and EEGNet according to claim 8, characterized in that In the DLKA module in step S2, adding results in new parameter quantity and new model complexity as follows: ; 。 10. The multimodal emotion recognition method based on DLKA and EEGNet according to claim 9, characterized in that, In the step S3, the EOG signal includes five types of feature data: pupil diameter, dispersion, fixation duration, saccade, and event statistics. The length of each emotion-triggering video is different, and the EOG data is also collected under different time period tasks. Therefore, the EOG data format for each subject is: the number of feature records × the total number of samples is 31×2505; in order to ensure the fusion of the EOG signal and the EEG signal, first perform dimension mapping on the EOG signal, and make the EOG feature data and the EEG signal maintain the same tensor dimension through data reconstruction. The EOG signal after the same-dimension transformation is concatenated with the EEG feature signal to obtain complete multi-modal emotion features, so as to be input into the subsequent deep learning model.

Citation Information

Patent Citations

  • Double-current self-adaptive convolution circulation mixed electroencephalogram emotion recognition method combined with attention mechanism

    CN118797410A

Cited By

  • Electroencephalogram and eye movement signal emotion recognition method based on modal generation

    CN122296897A