A Continuous Dynamic Emotion Recognition Method and System Integrating EEG and EOG Signals
Through the continuous dynamic emotion recognition method that fuses EEG and EEG signals, the convolution module, the two-way cross attention mechanism fusion network and the Transformer module are used for feature extraction and emotion recognition, which solves the accuracy and real-time problems of emotion recognition in dynamic situations, and achieves a more efficient emotion recognition effect.
Patent Information
- Application Number
- CN202510336025.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-03-21
AI Technical Summary
Existing emotions recognition research is usually limited to static laboratory situations and fails to fully solve the problem of changes in emotional states in dynamic situations, resulting in insufficient accuracy and real-timeness of emotion recognition.
A continuous dynamic emotion recognition method that combines EEG and EEG signals is adopted. By obtaining the EEG signal data and EEG signal data of the target subject, combining the Big Five personality scores, a convolution module, a two-way cross attention mechanism fusion network, a Transformer module and a full connection layer are used to perform feature extraction and emotion recognition, so as to achieve fusion and timing analysis of multimodal features.
It improves the accuracy and real-time nature of emotion recognition, and can better reflect emotional changes in dynamic situations.
Smart Images

Figure CN119848748B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of emotion recognition, and particularly to a continuous dynamic emotion recognition method and system that integrates electroencephalogram (EEG) and electrooculogram (EOG) signals. Background Art
[0002] Affective computing is an important branch in the field of artificial intelligence, which is dedicated to enhancing the capabilities of robots in emotional expression and perception, and has played an important role in various fields such as human-computer interaction, mental health, and educational science. Emotional expression involves synchronous changes in multiple physiological signals, and single-signal analysis often has data limitations, affecting the accuracy of emotion recognition. Existing emotion recognition research is usually limited to static laboratory settings and fails to fully address the problem of changes in emotional states in dynamic situations.
[0003] Therefore, based on the above problems, there is an urgent need to provide a continuous dynamic emotion recognition method that integrates EEG and EOG signals to improve the accuracy and real-time performance of continuous dynamic emotion recognition. Summary of the Invention
[0004] The purpose of the present application is to provide a continuous dynamic emotion recognition method and system that integrates EEG and EOG signals, which can improve the accuracy and real-time performance of continuous dynamic emotion recognition.
[0005] To achieve the above purpose, the present application provides the following solutions:
[0006] In a first aspect, the present application provides a continuous dynamic emotion recognition method that integrates EEG and EOG signals. The continuous dynamic emotion recognition method that integrates EEG and EOG signals includes:
[0007] Obtain the EEG signal data, EOG signal data, Big Five personality scores, and dynamic emotion labels of the target subject; the Big Five personality scores are obtained through the Big Five personality inventory; the dynamic emotion labels are obtained by the target subject scoring the emotional state from 1 to 9 every 5 seconds according to two dimensions of pleasure and arousal; the corresponding types of scores are calm to happy, happy to calm, sad to calm, calm to sad, tense to calm, and calm to tense; the target subjects include pilots and ordinary subjects; the EOG signal data includes horizontal EOG signals and vertical EOG signals;
[0008] Preprocess the EEG signal data and the EOG signal data respectively to obtain the preprocessed EEG signal data and the preprocessed EOG signal data;
[0009] Extract differential entropy features according to the preprocessed EEG signal data; and classify the target subjects into pilots and ordinary subjects according to the differential entropy features and the Big Five personality scores;
[0010] Based on the preprocessed electroencephalogram (EEG) signal data and preprocessed electrooculogram (EOG) signal data of pilots, as well as the preprocessed EEG signal data and preprocessed EOG signal data of ordinary subjects, emotion recognition is respectively carried out using a continuous dynamic emotion recognition model; the continuous dynamic emotion recognition model includes: a convolutional module, a bidirectional cross-attention mechanism fusion network, a Transformer module, and a fully connected layer; the convolutional module is used to extract EEG signal features and EOG signal features from the preprocessed EEG signal data and the preprocessed EOG signal data by using spatio-temporal convolution technology and a SE attention mechanism network module; the bidirectional cross-attention mechanism fusion network is used to fuse the EEG signal features and the EOG signal features to obtain multi-modal features; the Transformer module is used to extract the temporal features of the multi-modal features during the emotion change process; the fully connected layer is used to perform continuous dynamic emotion recognition based on the temporal features.
[0011] Optionally, the preprocessing of the EEG signal data and the EOG signal data respectively to obtain the preprocessed EEG signal data and the preprocessed EOG signal data specifically includes:
[0012] Performing filtering processing, artifact removal processing, and resampling processing on the EEG signal data in sequence to obtain the preprocessed EEG signal data;
[0013] Performing filtering processing and resampling processing on the EOG signal data in sequence to obtain the preprocessed EOG signal data.
[0014] Optionally, the extraction of differential entropy features according to the preprocessed EEG signal data specifically includes:
[0015] Using the formula To determine the differential entropy features;
[0016] Where Is the differential entropy feature of the EEG signal data, Is the total number of samples, Is the value of the EEG signal data at the nth sampling point, Is the value of the EEG signal data at the (n + 1)th sampling point.
[0017] Optionally, the frequency bands of the differential entropy features include: 1Hz - 3Hz frequency band, 4Hz - 7Hz frequency band, 8Hz - 13Hz frequency band, 14Hz - 30Hz frequency band, and 31Hz - 50Hz frequency band.
[0018] Optionally, before classifying the target subjects into pilots and ordinary subjects according to the differential entropy features and the Big Five personality scores, it also includes:
[0019] Standardize the differential entropy features and Big Five personality scores.
[0020] Optionally, the convolution module includes: the first-layer convolution, the second-layer convolution, the third-layer convolution, the fourth-layer convolution, the batch normalization layer between the first-layer convolution and the second-layer convolution, the batch normalization layer between the second-layer convolution and the third-layer convolution, the dropout layer after the third-layer convolution, and the SE attention mechanism network module between the third-layer convolution and the fourth-layer convolution;
[0021] The first-layer convolution uses 1D convolution to remove noise and extract features;
[0022] The activation function is excluded in the first-layer convolution; the second-layer convolution includes a spatial convolution layer; the third-layer convolution includes an average pooling layer; the SE attention mechanism network module adaptively learns the weights between channels; the fourth-layer convolution uses an advanced feature fusion layer.
[0023] In a second aspect, the present application provides a continuous dynamic emotion recognition system that fuses electroencephalogram and electrooculogram signals. The continuous dynamic emotion recognition system that fuses electroencephalogram and electrooculogram signals includes:
[0024] A data acquisition module, configured to acquire electroencephalogram signal data, electrooculogram signal data, Big Five personality scores, and dynamic emotion labels of a target subject; the Big Five personality scores are obtained through the Big Five personality inventory; the dynamic emotion labels are obtained by the target subject scoring the emotional state from 1 to 9 every 5 seconds according to two dimensions of pleasure and arousal; the corresponding types of the scores are calm to happy, happy to calm, sad to calm, calm to sad, tense to calm, and calm to tense; the target subjects include pilots and ordinary subjects; the electrooculogram signal data includes horizontal electrooculogram signals and vertical electrooculogram signals;
[0025] A preprocessing module, configured to preprocess the electroencephalogram signal data and the electrooculogram signal data respectively to obtain preprocessed electroencephalogram signal data and preprocessed electrooculogram signal data;
[0026] A subject division module, configured to extract differential entropy features according to the preprocessed electroencephalogram signal data; and divide the target subjects into pilots and ordinary subjects according to the differential entropy features and the Big Five personality scores;
[0027] An emotion recognition module, which is used to perform emotion recognition respectively on the preprocessed electroencephalogram (EEG) signal data and preprocessed electrooculogram (EOG) signal data of pilots, as well as the preprocessed EEG signal data and preprocessed EOG signal data of ordinary subjects, by using a continuous dynamic emotion recognition model; the continuous dynamic emotion recognition model includes: a convolution module, a bidirectional cross-attention mechanism fusion network, a Transformer module, and a fully connected layer; the convolution module is used to extract EEG signal features and EOG signal features from the preprocessed EEG signal data and the preprocessed EOG signal data by using spatio-temporal convolution technology and an SE attention mechanism network module; the bidirectional cross-attention mechanism fusion network is used to fuse the EEG signal features and the EOG signal features to obtain multimodal features; the Transformer module is used to extract the temporal features of the multimodal features during the emotion change process; the fully connected layer is used to perform continuous dynamic emotion recognition according to the temporal features.
[0028] According to the specific embodiments provided by the present application, the present application has the following technical effects:
[0029] The present application provides a continuous dynamic emotion recognition method and system that fuses EEG and EOG signals. By collecting the EEG signal data (Electroencephalogram, EEG), EOG signal data (Electrooculogram, EOG), Big Five personality scores, and dynamic emotion labels of the target subject, the EEG signal data and the EOG signal data are preprocessed, and then a hierarchical model architecture is used for recognition; among them, differential entropy features (Differential Entropy, DE) are extracted from the preprocessed EEG signal data in the first layer of the hierarchical model architecture, and the target subject is divided into pilots and ordinary subjects by combining the differential entropy features and the Big Five personality scores. Then, continuous dynamic emotion recognition is performed on the two types of subjects respectively in the second layer of the hierarchical model architecture. Specifically, a convolution module is used to extract features from the preprocessed EEG signal data and the preprocessed EOG signal data, a bidirectional cross-attention mechanism fusion network (Bidirectionalcross-attention mechanism fusion network, BCAF-Net) is used to fuse the EEG features and the EOG features, a Transformer module is used to learn global temporal features from the multimodal features, and finally a fully connected layer is used to determine the continuous dynamic emotion category of the target subject, thereby realizing the recognition of continuous dynamic emotions. The present application improves the accuracy and real-time performance of emotion recognition. Description of the Drawings
[0030] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0031] Figure 1 Schematic flowchart of a continuous dynamic emotion recognition method that fuses electroencephalogram and electrooculogram signals in an embodiment of the present application;
[0032] Figure 2 Detailed flowchart of the continuous dynamic emotion recognition method that fuses electroencephalogram and electrooculogram signals provided by the present application;
[0033] Figure 3 Schematic diagram of the acquisition process of electroencephalogram signal data, electrooculogram signal data, Big Five personality scores, and dynamic emotion labels in the present application;
[0034] Figure 4 Schematic diagram of the process of continuous dynamic emotion recognition that fuses electroencephalogram and electrooculogram signals in the present application. Detailed implementation manners
[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0036] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0037] In an exemplary embodiment, as Figure 1 and Figure 2 shown, a continuous dynamic emotion recognition method that fuses electroencephalogram and electrooculogram signals is provided, and this method includes the following S101 to S104. Among them:
[0038] S101. Obtain the electroencephalogram (EEG) signal data, electrooculogram (EOG) signal data, Big Five personality scores, and dynamic emotion labels of the target subjects. The Big Five personality scores are obtained through the Big Five Inventory (BFI). The dynamic emotion labels are obtained by the target subjects rating their emotional states from 1 to 9 every 5 seconds according to two dimensions: pleasure and arousal. The corresponding types of ratings include calm to happy (representing the dynamic emotional change from calm to happy), happy to calm (representing the dynamic emotional change from happy to calm), sad to calm (representing the dynamic emotional change from sad to calm), calm to sad (representing the dynamic emotional change from calm to sad), tense to calm (representing the dynamic emotional change from tense to calm), and calm to tense (representing the dynamic emotional change from calm to tense). The target subjects include pilots and ordinary subjects. The EOG signal data includes horizontal EOG signal and vertical EOG signal.
[0039] As important representatives of human physiological signals, EEG signals and EOG signals are unbiased and objective. EEG signals reflect the electrical signal changes in brain activities, while EOG signals record the trajectory and frequency of eye movements. Both of these signals can be used to reflect a person's emotional state.
[0040] The Big Five Inventory, also known as the Five-Factor Model of Personality, is a commonly used personality assessment tool in psychology. It evaluates an individual's personality traits through five dimensions (neuroticism, extraversion, openness, agreeableness, and conscientiousness), and each dimension has a corresponding score range. These scores can help us understand an individual's personality characteristics in different aspects.
[0041] Collect the original EEG signal data and original EOG signal data of the target subjects based on the set experimental paradigm. In this embodiment, the original EEG signal data and original EOG signal data are collected in real-time synchronization and stored in the database. The original EEG signal data includes: EEG signal data in the normal mode and EEG signal data in the target task state mode. The original EOG signal data includes: EOG signal data in the normal mode and EOG signal data in the target task state mode.
[0042] Among them, during the training process of the continuous dynamic emotion recognition model, the experimental paradigm is as follows: Select 48 movie clips from the material library as the stimuli used in the experiment. Each movie clip is carefully edited to generate coherent emotions. Each video should evoke two desired target emotions, including six types of stimulus movie clips such as happy calm and calm happy, sad calm and calm sad, and tense calm and calm tense. There are eight movie clips of each type, and the duration is about 3 minutes. As Figure 3As shown in the figure, first, an experimental paradigm for electroencephalogram (EEG) signals and electrooculogram (EOG) signals is created. Through acquisition devices (EEG caps and surface electrodes), the original EEG signal data, original EOG signal data, and dynamic emotion labels of the target subject are synchronously obtained. In actual applications, the experimental paradigm is similar to the training process of a deep learning model, except that dynamic emotion labels are not collected.
[0043] Organize experts to analyze the application scenarios and possible situations, and design a relatively complete experimental paradigm that can accurately reflect the EEG and EOG signals in this cognitive state. Under the condition of isolating external environmental interference, record and collect the discharge signals of the cerebral cortex of the subject during different cognitive activities. At the same time, as a comparison, record the EEG and EOG signals under normal conditions.
[0044] Comprehensively evaluate the quality of the collected EEG signal data and EOG signal data based on the performance of the subject at the acquisition site. Eliminate the original EEG signal data and original EOG signal data that perform poorly and cannot accurately reflect this process, and provide feedback and update the data in the database.
[0045] As Figure 3 shown, as a specific implementation method, the acquisition process of EEG signal data and EOG signal data is as follows.
[0046] 1) Wear EEG signal data and EOG signal data acquisition devices for the subject in a quiet room with appropriate temperature and brightness, and play a 5-second prompt and precautions for the subject. A total of two experiments are conducted. Each experiment is divided into 3 sections. For each section, play a video that induces two emotions for the subject. There are three types of videos that induce different emotional changes, namely happy to calm and calm to happy, calm to sad and sad to calm, calm to nervous and nervous to calm. The duration of each video is about 180 seconds, with each emotion lasting 90 seconds. The interval between the two experiments is more than one week to ensure the independence of each emotion experiment and improve the experimental accuracy.
[0047] 2) The subject conducts the experiment in a dedicated environment that can shield external interference. During the experiment, always set the brightness at a comfortable level so that the subject can clearly see the movie content. To ensure the smooth progress of the experiment, before the official start of the experiment, each subject is required to fill out a mental health assessment questionnaire, and each subject is trained with additional movie clips. In addition, the subject also needs to complete a personality measurement questionnaire to obtain the personality data of the target subject.
[0048] The personality measurement questionnaire is used to obtain data on the five personality traits of the target subject, which are used as the personality attributes of the target subject. When performing the subject classification task, the classification accuracy of the classification task is obtained. The personality measurement questionnaire uses the Chinese version of the Big Five Personality Inventory and includes a total of 44 questions.
[0049] 3) After the entire experiment starts, the presentation order of each movie clip is fixed. The subjects first adjust their emotions to a relatively relaxed state. When the brainwaves of the subjects are stable, the movie clips are played. During the task-free state, 2 seconds of electroencephalogram (EEG) signals and electrooculogram (EOG) signals are collected. During the cognitive activity state, 180 seconds of EEG signals and EOG signals are collected.
[0050] 4) Set the rest time. After each movie clip is watched, rest for 15 to 30 seconds to keep the brainwaves of the subjects in a stable state. Then repeat the above process continuously until the subjects complete all processes of the entire experimental paradigm. The subjects evaluate their self-state every 10 seconds to complete the experiment.
[0051] Five experts are selected to participate in the selection of movie clips. All of them have received professional film training and have experience in rehearsing films. Thirty healthy subjects are selected to participate in the experiment, including 10 pilots and 20 ordinary subjects. The average age is (22 ± 1.5) years old. All participants have normal hearing and vision and are right-handed, and are informed and consent to the experimental content.
[0052] S102, preprocess the EEG signal data and EOG signal data respectively to obtain the preprocessed EEG signal data and the preprocessed EOG signal data;
[0053] S102 specifically includes:
[0054] S21, perform filtering processing, artifact removal processing, and resampling processing on the EEG signal data in sequence to obtain the preprocessed EEG signal data;
[0055] Filter out the EEG signals in the required frequency band from the EEG signal data through a band-pass filter to obtain the filtered EEG signals. Remove the irrelevant information outside the effective EEG frequency band (1 Hz to 50 Hz).
[0056] Perform artifact removal processing on the filtered EEG signal data to obtain the EEG signal processed data. Remove the power frequency interference and artifacts of the filtered EEG signals, that is, remove the disturbances such as eyelid twitching, electrocardiogram, electromyogram, limb shaking, etc. that interfere with the EEG signal data, as well as the 50 Hz power frequency interference.
[0057] Resample the EEG signal processed data to obtain the EEG signal data of the target subject. Downsample the EEG signal processed data, and set the sampling frequency to 128 hz. It can improve the calculation efficiency while retaining the information related to the cognitive state.
[0058] Specifically, import the EEG signal data into ASA, and successively perform electrode positioning, electrode selection, rereferencing, filtering, bad segment rejection, bad channel interpolation, independent component analysis, segmentation, and baseline correction processing, and then save the data.
[0059] S22, perform filtering processing and resampling processing on the EOG signal data in sequence to obtain the preprocessed EOG signal data. That is, remove the 50Hz power frequency interference and artifacts from the EEG signal data, and remove the frequency bands outside the δ (1 - 3Hz) band, (4 - 7Hz) band, (8 - 13Hz) band, (14 - 30Hz) band and (31 - 50Hz) band. Similarly, remove the 50Hz power frequency interference and filtering from the EOG signal data, and remove the frequency bands outside the (0.3 - 30Hz) band. Then segment and classify the EEG signal data and EOG signal data under different conditions. Finally, window the preprocessed EEG signal data and EOG signal data, using a non-overlapping sliding time window with a window of 1s to obtain the preprocessed EEG signal data and EOG signal data.
[0060] Filter out the EOG signals in the required frequency band (0.3Hz - 30Hz) from the EOG signal data through a band-pass filter to obtain the EOG signal processing data. Remove the irrelevant information outside the effective EOG frequency band. Remove the power frequency interference (such as 50Hz or 60Hz power line interference) from the EOG signal data through a notch filter.
[0061] Resample the EOG signal processing data to obtain the EOG signal data of the target subject. Downsample the EOG signal processing data, and set the sampling frequency to 128hz. It can improve the calculation efficiency while retaining the information related to the cognitive state.
[0062] Specifically, import the EOG signal data into ASA, and successively perform filtering, notch filtering, and resampling processing, and then save the data.
[0063] After collecting the EEG signal data and EOG signal data, the subject is also required to score the stimulus film every five seconds in the valence - arousal two - dimensional space, with the numerical range being [1, 9], and perform normalization processing of the dynamic emotion labels.
[0064] S103, extract the differential entropy features according to the preprocessed EEG signal data; and classify the target subject into pilots and ordinary subjects according to the differential entropy features and the Big Five personality scores;
[0065] Calculate the differential entropy features of the EEG signal data in the time domain:
[0066] ;
[0067] Among them, is the differential entropy feature of the EEG signal data, is the total number of samples, is the value of the EEG signal data at the nth sampling point, is the value of the EEG signal data at the (n + 1)th sampling point.
[0068] This application adds a non-overlapping time window of 1 second, and extracts features from the 64-channel information of each 1 second. Extract features.
[0069] Among them, the frequency bands of the differential entropy feature include: 1Hz - 3Hz frequency band, 4Hz - 7Hz frequency band, 8Hz - 13Hz frequency band, 14Hz - 30Hz frequency band, and 31Hz - 50Hz frequency band.
[0070] Combining the differential entropy feature with the Big Five personality score feature of the target subject, input it into a multi-layer convolution for the binary classification task of pilots and ordinary subjects. Since each subject contains multiple samples, these samples may obtain different classification results, that is, some samples are classified as pilots and others are classified as ordinary subjects. To solve this problem, the majority voting method is adopted, and according to the classification results of all samples of each subject, the category of the subject is finally determined.
[0071] S104, respectively use the continuously dynamic emotion recognition model for emotion recognition according to the preprocessed EEG signal data and preprocessed electrooculogram signal data of pilots and the preprocessed EEG signal data and preprocessed electrooculogram signal data of ordinary subjects; as Figure 4 shown, the continuously dynamic emotion recognition model includes: a convolution module, a bidirectional cross-attention mechanism fusion network, a Transformer module, and a fully connected layer; the convolution module is used to extract EEG signal features and electrooculogram signal features from the preprocessed EEG signal data and preprocessed electrooculogram signal data by using spatio-temporal convolution technology and SE attention mechanism network module (Squeeze-and-Excitation Networks, SE); the bidirectional cross-attention mechanism fusion network is used to fuse the EEG signal features and the electrooculogram signal features to obtain multi-modal features; the Transformer module is used to extract the temporal features of the multi-modal features during the emotion change process; the fully connected layer is used to perform continuously dynamic emotion recognition according to the temporal features.
[0072] Specifically, the convolutional module is serially connected to the Transformer module; the bidirectional cross-attention mechanism fusion network performs feature fusion on the features extracted from the EEG signal and the EOG signal after passing through the convolutional module and outputs them as the input of the Transformer module; the fully connected layer is connected to the Transformer module.
[0073] The continuous dynamic emotion recognition model is pre-trained using a training sample set, and each training sample in the training sample set includes an EEG sample signal, an EOG sample signal, and a corresponding dynamic emotion label.
[0074] During the training process, methods such as Bayesian optimization and cross-validation are used to optimize the parameters of the deep learning model. The training set is used for model training, and the validation set is used for model performance evaluation. The optimal model parameters are selected according to the evaluation results. And it is deployed and used in actual applications.
[0075] The continuous dynamic emotion recognition model uses a deep learning model to extract spatial and temporal features from the preprocessed EEG signal and EOG signal to improve the accuracy of the dynamic emotion label of the target subject;
[0076] The convolutional module includes: the first-layer convolution, the second-layer convolution, the third-layer convolution, the fourth-layer convolution, the batch normalization layer between the first-layer convolution and the second-layer convolution, the batch normalization layer between the second-layer convolution and the third-layer convolution, the dropout layer after the third-layer convolution, and the SE attention mechanism network module between the third-layer convolution and the fourth-layer convolution;
[0077] The first-layer convolution uses 1D convolution to remove noise and extract features;
[0078] Among them, a batch normalization layer is inserted between the first-layer convolution and the second-layer convolution to solve the problem of vanishing gradients. Then, in order to maintain the interpretability of spatio-temporal convolution, an activation function is excluded from the first-layer convolution. The second-layer convolution consists of a spatial convolution layer that combines the valid information on all channels. The spatial convolution is separated in the depth dimension and implemented by depth convolution in the code. While the activation function is restored, a batch normalization layer is added between the second-layer convolution and the third-layer convolution. The third-layer convolution includes an average pooling layer with a pool length of 4. To avoid overfitting, a dropout layer is added after the third-layer convolution. An SE attention mechanism network module is added between the third-layer convolution and the fourth-layer convolution. The SE attention mechanism network module adaptively learns the weights between channels, automatically adjusts the responses of each channel, enabling the network to pay more attention to the features helpful for the emotion recognition task and improving the recognition ability of the model. After the SE attention mechanism network module, that is, the fourth-layer convolution adopts an advanced feature fusion layer, and the fourth-layer convolution is implemented using a separable convolution layer, minimizing the number of parameters. Finally, the output of the convolution module is dimension-reduced to adapt to the subsequent module.
[0079] The SE attention mechanism network module is the SE attention mechanism network, which adds an attention mechanism in the channel dimension, and the key operations are Squeeze and Excitation.
[0080] Specifically, for the Squeeze, the two-dimensional features of each channel are compressed into 1 real number; for the Excitation, a weight value is generated for each feature channel, and then the normalized weights are weighted to the features of each channel of the original input.
[0081] In this application, a non-overlapping window of 1 second is added, and the 64-channel information of the electroencephalogram signal data per 1 second and the 2-channel information of the electrooculogram signal data per 1 second are input into the convolution module to extract features.
[0082] The electroencephalogram signal data and the electrooculogram signal data are standardized to ensure the scale consistency between different data and avoid the impact of data being too large or too small on the training of the deep learning model. The standardization process is to use the Z-score standardization method to process the DE features and the Big Five personality scores.
[0083] The bidirectional cross-attention mechanism fusion network enables the model to better focus on the important information parts by establishing an attention mechanism between different modalities. Adjust the representation of the electrooculogram features according to the information of the electroencephalogram features and adjust the representation of the electroencephalogram features according to the information of the electrooculogram features to ensure that the information interaction between the two modalities is bidirectional.
[0084] In the two-way cross-attention mechanism fusion network of this application, first, through the Multi-Head Cross-Attention (MHCA), the electroencephalogram (EEG) signal data and electrooculogram (EOG) signal data are used as queries (Q), keys (K), and values (V) for each other to obtain EEG features and EOG features , and then and are dot-product operated to obtain features . Finally, global fusion is performed. The three obtained features are respectively applied with a fully connected layer FC(·), and is mapped to a 2D space and connected to obtain the final fused feature vector . This application can help the model better capture the mutual influence between EEG signal data and EOG signal data, and enhance the accuracy of continuous dynamic emotion recognition. Specifically as follows:
[0085] (1) Through MHCA, the EEG features and EOG features are fused at the modality level. Specifically, first, the EEG features are used as Q, the EOG features are used as K, and V is input into MHCA for fusion to obtain . Then, the EOG features are used as Q, the EOG features are used as K and V and input into MHCA to obtain ;
[0086] For the queries (Q), keys (K), and values (V) in MHCA, they are first mapped through their respective linear transformations:
[0087] ;
[0088] ;
[0089] ;
[0090] where , , are weight matrices, is the input of the query, is the input of the key and value (simultaneously as the input of the key and value).
[0091] In multi-head attention, the input is split into H heads (H is the number of heads), and each head processes a part of the features. For each head, the attention weight matrix is calculated as follows:
[0092] ;
[0093] where and are the vector representations of the i-th query and the i-th key respectively. The above scores will be used to calculate the attention weights.
[0094] To prevent the dot product result from being too large and causing unstable gradients, a scaling factor is used to scale the dot product result:
[0095] ;
[0096] The Softmax function is used to convert the scaled dot product attention scores into attention weights:
[0097] ;
[0098] where is the attention weight of the i-th query to the i-th key.
[0099] Finally, the attention weights are used to perform a weighted sum on the values to obtain the output of each head:
[0100] ;
[0101] where is the output vector of the i-th head.
[0102] After obtaining the output of each head, they are usually concatenated (in the feature dimension) and processed through an additional linear transformation (fully connected layer) to obtain the final output:
[0103] ;
[0104] where is the weight matrix of the output transformation, is the final output of the bidirectional cross-attention mechanism fusion network.
[0105] (2) Then, a dot product operation is performed on and to obtain ;
[0106] ;
[0107] (3) Finally, global fusion is performed, and the three features obtained previously , and then respectively apply them to the fully connected layer FC1(·), the fully connected layer FC2(·), and the fully connected layer FC3(·) to obtain the features Z1, Z2, and Z3 mapped to the 2D space; connect the features Z1, Z2, and Z3 mapped to the 2D space to obtain the final multi-modal feature .
[0108] ;
[0109] .
[0110] Among them, (·) is the activation function, , and are respectively the weight matrices corresponding to the three features , , and are respectively the bias terms corresponding to the three features .
[0111] Adopt the Transformer module to further extract the global temporal features of the multi-modal features, and finally obtain the recognition result of the preset dynamic emotion through the fully connected layer. The Transformer module includes positional encoding, multi-head self-attention mechanism, feed-forward neural network, and layer normalization. The positional encoding is used to retain the position information of the input sequence. The multi-head self-attention mechanism captures different subspaces of the input features through multiple attention heads. The feed-forward neural network performs feature transformation through two fully connected layers. The layer normalization is used to stabilize the training process and avoid gradient vanishing and explosion.
[0112] This application can preprocess online multi-modal data to obtain online multi-modal information, perform fusion classification processing based on the online multi-modal information, and determine the online mental state and online analysis result of the online multi-modal data.
[0113] Based on the same inventive concept, the embodiment of this application also provides a continuous dynamic emotion recognition system for fusing electroencephalogram and electrooculogram signals for implementing the above-mentioned continuous dynamic emotion recognition method for fusing electroencephalogram and electrooculogram signals. The implementation solution provided by this system to solve the problem is similar to the implementation solution recorded in the above method. Therefore, the specific limitations in one or more embodiments of the continuous dynamic emotion recognition system for fusing electroencephalogram and electrooculogram signals provided below can refer to the limitations on the continuous dynamic emotion recognition method for fusing electroencephalogram and electrooculogram signals in the above text, and will not be repeated here.
[0114] In an exemplary embodiment, a continuous dynamic emotion recognition system for fusing electroencephalogram and electrooculogram signals is provided, including:
[0115] A data acquisition module, configured to acquire electroencephalogram (EEG) signal data, electrooculogram (EOG) signal data, Big Five personality scores, and dynamic emotion labels of a target subject; the Big Five personality scores are obtained through a Big Five personality inventory; the dynamic emotion labels are obtained by the target subject scoring their emotional state from 1 to 9 every 5 seconds according to two dimensions of pleasure and arousal; the corresponding types of the scores are calm to happy, happy to calm, sad to calm, calm to sad, tense to calm, and calm to tense; the target subjects include pilots and ordinary subjects; the EOG signal data includes horizontal EOG signals and vertical EOG signals;
[0116] A preprocessing module, configured to preprocess the EEG signal data and the EOG signal data respectively to obtain preprocessed EEG signal data and preprocessed EOG signal data;
[0117] A subject classification module, configured to extract differential entropy features according to the preprocessed EEG signal data; and classify the target subjects into pilots and ordinary subjects according to the differential entropy features and the Big Five personality scores;
[0118] An emotion recognition module, configured to perform emotion recognition respectively on the preprocessed EEG signal data and the preprocessed EOG signal data of pilots and the preprocessed EEG signal data and the preprocessed EOG signal data of ordinary subjects by using a continuous dynamic emotion recognition model; the continuous dynamic emotion recognition model includes: a convolution module, a bidirectional cross-attention mechanism fusion network, a Transformer module, and a fully connected layer; the convolution module is configured to extract EEG signal features and EOG signal features from the preprocessed EEG signal data and the preprocessed EOG signal data by using spatio-temporal convolution technology and an SE attention mechanism network module; the bidirectional cross-attention mechanism fusion network is configured to fuse the EEG signal features and the EOG signal features to obtain multi-modal features; the Transformer module is configured to extract the temporal features of the multi-modal features during the emotion change process; the fully connected layer is configured to perform continuous dynamic emotion recognition according to the temporal features.
[0119] In an exemplary embodiment, a computer device is provided, which may be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a continuous dynamic emotion recognition method that fuses electroencephalogram and electrooculogram signals.
[0120] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0121] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0122] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0123] In the present application, all actions of obtaining signals, information, or data are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining the authorization given by the owner of the corresponding device.
[0124] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0125] In this article, specific examples are used to illustrate the principle and implementation of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. At the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A continuous dynamic emotion recognition method that fuses electroencephalogram and electrooculogram signals, characterized in that The continuous dynamic emotion recognition method that fuses electroencephalogram (EEG) and electrooculogram (EOG) signals includes: Obtaining the EEG signal data, EOG signal data, Big Five personality scores, and dynamic emotion labels of the target subjects; the Big Five personality scores are obtained through the Big Five personality inventory; the dynamic emotion labels are obtained by the target subjects scoring their emotional states from 1 to 9 every 5 seconds according to two dimensions of pleasure and arousal; the corresponding types of scores are calm to happy, happy to calm, sad to calm, calm to sad, tense to calm, and calm to tense; the target subjects include pilots and ordinary subjects; the EOG signal data includes horizontal EOG signals and vertical EOG signals; Preprocessing the EEG signal data and EOG signal data respectively to obtain the preprocessed EEG signal data and the preprocessed EOG signal data; Extracting differential entropy features based on the preprocessed EEG signal data; and classifying the target subjects into pilots and ordinary subjects according to the differential entropy features and Big Five personality scores; Performing emotion recognition respectively on the preprocessed EEG signal data and preprocessed EOG signal data of pilots and the preprocessed EEG signal data and preprocessed EOG signal data of ordinary subjects by using a continuous dynamic emotion recognition model; the continuous dynamic emotion recognition model includes: a convolutional module, a bidirectional cross-attention mechanism fusion network, a Transformer module, and a fully connected layer; the convolutional module is used to extract EEG signal features and EOG signal features from the preprocessed EEG signal data and the preprocessed EOG signal data by using spatio-temporal convolution technology and a SE attention mechanism network module; the bidirectional cross-attention mechanism fusion network is used to fuse the EEG signal features and the EOG signal features to obtain multi-modal features; the Transformer module is used to extract the temporal features of the multi-modal features during the emotion change process; the fully connected layer is used to perform continuous dynamic emotion recognition according to the temporal features.
2. The continuous dynamic emotion recognition method for fusing electroencephalogram and electrooculogram signals according to claim 1, wherein The step of preprocessing the EEG signal data and EOG signal data respectively to obtain the preprocessed EEG signal data and the preprocessed EOG signal data specifically includes: Performing filtering processing, artifact removal processing, and resampling processing on the EEG signal data in sequence to obtain the preprocessed EEG signal data; Performing filtering processing and resampling processing on the EOG signal data in sequence to obtain the preprocessed EOG signal data.
3. The continuous dynamic emotion recognition method for fusing EEG and EOG signals according to claim 1, characterized in that, The step of extracting differential entropy features based on the preprocessed EEG signal data specifically includes: Use the formula to determine the differential entropy feature; wherein, is the differential entropy feature of the electroencephalogram signal data, is the total number of samples, is the value of the electroencephalogram signal data at the nth sampling point, is the value of the electroencephalogram signal data at the (n + 1)th sampling point.
4. The continuous dynamic emotion recognition method for fusing EEG and EOG signals according to claim 1 or claim 3, characterized in that, The frequency bands of the differential entropy features include: 1Hz - 3Hz frequency band, 4Hz - 7Hz frequency band, 8Hz - 13Hz frequency band, 14Hz - 30Hz frequency band, and 31Hz - 50Hz frequency band.
5. The continuous dynamic emotion recognition method for fusing electroencephalogram and electrooculogram signals according to claim 1, wherein Before classifying the target subjects into pilots and ordinary subjects according to the differential entropy features and Big Five personality scores, it also includes: Performing standardization processing on the differential entropy features and Big Five personality scores.
6. The continuous dynamic emotion recognition method for fusing electroencephalogram and electrooculogram signals according to claim 1, wherein The convolutional module includes: the first-layer convolution, the second-layer convolution, the third-layer convolution, the fourth-layer convolution, the batch normalization layer between the first-layer convolution and the second-layer convolution, the batch normalization layer between the second-layer convolution and the third-layer convolution, the dropout layer after the third-layer convolution, and the SE attention mechanism network module between the third-layer convolution and the fourth-layer convolution; The first-layer convolution uses 1D convolution to remove noise and extract features; The activation function is excluded in the first-layer convolution; the second-layer convolution includes a spatial convolution layer; the third-layer convolution includes an average pooling layer; the SE attention mechanism network module adaptively learns the weights between channels; the fourth-layer convolution uses an advanced feature fusion layer.
7. A continuous dynamic emotion recognition system that fuses electroencephalogram and electrooculogram signals, characterized in that, The continuous dynamic emotion recognition system that fuses EEG and EOG signals includes: A data acquisition module, which is used to acquire the EEG signal data, EOG signal data, Big Five personality scores, and dynamic emotion labels of the target subject; the Big Five personality scores are obtained through the Big Five personality inventory; the dynamic emotion labels are scored by the target subject from 1 to 9 for the emotional state every 5 seconds according to two dimensions of pleasure and arousal; the corresponding types of the scores are calm to happy, happy to calm, sad to calm, calm to sad, tense to calm, and calm to tense; the target subjects include pilots and ordinary subjects; the EOG signal data includes horizontal EOG signals and vertical EOG signals; A preprocessing module, which is used to preprocess the EEG signal data and the EOG signal data respectively to obtain the preprocessed EEG signal data and the preprocessed EOG signal data; A subject division module, which is used to extract differential entropy features according to the preprocessed EEG signal data; and divide the target subjects into pilots and ordinary subjects according to the differential entropy features and the Big Five personality scores; An emotion recognition module, which is used to perform emotion recognition respectively on the preprocessed EEG signal data and the preprocessed EOG signal data of the pilots and the preprocessed EEG signal data and the preprocessed EOG signal data of the ordinary subjects by using a continuous dynamic emotion recognition model; the continuous dynamic emotion recognition model includes: a convolutional module, a bidirectional cross-attention mechanism fusion network, a Transformer module, and a fully connected layer; the convolutional module is used to extract EEG signal features and EOG signal features from the preprocessed EEG signal data and the preprocessed EOG signal data by using spatio-temporal convolution technology and the SE attention mechanism network module; the bidirectional cross-attention mechanism fusion network is used to fuse the EEG signal features and the EOG signal features to obtain multimodal features; the Transformer module is used to extract the temporal features in the process of emotion change of the multimodal features; the fully connected layer is used to perform continuous dynamic emotion recognition according to the temporal features.
Citation Information
Patent Citations
Recognition method and recognition system based on electroencephalogram and eye movement
CN116439706A
Multi-modal emotion recognition model training method and system and electronic equipment
CN117171626A
Cited By
Emotion monitoring system and method based on brain-computer interface
CN121101570A