A multimodal emotion recognition and awareness detection method based on electroencephalogram and micro-expression
A method combining EEG and micro-expression analysis using spatial and temporal attention mechanisms addresses the inaccuracy and complexity of current consciousness disorder diagnostics, enhancing detection accuracy and enabling personalized treatment.
Patent Information
- Application Number
- CN202410970432.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-07-19
AI Technical Summary
The prior art has a high misdiagnosis rate and lack of sensitivity in mood recognition and awareness detection in patients with awareness disorders. Traditional methods such as motor response assessment and fMRI devices are complex and expensive, and cannot accurately evaluate the state of consciousness.
Multimodal emotion recognition method based on EEG and micro-expression is adopted, and EEG signals and micro-expression features are fused through feature extraction and SwinTransformer model, feature fusion is used using the spatiotemporal attention mechanism, and emotion recognition and consciousness detection are combined with deep learning technology.
It improves the accuracy and comprehensiveness of emotional state detection in patients with awareness disorders, provides more accurate assessment of consciousness state, assists in clinical diagnosis and treatment decisions, and reduces the rate of misdiagnosis.
Smart Images

Figure CN118902458B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of artificial intelligence, brain-computer interface, and emotion recognition, and particularly relates to a multi-modal emotion recognition and consciousness detection method based on electroencephalogram and micro-expression. Background Art
[0002] Emotion recognition is to understand the emotional state of a person by analyzing their physiological behaviors and activities, and it involves fields such as psychology and medicine. This technology can not only help understand an individual's emotional experience, but also become an important tool for evaluating the state of consciousness. Consciousness detection is a technology for accurately monitoring and evaluating the consciousness state of patients with consciousness disorders, which is related to the selection of treatment methods and the judgment of prognosis for patients. Accurately evaluating the consciousness state is particularly important for such patients. In recent years, extensive research has been conducted on expressing emotions through facial expressions, gestures, and language, etc. Physiological signals such as electromyogram, electrocardiogram, and electroencephalogram are also objective manifestations that emotions are difficult to hide. The changes in facial muscles can reflect the changes in emotions, and the changes in the electrical signals of normal brain activities can also reflect the changes in various physiological electrical signals generated by brain nerve cells in the cerebral cortex, and then analyze the changes in emotions.
[0003] Currently, in the field of emotion recognition, researchers mainly use single-modal data, such as electroencephalogram (EEG), electrocardiogram, electromyogram, or facial expressions, gestures, and language, etc. for emotion recognition. However, there are limitations of single-modal methods in detecting emotions in patients with disorders of consciousness. These methods cannot capture the complex interactions of emotional states, resulting in inaccurate and incomplete assessment results, and the accuracy rate of assessment results is not high. Therefore, more accurate and comprehensive assessment tools are needed. In the field of disorders of consciousness, there are challenges in diagnosing the consciousness state of patients with disorders of consciousness, especially in three types of disorders of consciousness: coma, minimally conscious state (MCS), and unresponsive wakefulness syndrome (UWS). The challenging problems in diagnosis are becoming increasingly prominent. The methods for detecting the consciousness of patients with disorders of consciousness mainly include clinical assessment methods based on motor responses and functional magnetic resonance imaging (fMRI) technology. These methods also have some drawbacks. Traditional clinical assessment methods, such as the revised coma recovery scale (CRS-R) based on motor responses. This clinical assessment method has a high misdiagnosis rate and lacks sensitivity. There is a misdiagnosis rate of 37% - 43% in diagnosing the consciousness state, which may lead to some patients not obtaining an accurate assessment of the consciousness state and appropriate treatment. Therefore, more accurate and comprehensive assessment tools are needed. Functional magnetic resonance imaging (fMRI) can show the brain activity patterns of patients with disorders of consciousness during specific activity tasks, but it is expensive and complex to operate, requires professional operation and analysis, and its application in the clinical environment is limited, and the technical requirements for operators and analysts are relatively high. To reduce the misdiagnosis rate and improve the accuracy of assessment, more simple and effective methods need to be introduced to supplement the assessment of the consciousness state, making it simple enough to implement in the clinical environment and capable of providing more accurate consciousness state information.
[0004] To reduce the risk of misdiagnosis and improve the accuracy of assessment, more simplified and sensitive methods need to be introduced to supplement the assessment of consciousness. For patients with unresponsive wakefulness syndrome (UWS), they have no awareness of the surrounding environment and themselves, while patients in the minimally conscious state (MCS) have weak consciousness. Therefore, it can be hypothesized that patients with emotional changes in the passive paradigm have better chances of recovery. In recent years, due to the rapid development of computer vision and brain-computer interfaces (BCIs), the accuracy of emotion recognition based on facial images and electroencephalogram has improved. Compared with assessment methods based on motor responses and fMRI, multi-modal emotion recognition based on electroencephalogram and micro-expression can more simply diagnose the residual consciousness state of patients with disorders of consciousness and serve as an auxiliary diagnostic tool. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi-modal emotion recognition and consciousness detection method based on electroencephalogram and micro-expression to solve the problems of high misdiagnosis rate and lack of sensitivity in traditional diagnostic methods in the field of disorders of consciousness mentioned in the above background technology.
[0006] To achieve the above object, the present invention provides the following technical solutions: A multi-modal emotion recognition and consciousness detection method based on electroencephalogram and micro-expression, specifically including the following steps:
[0007] Step 1: Feature extraction, pay attention to the correlation of multi-channel electroencephalogram signals and extract micro-expression features, and perform feature fusion on electroencephalogram and micro-expression through the feature layer to form multi-modal features;
[0008] For electroencephalogram features, the processing steps are as follows:
[0009] a. Preprocess the original electroencephalogram signals, including removing artifacts, filtering, and component removal;
[0010] b. After preprocessing, segment the electroencephalogram data of each channel into non-overlapping 1-second intervals according to the output size to extract electroencephalogram features;
[0011] Step 2: Build an emotion recognition model, select SwinTransformer as the backbone network of the model. When customizing SwinTransformer, optimize the patch size and network depth of the model to make them consistent with the dimensions of the feature tensor; ensure that the STST model meets the requirements of the emotion recognition task and can fully exert the potential of the SwinTransformer architecture;
[0012] Step 3: Emotion comparison metrics, use accuracy as the benchmark, compare the model performance with other existing methods, and provide a quantitative measurement standard for the efficacy in emotion classification;
[0013] Step 4: Experimental dataset, on the public dataset, detect the emotion recognition accuracy on the multi-modal dataset MAHNOB-HCI;
[0014] Step 5: Experimental data preprocessing, use the public dataset and self-collected dataset, and adopt a consistent data preprocessing process for these two datasets to ensure the comparability and robustness of the results;
[0015] Step 6: Experimental parameter settings, extract DE as the input of the electroencephalogram modality, and extract continuous sequences of face images as the input of the micro-expression modality; their input sizes are 56×56×5 respectively; the electroencephalogram modality represents 5 electroencephalogram frequency bands, and the micro-expression modality represents 5 face images that change continuously and uniformly within 0.2 seconds; verify the feasibility of the method on the public dataset, then pre-train on the data of healthy subjects, and fine-tune on the data of MCS patients; finally, calculate the evaluation metrics on the data of UWS patients.
[0016] As a preferred technical solution in the present invention, in step one, the EEG signal after passing through the band-pass filter follows a Gaussian distribution N(μ, σ 2 ), and the DE feature of the electroencephalogram is calculated as follows:
[0017]
[0018] where f(x) is the probability density function of x, and σ 2 is the variance of this electroencephalogram; the DE feature belongs to the frequency domain feature. The EEG signals usually used for emotion recognition are divided into five different frequency bands: δ(0.1 - 4Hz), θ(4 - 8Hz), α(8 - 12Hz), β(12 - 30Hz), and γ(30 - 45Hz); in this method, the DE features of each frequency band are calculated respectively; the final feature vector of each sample contains 56 channels (or 28 channels × 2) × 56 DE features × 5 frequency bands to ensure matching the size of the micro-expression features.
[0019] As a preferred technical solution in the present invention, in the said step one, for the multi-modal feature fusion of electroencephalogram and micro-expression, a spatio-temporal attention mechanism multi-modal feature fusion architecture for electroencephalogram and micro-expression data is used, aiming to separate the emotionally significant features from the multi-modal data stream;
[0020] In the time attention module, the input micro-expression feature X me is initially processed through a 1×1 convolutional layer (denoted as W v ), generating a transformed feature; subsequently, these features are reshaped to a dimension of C me ×2HW, where H and W represent the height and width of the feature map respectively; another 1×1 convolutional layer W q is used to generate the query feature, and then the attention distribution is calculated through the softmax layer; this attention distribution map is multiplied by the transposed feature elements to obtain the time attention feature map Z ta ; the time attention mechanism aims to capture the time information in the micro-expression image sequence; it uses the attention weights to capture the changes of the same spatial position over time to describe the features of the micro-expression; the formula representing this process is as follows:
[0021]
[0022] where, W q , W v and W z are 1×1 convolutional layers, σ1 and σ2 are two tensor reshaping operators, is the matrix dot product operation, ⊙ is the channel multiplication; F sm is the Softmax operator; includes Wz Convolution operation, followed by layer regularization and Sigmoid operator;
[0023] Meanwhile, the spatial attention mechanism processes the input EEG feature X in a similar way eeg ; this mechanism calculates the spatial attention distribution through global pooling and reshaping operations; subsequently, this distribution generates the spatial attention feature Z through the softmax and sigmoid layers sa ; the spatial attention mechanism is good at capturing the spatial attention information of different channels within the same frequency band of EEG signals; the formula representing this process is as follows:
[0024]
[0025] where, W q and W v are 1×1 convolutional layers, σ1 and σ2 are two tensor reshaping operators, is the matrix dot product operation, ⊙ is the channel multiplication, and G is the global pooling operator; F sm is the Softmax operator; σ s consists of a tensor reshaping operator and a Sigmoid operator.
[0026] As a preferred technical solution in the present invention, in the fusion model of EEG and micro-expression features, the attention mechanism is integrated to adapt to the single-modal input X;
[0027] In the model, first, the EEG and micro-expression features are combined; to strengthen the feature representation, multiple layers of single-modal inputs are stacked onto the spatio-temporal attention mechanism; in the domain coordination mechanism, the temporal attention feature Z ta and the spatial attention feature Z sa are non-linearly transformed through the PReLU activation function; PReLU adjusts the slope of the negative part of the input by introducing a learnable parameter, thereby enhancing the model's ability to learn complex patterns from the data; then, the respective results are subjected to channel multiplication operations with the added fusion features, and the sum result is processed through the Sigmoid function to output the final fusion feature Z dc ; the calculation formula of the domain coordination mechanism is as follows:
[0028] Z dc_1 = x⊙F sg ((F prl (Z ta )+F prl (Z sa ))⊙(Z ta +Z sa ))
[0029] Z dc_2 = Fsg ((F prl (Z ta ) + F prl (Z sa )) ⊙ (Z ta + Z sa ))
[0030] Among them, Z dc_1 and Z dc_2 are the outputs obtained when inputting two dimensions of multi-modal data or electroencephalogram and micro-expression respectively; F sg represents the sigmoid operator, and F prl represents the PReLU operator;
[0031] The feed-forward network design used in the spatio-temporal attention module is based on Swin Transformer, which is suitable for processing electroencephalogram signal time data and micro-expression spatial data, can capture local spatial features in micro-expression data, and adapt to the time features of electroencephalogram signals, so as to achieve multi-modal data fusion; in addition, the sliding window mechanism changes the calculation method of self-attention.
[0032] As a preferred technical solution in the present invention, in the second step, the input of the backbone network is a multi-modal feature tensor with a size of 56×56×n. Then, after dividing these tensors, linear embedding is performed to convert the features into a size suitable for STST input; each module consists of a multi-layer self-attention mechanism and a multi-layer perceptron (MLP). Each module will perform layer normalization and then apply residual connections; the self-attention within the block runs on local windows that move between layers, enabling the model to effectively capture local and global context information.
[0033] As a preferred technical solution in the present invention, in the third step, for the self-collected dataset, first, the data of 10 healthy subjects are used to evaluate the accuracy of the model, and this step establishes a baseline for the expected emotional response; subsequently, the model is applied to the emotional classification of 21 patients collected; the emotional classification probability of the model is mapped to a continuous emotional score interval from -1 to 1, where -1 represents negative, 0 represents neutral, and 1 represents positive; in order to calculate the overall emotional state score, a weighted method is adopted, which includes multiplying the predicted probability of each category by its related label value to obtain a continuous score; the formula representing this process is as follows:
[0034] Score = C neg × C neg + C neu × P neu + C pos × P pos
[0035] Among them, C neg, C neu and C pos represent the contribution values related to negative, neutral, and positive emotional states respectively, with the assigned values of -1, 0, and 1; P neg , P neu and P pos represent the probabilities of negative, neutral, and positive emotions predicted by the deep learning model respectively; this formula aims to convert the probability output of emotion classification into a continuous score from -1 to 1, providing a more accurate comprehensive assessment of emotional tendency; to evaluate the output results of the model, Euclidean distance and cosine similarity are used as the metrics for distance calculation and similarity assessment respectively; these metrics measure the similarity between the emotional patterns of patients with disorders of consciousness and those of healthy people;
[0036] Accuracy is used to evaluate the classification performance of emotion recognition; accuracy is a general measure of the model's ability to correctly classify positive and negative instances in different emotion categories, and the definition of the metric is as follows:
[0037]
[0038] where TP represents true positive, TN represents true negative, FP represents false positive, and FN represents false negative; the accuracy score allows quantifying the overall classification accuracy of the model on the entire dataset, providing a global performance metric; when performing emotion recognition on patients with disorders of consciousness, the emotion labels of patients with disorders of consciousness cannot be directly used; it cannot be ensured that the corresponding emotions have been induced in the patients, so traditional accuracy assessment cannot be carried out, and Euclidean distance and cosine similarity are used as the measurement criteria;
[0039] First, Euclidean distance is used to measure the spatial distance between emotional expressions; for each sample, its emotional score is represented as a vector, and the Euclidean distance is calculated to measure the difference between samples; the smaller the Euclidean distance, the closer the emotional intensity is;
[0040] Second, cosine similarity is used to measure the angular relationship between emotional vectors; specifically, the closer the cosine similarity value is to 1, the more similar the emotional change trends between the two are; conversely, the closer this value is to -1, the more opposite the emotional change trends between the two are; for the emotional score vector A=(a1, a2,..., a n ) of healthy subjects during the playback of induced videos and the emotional score vector B=(b1, b2,..., b n ) of patients with disorders of consciousness, their Euclidean distance d and cosine similarity s are calculated according to the following formula:
[0041]
[0042] The Euclidean distance is the square root of the sum of the squared differences of the corresponding components of two vectors, and the value range of the Euclidean distance is [0, 2]; the value range of the cosine similarity is [-1, 1]. The closer the value is to 1, the more similar the two vectors are, and the closer the value is to -1, the more opposite the two vectors are.
[0043] As a preferred technical solution in the present invention, in step four, the data in the self-collected dataset includes electroencephalogram (EEG) signals and facial expression images; in order to collect high-quality scalp EEG signals, NuAmps amplifiers and an EEG cap were used to record the scalp EEG signals of 32 electrodes, and the impedance of all 32 electrodes was maintained below 30 kΩ; the EEG signals were collected at a sampling rate of 1000 Hz and then band-pass filtered (0.1 - 50 Hz); the facial expression images were exported from the videos recorded by the camera at 1080p 30fps; the subjects were in a controlled environment with minimal external stimuli; the experimental design included a passive paradigm of watching videos to induce emotions, including image and audio clips.
[0044] As a preferred technical solution in the present invention, in step five, when preprocessing the facial images, several steps were taken to ensure the quality and adaptability of the data; the face images were loaded from the dataset and grayscale images were generated to reduce the computational complexity; histogram equalization was applied to enhance the contrast of the images, thereby improving the image quality, and then the MTCNN model was used to detect the faces in the images and align the detected faces to ensure that they had similar positions and sizes; image normalization processing was adopted to make the mean of the pixel values 0 and the variance 1.
[0045] For the preprocessing of EEG signals:
[0046] First, the signals were filtered using a 0.1 to 60 Hz band-pass filter and a 50 Hz notch filter to remove noise; the variance of the leads was also checked to see if it was too large, if it contained too many zeros, and whether the data acquisition of each lead was normal. If there were leads with abnormal acquisition, the mean value of the channels around the bad lead was used for interpolation; independent component analysis (ICA) was used to identify and remove the noise in the electrooculogram and electromyogram signals; the EEG signals were also normalized to ensure the consistency of the amplitudes of each channel.
[0047] As a preferred technical solution in the present invention, in the experiment of step six, the training set and the test set were divided in a ratio of 4:1, and a five-fold cross-validation strategy was used to train the emotion recognition model; the model was implemented using the PyTorch framework and trained using the AdamW optimizer with a learning rate of 3e-4, and the learning rate decay used the cosine annealing method; the batch size used for model training was 32, and the model was trained 100 times.
[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0049] Emotion recognition is performed through multi-modal fusion features of electroencephalogram (EEG) and micro-expression, and the distance is used to measure the gap in emotion expression between patients with disorders of consciousness and normal people, so as to distinguish the degree of the quality of the state of consciousness, achieving the purpose of assisting in the detection of the state of consciousness. For multi-modal emotion recognition, the present invention proposes a multi-modal feature fusion method, which fuses EEG and micro-expression data through an innovative spatio-temporal attention mechanism, and introduces emotion recognition technology into the assessment of the state of consciousness. This method uses deep learning technology to fuse the differential entropy feature of EEG and the time-domain feature of micro-expression, significantly improving the detection and classification ability of the emotional state. The multi-modal method combining EEG and micro-expression provides a more comprehensive and in-depth understanding of the emotional state of patients with disorders of consciousness. EEG is a direct reflection of brain activity, while micro-expression provides a more comprehensive and accurate description of the emotional state of patients with disorders of consciousness through minute facial change information. This method not only reaches the level of mainstream models in terms of accuracy, but also demonstrates broad application potential in medical diagnosis, provides a basis for personalized patient care, and opens up new avenues for clinical practice research. The results in practical applications show that for patients whose detected emotional fluctuations are more similar to those of normal people, they have a higher level of brain activity, indicating better treatment outcomes, which is consistent with the clinical results one month later. By effectively combining data from multiple modalities, this method has established a new paradigm for emotion recognition in the assisted detection of disorders of consciousness, contributing to more accurate clinical assessment, timely adjustment of the treatment plan for patients, and improvement of the quality of life and treatment outcomes of patients.
[0050] (1) In the present invention, an innovative multi-modal fusion framework is proposed, which combines EEG and micro-expression data to achieve the detection and classification of emotional states through deep learning technology. The key to multi-modal fusion in the framework is the proposal of a spatio-temporal attention module for multi-modal fusion, which can independently process the spatial and temporal data from electroencephalogram and micro-expression, and fuse EEG and micro-expression data at the feature level through a spatio-temporal attention mechanism based on polarization attention and domain fusion, improving the accuracy of emotion recognition.
[0051] (2) This method combines the low-dimensional artificial features and high-dimensional deep learning features of electroencephalogram and micro-expression, improving the accuracy of emotion recognition. The low-dimensional artificial features extract more relevant information through the spatio-temporal attention module and generate high-dimensional features through the neural network. This hybrid method can combine specific low-dimensional features and complex high-dimensional features to improve the accuracy of emotion recognition.
[0052] (3) The present invention proposes a new multi-modal method for evaluating the state of consciousness by combining EEG and micro-expressions, providing a more comprehensive and accurate description of the emotional state of patients with disorders of consciousness, and providing a basis for personalized medical and nursing decisions for healthcare professionals. Through multi-modal feature fusion and deep learning methods, an evaluation index for evaluating the emotional expression of patients with disorders of consciousness is established, providing a new quantitative means for evaluating the state of consciousness, and also providing a new method and concept for introducing multi-modal emotion recognition into the clinical practice of the field of consciousness detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a framework diagram of the multi-modal emotion recognition and consciousness detection method based on electroencephalogram and micro-expression of the present invention;
[0054] Figure 2 is a multi-modal feature fusion architecture diagram of the present invention for using spatio-temporal attention mechanism for electroencephalogram and micro-expression data;
[0055] Figure 3 is a feature fusion architecture diagram of the present invention for using spatio-temporal attention mechanism for multi-modal features. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0057] Please refer to Figures 1 to 3 , the present invention provides a technical solution: a multi-modal emotion recognition and consciousness detection method based on electroencephalogram and micro-expression, specifically including the following steps:
[0058] Step 1: Feature extraction, paying attention to the correlation of multi-channel electroencephalogram signals and extracting micro-expression features, and performing feature fusion on electroencephalogram and micro-expression through the feature layer to form multi-modal features;
[0059] For electroencephalogram features, the processing steps are as follows:
[0060] a. Preprocess the original electroencephalogram signal, including removing artifacts, filtering, and component removal;
[0061] b. After preprocessing, segment the electroencephalogram data of each channel into non-overlapping 1-second intervals according to the output size to extract electroencephalogram features. Previous studies have confirmed that in the emotion recognition task, the differential entropy (DE) has better effects than other artificial features;
[0062] Step 2: Construct an emotion recognition model. Select Swin Transformer as the backbone network of the model because it demonstrates excellent performance, adaptability, and scalability in various vision-related tasks when dealing with multimodal data. It innovatively uses shifted windows to simulate local and global interactions and has obvious advantages in handling the complexity of electroencephalogram (EEG) features and micro-expression data, which is crucial in the research. When customizing Swin Transformer, optimize the patch size and network depth of the model to make them consistent with the dimensions of the feature tensor, ensuring that the STST model meets the requirements of the emotion recognition task and can fully exploit the potential of the Swin Transformer architecture.
[0063] Step 3: Emotion comparison metrics. Take accuracy as the benchmark and compare the model performance with other existing methods, providing a quantitative measure of its efficacy in emotion classification.
[0064] Step 4: Experimental dataset. On a public dataset, detect the emotion recognition accuracy on the multimodal dataset MAHNOB-HCI, which is a dataset focused on multimodal emotion recognition. This dataset consists of 30 healthy subjects and includes head-mounted camera videos of different subjects and corresponding 32-channel EEG signals. This dataset provides detailed facial expression and EEG data, enabling in-depth exploration of multimodal emotion recognition.
[0065] Step 5: Experimental data preprocessing. Use both public datasets and self-collected datasets and adopt a consistent data preprocessing process for these two types of datasets to ensure the comparability and robustness of the results.
[0066] Step 6: Experimental parameter settings. Extract DE as the input of the EEG modality and extract a continuous sequence of face images as the input of the micro-expression modality. Their input sizes are 56×56×5 respectively. The EEG modality represents 5 EEG frequency bands, and the micro-expression modality represents 5 face images that change continuously and uniformly within 0.2 seconds. Verify the feasibility of the method on the public dataset, then pre-train on the data of healthy subjects and fine-tune on the data of MCS patients. Finally, calculate the evaluation metrics on the data of UWS patients.
[0067] In this embodiment, in Step 1, the EEG signal after passing through the band-pass filter follows a Gaussian distribution N(μ, σ 2 ), and the DE feature of the EEG is calculated as follows:
[0068]
[0069] where f(x) is the probability density function of x, and σ 2is the variance of this electroencephalogram; DE features belong to frequency-domain features. Electroencephalogram signals commonly used for emotion recognition are divided into five different frequency bands: δ (0.1 - 4 Hz), θ (4 - 8 Hz), α (8 - 12 Hz), β (12 - 30 Hz), and γ (30 - 45 Hz); in the method, the DE features of each frequency band are calculated respectively; the final feature vector of each sample contains 56 channels (or 28 channels × 2) × 56 DE features × 5 frequency bands to ensure matching the size of micro-expression features;
[0070] Micro-expression is a transient and involuntary facial expression, usually lasting from 1 / 25 second to 1 / 5 second; for micro-expression features, a new method for extracting micro-expression features is adopted; after exporting the original images of video frames, the MTCNN model is used for face detection and alignment; to ensure efficient processing, five facial images are selected with a time window of 1 / 5 second, evenly distributed throughout the duration of the micro-expression, so as to capture the basic dynamics of facial expressions within the limited time of the micro-expression; each image is adjusted to a consistent size of 56×56 pixels and then stacked to form a 5×56×56 feature tensor; this tensor not only reduces the computational requirements but also retains the key details required for effective emotion recognition; by generating micro-expression features in the feature fusion stage, it is ensured that the key information about micro-expressions is encapsulated, which helps to improve the accuracy and computational efficiency of the analysis.
[0071] In this embodiment, in step one, for the multi-modal feature fusion of electroencephalogram and micro-expression, a spatio-temporal attention mechanism multi-modal feature fusion architecture for electroencephalogram and micro-expression data is used, aiming to separate the emotionally significant features from the multi-modal data stream;
[0072] In the time attention module, the input micro-expression feature X me is initially processed through a 1×1 convolutional layer (denoted as W v ) to generate a transformed feature; subsequently, these features are reshaped to a dimension of C me ×2HW, where H and W represent the height and width of the feature map respectively; another 1×1 convolutional layer W q is used to generate query features, and then the attention distribution is calculated through a softmax layer; this attention distribution map is multiplied by the transposed feature elements to obtain the time attention feature map Z ta ; the time attention mechanism aims to capture the time information in the micro-expression image sequence; it uses attention weights to capture the changes over time at the same spatial position to describe the features of micro-expressions; the formula representing this process is as follows:
[0073]
[0074] where, W q 、Wv and W z are 1×1 convolutional layers, σ1 and σ2 are two tensor reshaping operators, is the matrix dot product operation, ⊙ is the channel multiplication; F sm is the Softmax operator; includes W z convolution operation, followed by layer regularization and Sigmoid operator;
[0075] Meanwhile, the spatial attention mechanism processes the input EEG feature X in a similar way eeg ; this mechanism calculates the spatial attention distribution through global pooling and reshaping operations; subsequently, this distribution generates the spatial attention feature Z through the softmax and sigmoid layers sa ; the spatial attention mechanism is good at capturing the spatial attention information of different channels within the same frequency band of the EEG signal; the formula representing this process is as follows:
[0076]
[0077] where, W q and W v are 1×1 convolutional layers, σ1 and σ2 are two tensor reshaping operators, is the matrix dot product operation, ⊙ is the channel multiplication, G is the global pooling operator; F sm is the Softmax operator; σ s consists of a tensor reshaping operator and a Sigmoid operator.
[0078] In this embodiment, in the fusion model of EEG and micro-expression features, the attention mechanism is integrated to adapt to the single-modal input X;
[0079] In the model, first the EEG and micro-expression features are combined; to strengthen the feature representation, multiple layers of single-modal inputs are stacked on the spatio-temporal attention mechanism; in the domain coordination mechanism, the temporal attention feature Z ta and the spatial attention feature Z sa are non-linearly transformed through the PReLU activation function; PReLU adjusts the slope of the negative part of the input by introducing a learnable parameter, thus enhancing the model's ability to learn complex patterns from the data; then, the respective results are subjected to channel multiplication operations with the added fusion features, and the sum result is processed through the Sigmoid function to output the final fusion feature Z dc ; the calculation formula of the domain coordination mechanism is as follows:
[0080] Z dc_1 = x⊙F sg ((F prl (Z ta)+F prl (Z sa ))⊙(Z ta +Z sa ))
[0081] Z dc_2 =F sg ((F prl (Z ta )+F prl (Z sa ))⊙(Z ta +Z sa ))
[0082] Among them, Z dc_1 and Z dc_2 are the outputs obtained when inputting multi-modal data or two dimensions of electroencephalogram and micro-expression respectively; F sg represents the sigmoid operator, and F prl represents the PReLU operator;
[0083] The feed-forward network design used in the spatio-temporal attention module is based on Swin Transformer, which is suitable for processing electroencephalogram signal time data and micro-expression spatial data, can capture local spatial features in micro-expression data, and adapt to the time features of electroencephalogram signals, so as to realize multi-modal data fusion; In addition, the sliding window mechanism changes the calculation method of self-attention and improves the calculation efficiency, which is particularly important when processing multi-modal data containing complex features; Using this method can enable the spatio-temporal attention mechanism to focus on the spatial and time information in electroencephalogram and micro-expression signals respectively, and can capture the interaction and correlation between different modalities; It can not only effectively extract the shared features between different modalities, but also extract the features that can best express the emotional state, thus improving the accuracy of emotion recognition.
[0084] In this embodiment, in step two, the input of the backbone network is a multi-modal feature tensor with a size of 56×56×n. Then, these tensors are segmented and linearly embedded to convert the features into a size suitable for the input of STST. Each module consists of multiple layers of self-attention mechanisms and multi-layer perceptrons (MLPs). Each module performs layer normalization and then applies residual connections. The self-attention within the block operates on local windows that move between layers, enabling the model to effectively capture local and global context information. This hierarchical processing enables the network to effectively integrate spatio-temporal information from facial images and electroencephalogram data, thereby generating a semantically rich feature set with high discriminative ability for emotion recognition. The MLP integrates the features into probability scores corresponding to positive, neutral, and negative emotion categories respectively. This end-to-end trainable system ensures that the classification process can utilize both spatio-temporal patterns and refined features obtained from the Swin Transformer module, ultimately forming a high-accuracy and robust emotion recognition model.
[0085] In this embodiment, in step three, for the self-collected dataset, first, the data of 10 healthy subjects are used to evaluate the accuracy of the model, which establishes a baseline for the expected emotional responses. Subsequently, the model is applied to the emotion classification of 21 patients collected. The emotion classification probability of the model is mapped to a continuous emotion score range from -1 to 1, where -1 represents negative, 0 represents neutral, and 1 represents positive. To calculate the overall emotional state score, a weighted method is adopted, which includes multiplying the predicted probability of each category by its relevant label value to obtain a continuous score. The formula representing this process is as follows:
[0086] Score = C neg ×P neg +C neu ×P neu +C pos ×P pos
[0087] where C neg , C neu and C pos represent the contribution values related to negative, neutral, and positive emotional states respectively, with the assigned values of -1, 0, and 1; P neg , P neu and P pos represent the probabilities of negative, neutral, and positive emotions predicted by the deep learning model respectively. This formula aims to convert the probability output of emotion classification into a continuous score from -1 to 1, providing a more accurate comprehensive assessment of emotional tendencies. To evaluate the output results of the model, Euclidean distance and cosine similarity are used as indicators for distance calculation and similarity evaluation respectively. These indicators measure the similarity between the emotional patterns of patients with disorders of consciousness and those of healthy people.
[0088] The accuracy rate is used to evaluate the classification performance of emotion recognition; accuracy is a general measure of the model's ability to correctly classify positive and negative instances in different emotion categories, and the definition of the metric is as follows:
[0089]
[0090] Among them, TP represents true positive, TN represents true negative, FP represents false positive, and FN represents false negative; the accuracy score allows quantifying the overall classification accuracy of the model on the entire dataset, providing a global performance metric; when performing emotion recognition on patients with disorders of consciousness, the emotion labels of patients with disorders of consciousness cannot be directly used; it is impossible to ensure that the corresponding emotions have been induced in the patients, so traditional accuracy assessment cannot be carried out, and the Euclidean distance and cosine similarity are used as the measurement criteria; these two metrics can quantify the similarity of emotional expressions between healthy subjects and patients with disorders of consciousness, which provides a feasible performance evaluation method;
[0091] First, the Euclidean distance is used to measure the spatial distance between emotional expressions; for each sample, its emotion score is represented as a vector, and the Euclidean distance is calculated to measure the difference between samples; the smaller the Euclidean distance, the closer the emotional intensities are.
[0092] Second, the cosine similarity is used to measure the angular relationship between emotional vectors; specifically, the closer the cosine similarity value is to 1, the more similar the emotional change trends between the two are; conversely, the closer the value is to -1, the more opposite the emotional change trends between the two are; for the emotion score vector A=(a1, a2,..., a n ) of healthy subjects during the playback of the induced video and the emotion score vector B=(b1, b2,..., b n ) of patients with disorders of consciousness, their Euclidean distance d and cosine similarity s are calculated according to the following formulas:
[0093]
[0094] The Euclidean distance is the square root of the sum of the squared differences of the corresponding components of two vectors, and the value range of the Euclidean distance is [0,2]; the value range of the cosine similarity is [-1,1], and the closer the value is to 1, the more similar the two vectors are, and the closer the value is to -1, the more opposite the two vectors are.
[0095] In this embodiment, in step four, the data in the self - collected dataset includes electroencephalogram (EEG) signals and facial expression images; To collect high - quality scalp EEG signals, a NuAmps amplifier and an EEG cap were used to record scalp EEG signals of 32 electrodes, and the impedance of all 32 electrodes was maintained below 30 kΩ; The EEG signals were collected at a sampling rate of 1000 Hz and then band - pass filtered (0.1 - 50 Hz); The facial expression images were exported from the video recorded by the camera at 1080p 30fps; The subjects were in a controlled environment with minimal external stimuli; The experimental design included a passive paradigm of inducing emotions by watching videos, including image and audio segments; There were 31 subjects in this dataset, and the research objects included 10 healthy subjects from South China Normal University and 21 patients with disorders of consciousness from Zhujiang Hospital of Southern Medical University (including 13 MCS patients and 8 UWS patients); Each participant or their family member signed an informed consent form before participating in the study to ensure that they understood the nature, purpose, and potential risks of the experiment; This study was approved by the hospital ethics committee, and all experimental procedures met ethical standards;
[0096] Table 1 Comparison of emotion recognition accuracies on the MAHNOB - HCI dataset
[0097]
[0098]
[0099] In the experiment, the participants completed a series of emotion - inducing tasks; The experiment involved watching 6 video segments that evoked 3 emotions: positive, neutral, and negative emotions; The video segments used in the experiment are shown in Table 1. During the experiment, the EEG and facial images of the subjects were recorded; To protect the privacy of the participants, the dataset was de - identified; All personally identifiable information was deleted to ensure the privacy and confidentiality of the data.
[0100] In this embodiment, in step five, when pre - processing the facial images, several steps were taken to ensure the quality and adaptability of the data; The face images were loaded from the dataset and grayscale images were generated to reduce the computational complexity; Histogram equalization was applied to enhance the contrast of the images, thereby improving the image quality. Then, the MTCNN model was used to detect the faces in the images, and the detected faces were aligned to ensure that they had similar positions and sizes; Image normalization processing was adopted to make the mean of the pixel values 0 and the variance 1;
[0101] For the pre - processing of EEG signals:
[0102] First, the signal is filtered using a 0.1 to 60 Hz band-pass filter and a 50 Hz notch filter to remove noise; the variance of the leads is also checked to see if it is too large, if it contains too many zeros, and the data acquisition of each lead is checked to be normal. If there are leads with abnormal acquisition, interpolation is performed using the average value of the channels around the bad lead; independent component analysis (ICA) is used to identify and remove noise in electrooculogram and electromyogram signals; the electroencephalogram signal is also normalized to ensure consistent amplitudes across channels; these preprocessing steps lay a solid foundation for manual feature extraction and emotion recognition.
[0103] In this embodiment, in the experiment of step six, the training set and the test set are divided in a ratio of 4:1, and a five-fold cross-validation strategy is used to train the emotion recognition model; this model is implemented using the PyTorch framework and trained using the AdamW optimizer with a learning rate of 3e-4, and the learning rate decay uses the cosine annealing method; the batch size used for model training is 32, and the model is trained 100 times.
[0104] In Figure 1 it starts with feature extraction. The raw electroencephalogram data extracts relevant features through electroencephalogram preprocessing steps, and the facial images are cropped using the MTCNN model. Then the features extracted from the two modalities are input into the emotion recognition model, which includes a spatio-temporal attention encoder (STAE) that can fuse multi-modal features into a cohesive representation. Swin-T is used as the backbone network. Emotions are classified by an MLP layer that outputs the probabilities of positive, neutral, or negative emotions. When applying emotion recognition to consciousness detection, the emotions obtained from the classification of patients with disorders of consciousness are used to calculate distances and similarities, and by comparing the gaps between patients with disorders of consciousness and healthy people, it helps to detect disorders of consciousness.
[0105] Although the embodiments of the present invention have been shown and described (see the detailed description above), for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A multi-modal emotion recognition and consciousness detection method based on electroencephalogram and micro-expression, characterized in that: Specifically, it includes the following steps: Step 1: Feature extraction. Pay attention to the correlation of multi-channel EEG signals and extract micro-expression features. Through the feature layer, feature fusion of EEG and micro-expressions is performed to form multi-modal features; For EEG features, the processing steps are as follows: a. Preprocess the original EEG signals, including artifact removal, filtering, and component removal; b. After preprocessing, segment the EEG data of each channel into non-overlapping 1-second intervals according to the output size to extract EEG features; Step 2: Build an emotion recognition model. Select Swin Transformer as the backbone network of the model. When customizing Swin Transformer, optimize the patch size and network depth of the model to make them consistent with the dimensions of the feature tensor; ensure that the STST model meets the requirements of the emotion recognition task and can fully exert the potential of the Swin Transformer architecture; Step 3: Emotion comparison metrics. Take accuracy as the benchmark and compare the model performance with other existing methods, providing a quantitative measure of its efficacy in emotion classification; Step 4: Experimental dataset. On the public dataset, detect the emotion recognition accuracy on the multi-modal dataset MAHNOB-HCI; Step 5: Experimental data preprocessing. Use the public dataset and the self-collected dataset, and adopt a consistent data preprocessing process for these two datasets to ensure the comparability and robustness of the results; Step 6: Experimental parameter settings. Extract DE as the input of the EEG modality and extract a continuous sequence of face images as the input of the micro-expression modality; their input sizes are 56×56×5 respectively; the EEG modality represents 5 EEG frequency bands, and the micro-expression modality represents 5 face images that change continuously and uniformly within 0.2 seconds; verify the feasibility of the method on the public dataset, then pre-train on the data of healthy subjects and fine-tune on the data of MCS patients; finally, calculate the evaluation metrics on the data of UWS patients; In the above Step 1, for the multi-modal feature fusion of EEG and micro-expressions, a spatio-temporal attention mechanism multi-modal feature fusion architecture for EEG and micro-expression data is used, aiming to separate the emotionally significant features from the multi-modal data stream; In the temporal attention module, the input micro-expression features X me are initially processed by a 1×1 convolutional layer (denoted as W v ) to generate a transformed feature; subsequently, these features are reshaped to a dimension of C me ×2HW, where H and W represent the height and width of the feature map respectively; another 1×1 convolutional layer W q is used to generate query features, and then the attention distribution is calculated through a softmax layer; this attention distribution map is multiplied by the transposed feature elements to obtain the temporal attention feature map Z ta ; the temporal attention mechanism aims to capture the temporal information in the micro-expression image sequence; it uses attention weights to capture the changes over time at the same spatial location to describe the features of micro-expressions; the formula representing this process is as follows: Among them, W q , W v and W z are 1×1 convolutional layers, σ1 and σ2 are two tensor reshaping operators, is matrix dot product operation, ⊙ is channel multiplication; F sm is the Softmax operator; includes the convolution operation of W z , followed by layer regularization and the Sigmoid operator; Meanwhile, the spatial attention mechanism processes the input EEG feature X in a similar way eeg ; this mechanism calculates the spatial attention distribution through global pooling and reshaping operations; subsequently, this distribution generates the spatial attention feature Z through the softmax and sigmoid layers sa ; the spatial attention mechanism is good at capturing the spatial attention information of different channels within the same frequency band of EEG signals; the formula representing this process is as follows: Among them, W q and W v are 1×1 convolutional layers, σ1 and σ2 are two tensor reshaping operators, is matrix dot product operation, ⊙ is channel multiplication, G is global pooling operator; F sm is Softmax operator; σ s consists of a tensor reshaping operator and a Sigmoid operator; In the fusion model of EEG and micro-expression features, the attention mechanism is integrated to adapt to the single-modal input X; In the model, first, electroencephalogram and micro-expression features are combined; to strengthen feature representation, multiple layers of unimodal inputs are stacked onto a spatio-temporal attention mechanism; in the domain coordination mechanism, the temporal attention feature Z ta and the spatial attention feature Z sa are non-linearly transformed through the PReLU activation function; PReLU enhances the model's ability to learn complex patterns from data by introducing a learnable parameter to adjust the slope of the negative part of the input; then, the respective results are subjected to channel multiplication with the fused feature obtained by summation, and the result of the summation is processed through the Sigmoid function to output the final fused feature Z dc ; the calculation formula of the domain coordination mechanism is as follows: Z dc_1 = x ⊙ F sg ((F prl (Z ta ) + F prl (Z sa )) ⊙ (A ta + Z sa )) Z dc_2 = F sg ((F prl (Z ta ) + F prl (Z sa )) ⊙ (Z ta + Z sa )) Among them, Z dc_1 and Z dc_2 are the outputs obtained when inputting multi-modal data or two dimensions of electroencephalogram and micro-expression respectively; F sg represents the sigmoid operator, and F prl represents the PReLU operator. The feed-forward network design used in the spatio-temporal attention module is based on Swin Transformer, which is suitable for processing the time data of EEG signals and the spatial data of micro-expressions, can capture the local spatial features in the micro-expression data, and adapt to the time features of EEG signals, thus realizing multi-modal data fusion; in addition, the sliding window mechanism changes the calculation method of self-attention.
2. A multimodal emotion recognition and awareness detection method based on electroencephalogram and micro-expression according to claim 1, characterized in that: In Step 1, the EEG signal after passing through the band-pass filter follows a Gaussian distribution N(μ, σ 2 ), and the DE features of the electroencephalogram are calculated as follows: where f(x) is the probability density function of x, and σ 2 is the variance of this electroencephalogram; DE features belong to frequency-domain features. Electroencephalogram (EEG) signals commonly used for emotion recognition are divided into five different frequency bands: δ (0.1 - 4 Hz), θ (4 - 8 Hz), α (8 - 12 Hz), β (12 - 30 Hz), and γ (30 - 45 Hz); in the method, the DE features of each frequency band are calculated respectively; the final feature vector of each sample contains 56 channels (or 28 channels × 2) × 56 DE features × 5 frequency bands to ensure matching the size of micro-expression features.
3. A multimodal emotion recognition and awareness detection method based on electroencephalogram and micro-expression according to claim 1, characterized in that: In the second step, the input of the backbone network is a multi-modal feature tensor of size 56×56×n. Then, after splitting these tensors, linear embedding is performed to convert the features into a size suitable for STST input; each module consists of multiple layers of self-attention mechanisms and multi-layer perceptrons (MLPs). Each module performs layer normalization and then applies residual connections; the self-attention within the block operates on local windows that move between layers, enabling the model to effectively capture local and global context information.
4. A multimodal emotion recognition and awareness detection method based on electroencephalogram and micro-expression according to claim 1, characterized in that: In the third step, for the self-collected dataset, first, the data of 10 healthy subjects are used to evaluate the accuracy of the model. This step establishes a baseline for the expected emotional response; subsequently, the model is applied to the emotion classification of 21 patients collected; the emotion classification probability of the model is mapped to a continuous emotion score interval from -1 to 1, where -1 represents negative, 0 represents neutral, and 1 represents positive; to calculate the overall emotional state score, a weighted method is adopted, which includes multiplying the predicted probability of each category by its related label value to obtain a continuous score; the formula representing this process is as follows: Score=C neg ×P neg +C neu ×P neu +C pos ×P pos Among them, C neg , C neu and C pos represent the contribution values related to negative, neutral, and positive emotional states respectively, with the assigned values of -1, 0, and 1; P neg , P neu and P pos represent the probabilities of negative, neutral, and positive emotions predicted by the deep learning model respectively; this formula aims to convert the probability output of emotion classification into a continuous score from -1 to 1 to provide a more accurate comprehensive assessment of emotional tendency; to evaluate the output results of the model, Euclidean distance and cosine similarity are used as the indicators for distance calculation and similarity evaluation respectively; these indicators measure the similarity degree between the emotional patterns of patients with disorders of consciousness and those of healthy people; Accuracy is used to evaluate the classification performance of emotion recognition; accuracy is a general measure of the model's ability to correctly classify positive and negative instances in different emotion categories, and the definition of the metric is as follows: Where, TP represents true positive, TN represents true negative, FP represents false positive, and FN represents false negative; the accuracy score allows quantifying the overall classification accuracy of the model on the entire dataset, providing a global performance metric; when performing emotion recognition on patients with disorders of consciousness, the emotion labels of patients with disorders of consciousness cannot be directly used; it is impossible to ensure that the corresponding emotions have been induced in the patients, so traditional accuracy assessment cannot be performed, and Euclidean distance and cosine similarity are used as measurement criteria; First, the Euclidean distance is used to measure the spatial distance between emotional expressions; for each sample, its emotion score is represented as a vector, and the Euclidean distance is calculated to measure the difference between samples; the smaller the Euclidean distance, the closer the emotional intensity. Secondly, cosine similarity is used to measure the angular relationship between sentiment vectors; specifically, the closer the cosine similarity value is to 1, the more similar the sentiment change trends between the two; conversely, the closer the value is to -1, the more opposite the sentiment change trends between the two; for the emotion score vector A=(a1, a2,..., a n ) of healthy subjects during the induced video playback and the emotion score vector B=(b1, b2,..., b n ) of patients with disorders of consciousness, their Euclidean distance d and cosine similarity s are calculated according to the following formula: The Euclidean distance is the square root of the sum of the squared differences of the corresponding components of two vectors. The value range of the Euclidean distance is [0, 2]; the value range of the cosine similarity is [-1, 1]. The closer the value is to 1, the more similar the two vectors are, and the closer the value is to -1, the more opposite the two vectors are.
5. A multimodal emotion recognition and awareness detection method based on electroencephalogram and micro-expression according to claim 1, characterized in that: In the step 4, the data in the self-collected data set include EEG signals and facial expression images; in order to collect high-quality scalp EEG signals, the scalp EEG signals of 32 electrodes were recorded using a NuAmps amplifier and an EEG cap, and the impedance of the 32 electrodes was kept below 30 kΩ; the EEG signals were collected at a sampling rate of 1000 Hz and then band-pass filtered (0.1-50 Hz); the facial expression images were derived from the video recorded by the camera at 1080p 30fps; the subjects were in a controlled environment with minimal external stimulation; the experimental design included a passive paradigm of watching videos to induce emotions, including images and audio clips.
6. A multi-modal emotion recognition and awareness detection method based on electroencephalogram and micro-expression according to claim 1, characterized in that: In the step 5, several steps are taken to ensure the quality and adaptability of the data when preprocessing the facial image; the face image is loaded from the data set and a grayscale image is generated to reduce the computational complexity; histogram equalization is applied to enhance the contrast of the image to improve the image quality, and then the MTCNN model is used to detect the face in the image, and the detected faces are aligned to ensure that they have similar positions and sizes; image normalization is performed to make the mean of the pixel value 0 and the variance 1; Preprocessing of EEG signals: First, the signal was filtered using a 0.1 to 60 Hz bandpass filter and a 50 Hz notch filter to remove noise. It was also checked whether the variance of the lead was too large and whether it contained too many zeros. It was also checked whether the data acquisition of each lead was normal. If there were leads with abnormal acquisition, the average value of the channels around the bad lead was used for interpolation. Independent component analysis (ICA) was used to identify and remove noise from the electrooculogram and electromyography signals. The EEG signals were also standardized to ensure that the amplitude of each channel was consistent.
7. A multimodal emotion recognition and consciousness detection method based on electroencephalogram and micro-expression according to claim 1, characterized in that: In the experiment of step six, the training set and the test set are divided into a ratio of 4:1, and the emotion recognition model is trained using a five-fold cross-validation strategy; the model is implemented using the PyTorch framework and trained using the AdamW optimizer with a learning rate of 3e-4, and the learning rate decay uses the cosine annealing method; the batch size used for model training is 32, and the model is trained 100 times.
Citation Information
Patent Citations
Multi-modal emotion recognition method and system based on wearable device
CN117520826A