Mixed emotion recognition method based on skin electric signal multi-task feature fusion

Through the multi-task feature fusion method of skin electrical signals, the CNN-LSTM model and gated network are used to achieve accurate recognition of mixed emotions, solving the problem of insufficient mixed emotions recognition performance in the prior art, and improving classification accuracy and generalization capabilities.

CN120296475APending Publication Date: 2025-07-11SOUTHEAST UNIV

Patent Information

Application Number
CN202510454220.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately identify mixed emotions, especially complex emotions with multiple emotional states at the same time, and the existing fusion mechanism of electro-skin signal characteristics is inefficient, resulting in insufficient recognition performance.

Method used

By constructing a multi-task feature fusion method based on skin electrical signals, including data acquisition, preprocessing, feature extraction and multi-task classification models, deep spatiotemporal features are extracted using the CNN-LSTM hybrid model, and dynamically fusion weights are used to realize collaborative learning of emotion classification and valence/awakening degree prediction through the gated network.

Benefits of technology

It significantly improves the classification accuracy and generalization ability of single-modal GSR in complex emotional scenarios, can accurately identify positive and negative emotions and mixed emotions, and provides efficient and lightweight emotion recognition methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296475A_ABST
    Figure CN120296475A_ABST
Patent Text Reader

Abstract

The invention discloses a skin electric signal multi-task feature fusion-based mixed emotion recognition method, which comprises the following steps of: acquiring a skin electric signal in a mixed emotion state, performing preprocessing such as filtering and baseline correction, extracting time domain, frequency domain and nonlinear manual features, and extracting deep spatial-temporal features in combination with a CNN-LSTM model; a gating network is designed to realize self-adaptive fusion of manual features and deep spatial-temporal features, so that a multi-task collaborative learning framework with emotion classification as a main task and titer / wakeup degree prediction as an auxiliary task is constructed, and accurate distinguishing of positive, negative and mixed emotions is realized through joint optimization of model parameters. According to the method, classification of mixed emotions is realized, and an extensible technical reference can be provided for the fields of emotion calculation, intelligent health monitoring and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an emotion recognition method, and particularly to a method for recognizing mixed emotions based on skin electrical signals. Background Art

[0002] Emotions are complex psychological and physiological reactions that humans generate in specific situations. Their accurate recognition is of great significance for understanding human behavior patterns, evaluating mental health status, and optimizing human-computer interaction systems. As a special phenomenon with multiple emotional states existing simultaneously, such as being both excited and anxious or having a sense of frustration in calmness, mixed emotions can more truly reflect the complexity of human emotions. Precise recognition of them can break through the limitations of traditional emotion recognition, such as distinguishing discrete emotion categories like happiness and sadness, or predicting single dimensions of valence and arousal, and provide a more refined decision-making basis for personalized emotion computing.

[0003] Current mainstream emotion recognition technologies mainly rely on external manifestation signals such as facial expression analysis, speech feature extraction, and gesture behavior recognition. However, such methods are vulnerable to subjective consciousness control and environmental interference, and have inherent defects such as strong disguisability and significant cross-cultural differences. For example, facial expressions may be deliberately suppressed due to social etiquette, and the signal-to-noise ratio of speech signals drops sharply in a noisy environment, resulting in a significant deviation between the recognition result and the true emotional state.

[0004] Skin electrical signals can objectively reflect the characteristics of emotion arousal and valence by detecting changes in sweat gland conductivity caused by sympathetic nerve activity. It has three core advantages. First, it has high sensitivity. Skin electrical signals can capture microsecond-level emotion fluctuations, and the response to high-arousal emotions such as anxiety and excitement is particularly significant. Second, it is non-invasive. The acquisition of skin electrical signals only requires contact electrodes and is suitable for long-term continuous monitoring. Finally, it has low disguisability and is not affected by language, culture, and subjective concealment behaviors.

[0005] Although there have been many emotion recognition studies based on GSR (Galvanic Skin Response), the existing technologies still have two major bottlenecks: insufficient performance in classifying mixed emotions and inefficient feature fusion mechanisms. For example, CN112006696A only supports discrete emotion classification and does not define the quantization standard for mixed emotions, resulting in a recognition blind spot in the scenario of "coexistence of positive and negative emotions". For example, CN115813389A uses manual feature stacking or simple cascading of deep learning features and does not establish a dynamic association model between features, resulting in information redundancy and the drowning of discriminative features. Summary of the Invention

[0006] Objective of the Invention: Aiming at the above-mentioned existing technologies, a hybrid emotion recognition method based on multi-task feature fusion of skin electrical signals is proposed, which can accurately distinguish positive and negative emotions and mixed emotions, and significantly improve the classification accuracy and generalization ability of single-modal GSR in complex emotion scenarios.

[0007] Technical Solution: A hybrid emotion recognition method based on multi-task feature fusion of skin electrical signals, comprising: Step 1: Collect skin electrical signal data and emotion self-assessment data of the sample population in induced positive, negative, and mixed emotion states to form a first sample set; Step 2: Preprocess the collected original skin electrical signals to eliminate motion artifacts and high-frequency noise; Step 3: Extract manual features from the preprocessed skin electrical signals, and extract deep spatio-temporal features through a CNN-LSTM hybrid model; Step 4: Construct a multi-task classification model with emotion classification as the main task and valence / arousal prediction as the auxiliary task, and train the multi-task classification model with the first sample set. After training, the multi-task classification model is used for hybrid emotion recognition based on skin electrical signals.

[0008] Further, the Step 1 includes: Design an emotion induction experiment, recruit subjects and select video clips covering positive, negative, and mixed emotions as stimulus materials, synchronously collect skin electrical signals through wearable devices, and mark the video trigger time points; Design valence / arousal, positive, negative emotion, and mixed emotion self-assessment scales, and have the subjects complete the scale annotations to generate emotion labels corresponding to the stimulus events.

[0009] Further, the Step 2 includes: Step 2.1: Adopt a multi-stage filtering strategy, eliminate low-frequency baseline drift through a high-pass filter, and then suppress high-frequency noise through a low-pass filter; Step 2.2: Use the average conductance level of the resting period signal as the individual baseline value, subtract the baseline value from the filtered skin electrical signal data to eliminate the systematic influence of individual inherent physiological differences on the signal amplitude; Step 2.3: Perform Z-score normalization on the baseline-corrected signal to unify the signal amplitude scale.

[0010] Further, in the Step 3, the manual features include time-domain features, frequency-domain features, and non-linear features; extracting deep spatio-temporal features through a CNN-LSTM hybrid model specifically includes: extracting local waveform features through a residual CNN module, and inputting them into a BiLSTM module to model temporal dependencies, capturing the phase delay and baseline correlation of emotion responses, and outputting deep spatio-temporal features.

[0011] Further, in the multi-task classification model of the Step 4, a gating network is used to dynamically allocate fusion weight α to perform weighted fusion on the manual features and the deep spatio-temporal features, and then the Softmax layer outputs the probabilities of positive, negative, and mixed emotion categories.

[0012] Further, in the model training of step 4, through the weighted loss function L total = a·L emotion + b·L valence + c·L arousal jointly optimize the model, where L total is the total loss function value, L emotion is the cross-entropy loss of the mixed emotion classification task, L valence is the mean square error loss of the valence dimension emotion prediction task, L arousal is the mean square error loss of the arousal dimension emotion prediction task, and a, b, and c are task importance coefficients that balance the contribution weights of different tasks to the total loss.

[0013] Beneficial effects: The method of the present invention constructs a multi-task framework of valence-arousal-mixed emotion, jointly optimizes manual features and deep spatio-temporal features, and designs a dynamic gating network to adaptively fuse heterogeneous features, significantly improving the classification accuracy and generalization ability of single-modal GSR in complex emotion scenarios. This method provides an efficient and lightweight emotion recognition method for scenarios such as online learning and telemedicine.

[0014] Specifically, by constructing a collaborative learning framework of the main task of mixed emotion classification and the auxiliary tasks of valence / arousal dimension emotion prediction, it overcomes the recognition blind spots of traditional models for complex emotions, significantly improves the classification accuracy in scenarios where positive and negative emotions coexist, and provides reference value for future research on mixed emotion recognition in the field of emotion recognition.

[0015] By dynamically allocating the weights of manual features and automatically extracted deep spatio-temporal features through a gating network, using the Sigmoid function to generate an adaptive fusion coefficient, suppressing high-frequency noise interference, strengthening key emotion response features, and achieving effective complementarity between different features, the recognition method maintains high robustness in scenarios of individual differences and signal drift. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a flowchart of the mixed emotion recognition method based on multi-task feature fusion of skin electrical signals of the present invention;

[0017] Figure 2 is a flowchart of signal preprocessing in step 2 of the present invention;

[0018] Figure 3 is a multi-task feature fusion model diagram based on CNN-LSTM of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The following further explains the present invention with reference to the accompanying drawings.

[0020] AsFigure 1 As shown, a hybrid emotion recognition method based on multi-task feature fusion of skin electrical signals includes four steps: data acquisition, skin electrical signal preprocessing, feature extraction and fusion, and hybrid emotion classification.

[0021] Step 1: Collect skin electrical signal data and emotion self-assessment data of the sampled population under induced positive, negative, and mixed emotion states to form the first sample set.

[0022] Specifically, Step 1 is as follows: Design an emotion induction experiment, recruit subjects and select video clips covering positive, negative, and mixed emotions as stimulus materials. Use wearable devices to collect skin electrical signals, and mark the playback time points of the video clips through triggers to ensure that the signals are strictly aligned with the emotion stimulation events. Design a valence / arousal self-assessment scale and a discrete emotion self-assessment scale, and let the subjects complete the self-assessment scales immediately after each video clip is played, and mark the playback time points of the video clips to complete the final emotion annotation corresponding to different stimulus materials.

[0023] In this embodiment, a number of subjects with good mental states are recruited, and 32 video clips are selected as emotion induction materials, including 10 for inducing positive, negative, and mixed emotions respectively, and each clip is about 25 seconds long. While the subjects are watching the videos, use the multi-modal human factor perception terminal Ergosensing ES1 (ES1 bracelet) to collect skin electrical signals. Give the subjects enough time to calm down before watching the emotion induction videos, and collect skin electrical signals 3 minutes before starting to watch the emotion induction videos as resting state skin electrical signal data, and then start watching the emotion induction videos and collect skin electrical signals as the skin electrical signal dataset. After each video clip is played, the subjects immediately complete the VA self-assessment scale and the discrete emotion self-assessment scale. The VA self-assessment scale marks the valence (Valence) and arousal (Arousal) level scores, and the discrete emotion self-assessment scale marks positive, negative, and mixed emotions, record the timestamp and associate it with the corresponding skin electrical signal segment.

[0024] Step 2: Since the skin electrical signals are interfered by factors such as limb micro-movement and environmental interference during the acquisition process, generating noise and motion artifacts, the collected original skin electrical signals are preprocessed to eliminate motion artifacts and high-frequency noise.

[0025] Such as Figure 2As shown in the figure, first, multi-level filtering is adopted to eliminate noise interference. Specifically, first, a high-pass filter is used to remove the low-frequency baseline drift caused by limb micro-movement or environmental interference. Subsequently, a low-pass filter is used to suppress high-frequency noise components and retain the signal characteristics within the effective physiological response frequency band. In this embodiment, an FIR filter with a cut-off frequency of 0.1 Hz is used to eliminate the low-frequency baseline drift, and a Butterworth filter with a cut-off frequency of 0.3 Hz is used to suppress high-frequency noise.

[0026] Then, baseline correction is performed on the denoised skin electrical signal to reduce the differences between individuals. Specifically, based on the skin electrical signal collected from the subject in the resting state, its average conductance level is calculated as the individual baseline value. The filtered skin electrical signal data is subtracted by this baseline value to eliminate the systematic influence of individual inherent physiological differences on the signal amplitude. The mathematical expression of baseline correction is shown in Equation (1):

[0027] X corrected =X raw -X baseline (1)

[0028] Where X raw is the filtered skin electrical signal data, X baseline is the average value data of the filtered skin electrical signal data in the resting state, and X corrected is the skin electrical signal value obtained after baseline correction.

[0029] Finally, normalization processing is performed on the skin electrical signal after baseline correction. Specifically, Z-score normalization processing is performed on X corrected to unify the signal amplitude scale and eliminate the heterogeneity of the response intensity across subjects.

[0030] Step 3: As Figure 3 shown, manual features are extracted from the preprocessed skin electrical signal, and deep spatio-temporal features are extracted through a CNN-LSTM hybrid model.

[0031] Step 3.1: Manual feature extraction. The following three types of manual features are extracted from the preprocessed skin electrical signal: time-domain features, frequency-domain features, and non-linear features. Specifically, time-domain features such as mean, standard deviation, maximum value of the first / second difference, minimum value, variance, etc. are extracted, frequency-domain features such as skewness, kurtosis, power spectral density, etc. of the energy spectrum in the frequency band of 0.05 - 0.5 Hz are extracted through the fast Fourier transform (FFT), and non-linear features such as Mel frequency cepstral coefficients (MFCC) are extracted.

[0032] Step 3.2: Extract deep spatio-temporal features using a hybrid architecture of convolutional neural network and long short-term memory network (CNN-LSTM model). Specifically, the preprocessed electrodermal activity signal X(t) is input, where t ∈ [1, T] is the time step and T is the signal length. Local waveform features are extracted through the residual convolutional layer, and residual connections are added to prevent gradient vanishing, and the output feature map F CNN (t) is shown in the mathematical representation of this process as Equation (2):

[0033] F CNN (t) = ReLU(X(t) * W conv + b) + X(t) (2)

[0034] In the formula, the convolutional kernel where 5 is the convolutional kernel size, 64 is the number of input channels (consistent with the input signal dimension), and 128 is the number of output channels; the bias parameter b is The activation function is ReLU.

[0035] Then, the BiLSTM (bidirectional long short-term memory network) layer is used to model the long-range temporal dependence, capture the phase delay and baseline correlation of the emotional response, and output the deep spatio-temporal features. The mathematical representation of this process is shown in Equation (3):

[0036] h t = BiLSTM(F CNN (t), h t-1 ) (3)

[0037] where is the deep spatio-temporal feature at the t-th moment, with dimension D. The initial state h0 is a zero vector.

[0038] Step 4: Construct a multi-task classification model with the main task of emotion classification and the auxiliary task of valence / arousal prediction. Emotion recognition is achieved by dynamically fusing the two types of features through a gating network, and then the recognition effect is determined through relevant metrics.

[0039] Step 4.1: Construction of the multi-task classification model. Based on the emotion annotation results of the subjects, three categories of positive emotions, negative emotions, and mixed emotions are divided as the main task of the multi-task classification model, and predicting the valence and arousal levels of emotions is used as the auxiliary task of the multi-task classification model.

[0040] Specifically, the manual features and deep spatio-temporal features are dynamically fused through a gating network, and the fusion weights of the manual features and deep spatio-temporal features are dynamically allocated. The weighted fusion process is shown in Equation (4):

[0041] F fused = α · F handcrafted + (1 - α) · F deep (4)

[0042] Among them, F fused is the fused feature, F handcrafted is the manual feature, F deep is the deep spatio-temporal feature, and α is the fusion weight generated by the Sigmoid function.

[0043] Step 4.2: Joint training and output of emotion recognition results.

[0044] The fused feature F fused is input into the Softmax layer through the fully connected layer, and the final emotion category is output, that is, the probabilities of positive, negative, and mixed emotion categories are output by the Softmax layer, so as to obtain the result of mixed emotion recognition.

[0045] In the joint training, the total loss function for balancing the main task and the auxiliary task is shown in Equation (5):

[0046] L total = a·L emotion + b·L valence + c·L arousal (5)

[0047] Among them, L total is the value of the total loss function, L emotion is the cross-entropy loss of the mixed emotion classification task, L valence is the mean square error loss of the valence dimension emotion prediction task, L arousal is the mean square error loss of the arousal dimension emotion prediction task, and a, b, c are task importance coefficients for balancing the contributions of different tasks to the total loss. In this embodiment, they are taken as 0.6, 0.3, and 0.1 in sequence.

[0048] A method for mixed emotion recognition based on multi-task feature fusion of skin electrical signals according to the present invention collects skin electrical signals in a mixed emotion state, performs preprocessing such as filtering and baseline correction, extracts time-domain, frequency-domain, and non-linear manual features, and combines a CNN-LSTM model to extract deep spatio-temporal features. A gated network is designed to realize the adaptive fusion of manual features and deep spatio-temporal features, thereby constructing a multi-task collaborative learning framework with emotion classification as the main task and valence / arousal prediction as the auxiliary task, and accurately distinguishing positive, negative, and mixed emotions by jointly optimizing model parameters. This method realizes the possibility of mixed emotion classification, provides an expandable technical reference for fields such as affective computing and intelligent health monitoring, and provides more reference value for future research on complex emotion recognition.

[0049] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A hybrid emotion recognition method based on multi-task feature fusion of skin electrical signals, characterized in that, Including: Step 1: Collect the skin electrophysiological signal data and emotion self-assessment data of the sample population under induced positive, negative, and mixed emotional states to form a first sample set; Step 2: Preprocess the collected original skin electrophysiological signals to eliminate motion artifacts and high-frequency noise; Step 3: Extract manual features from the preprocessed skin electrophysiological signals and extract deep spatio-temporal features through a CNN-LSTM hybrid model; Step 4: Construct a multi-task classification model with emotion classification as the main task and valence / arousal prediction as the auxiliary task, and train the multi-task classification model with the first sample set. After training, the multi-task classification model is used to identify mixed emotions based on skin electrophysiological signals.

2. The mixed emotion recognition method according to claim 1, wherein The said Step 1 includes: Design an emotion induction experiment, recruit subjects and select video clips covering positive, negative, and mixed emotions as stimulus materials, synchronously collect skin electrophysiological signals through wearable devices, and mark the video trigger time points; Design valence / arousal, positive, negative emotion, and mixed emotion self-assessment scales, and have the subjects complete the scale annotation to generate emotion labels corresponding to the stimulus events.

3. The mixed emotion recognition method according to claim 1, wherein, The said Step 2 includes: Step 2.1: Adopt a multi-level filtering strategy, eliminate low-frequency baseline drift through a high-pass filter, and then suppress high-frequency noise through a low-pass filter; Step 2.2: Use the average conductance level of the resting period signal as the individual baseline value, subtract the filtered skin electrophysiological signal data from this baseline value to eliminate the systematic influence of individual inherent physiological differences on the signal amplitude; Step 2.3: Perform Z-score normalization on the signals after baseline correction to unify the signal amplitude scale.

4. The mixed emotion recognition method according to claim 1, wherein In the said Step 3, the manual features include time-domain features, frequency-domain features, and non-linear features; Extracting deep spatio-temporal features through a CNN-LSTM hybrid model specifically includes: extracting local waveform features through a residual CNN module and inputting them into a BiLSTM module to model temporal dependencies, capturing the phase delay and baseline correlation of emotional responses, and outputting deep spatio-temporal features.

5. The mixed emotion recognition method according to claim 1, characterized in that In the multi-task classification model of the said Step 4, a gating network is used to dynamically allocate fusion weights α to perform weighted fusion on the manual features and the deep spatio-temporal features, and then the Softmax layer outputs the probabilities of positive, negative, and mixed emotion categories.

6. The mixed emotion recognition method according to claim 5, wherein In the model training of step 4, the weighted loss function L total = a·L emotion + b·L valence + c·L arousal is used to jointly optimize the model, where L total is the total loss function value, L emotion is the cross-entropy loss of the mixed emotion classification task, L valence is the mean square error loss of the valence dimension emotion prediction task, L arousal is the mean square error loss of the arousal dimension emotion prediction task, and a, b, and c are task importance coefficients that balance the contribution weights of different tasks to the total loss.

Citation Information

Patent Citations

  • Emotion recognition method based on skin electric signals

    CN112006696A

  • Pregnant woman delivery fear detection method and system based on skin electric signal analysis

    CN115813389A

Cited By

  • Method and system for decoding emotion awakening degree

    CN121614955A

  • A method and system for decoding arousal

    CN121614955B