Methods, systems, and electronic devices for training emotion recognition models for multimodal applications.

By constructing a multimodal emotion dataset and a multimodal adaptive emotion Transformer model, the problems of limited emotion categories and short time duration in EEG emotion datasets are solved, thereby improving the accuracy and cross-domain adaptability of emotion recognition.

CN117171626BActive Publication Date: 2025-11-14SHANGHAI ZERO UNIQUE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311377602.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-23
Publication Date
2025-11-14
Estimated Expiration
2043-10-23

AI Technical Summary

Technical Problem

Existing EEG emotion datasets have limited emotion categories, short durations, and lack multimodal signals, resulting in limited effectiveness of emotion recognition models.

Method used

A multimodal emotion dataset was constructed, including EEG signals and eye movement signals. The dataset was trained using a multimodal adaptive emotion Transformer model, and feature extraction and classification were performed using a multi-view embedding module, an adaptive Transformer module, and a hybrid Transformer module.

Benefits of technology

It improves the accuracy of emotion recognition, solves the cross-domain problem, and exhibits good performance in both single-modal and multimodal scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117171626B_ABST
    Figure CN117171626B_ABST
Patent Text Reader

Abstract

This invention provides a method, system, and electronic device for training a multimodal emotion recognition model. The method includes: acquiring multimodal data from a subject viewing first and second emotional content; inputting EEG differential entropy features and eye-tracking features extracted from EEG signals and eye-tracking signals into a multimodal adaptive emotion Transformer model for emotion recognition; obtaining multimodal embeddings using multiple parallel linear layers within a multi-view embedding module; inputting the multimodal embeddings into an adaptive Transformer module and a hybrid Transformer module to output multimodal features; determining at least the predicted emotion label and emotion loss of the multimodal features using a classifier, and training based on the emotion loss to obtain the multimodal adaptive emotion Transformer model for multimodal emotion recognition. This invention constructs a multimodal emotion dataset and a multimodal adaptive emotion Transformer model, improving the accuracy of emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of affective computing, and more particularly to a method, system, and electronic device for training a multimodal emotion recognition model. Background Technology

[0002] Emotion recognition enables human-computer interaction machines to achieve emotional intelligence, allowing them to recognize, understand, and respond to human emotions and feelings. Current technologies utilize various physiological and non-physiological signals for emotion recognition. Regarding non-physiological signals, speech, facial expressions, and body language have been used by researchers to identify human emotions. However, non-physiological signals are easily faked and therefore unreliable, as individuals may conceal their true emotions. In contrast, physiological signals such as EEG (electroencephalography), EMG (electromyogram), and electrocardiogram (ECG) offer more reliable and stable options than non-physiological signals. Among these, EEG excels in emotion recognition because it is intrinsically linked to fundamental neural mechanisms, and EEG emotion datasets are used in emotion research.

[0003] In the process of realizing this invention, the inventors discovered at least the following problems in the related technology:

[0004] Existing EEG emotion datasets used in emotion research include MAHNOB-HCI, DEAP, and SEED. However, these commonly used datasets typically include positive, neutral, and negative emotion categories, resulting in a relatively limited range of emotion types. Furthermore, emotion is a complex and dynamic physiological process, with its intensity and state changing over time. Current EEG emotion datasets are relatively short. There is a lack of more diverse and multimodal EEG emotion datasets, as well as methods for training multimodal complementary emotion recognition models to fit these datasets. Summary of the Invention

[0005] To address at least the problem that existing EEG emotion datasets have limited emotion categories, short durations, and a lack of multimodal signals, resulting in limited effectiveness of emotion recognition models trained on existing EEG emotion datasets.

[0006] In a first aspect, embodiments of the present invention provide a method for training a multimodal emotion recognition model, comprising:

[0007] A multimodal emotion dataset was obtained from the subjects during the viewing of the first and second emotional materials. The multimodal emotion dataset included: electroencephalogram (EEG) signals and eye movement signals.

[0008] The EEG differential entropy features and eye movement features extracted from the EEG signals and eye movement signals are input into a multimodal adaptive emotion Transformer model for emotion recognition. The multimodal adaptive emotion Transformer model includes: a multi-view embedding module, an adaptive Transformer module, a hybrid Transformer module, and a classifier.

[0009] Using multiple parallel linear layers within the multi-view embedding module, the EEG differential entropy features and eye movement features are converted into EEG embedding sequences and eye movement embedding sequences under multi-view conditions. The EEG embedding sequences and eye movement embedding sequences are then input into the batch normalization layer within the multi-view embedding module to obtain multimodal embeddings.

[0010] The multimodal embedding is input to the adaptive Transformer module and the hybrid Transformer module, and the multimodal features are output through the EEG feedforward network or eye-tracking feedforward network in the adaptive Transformer module and the hybrid feedforward network in the hybrid Transformer module.

[0011] The classifier determines at least the predicted emotion label of the multimodal features, determines the emotion loss based on the predicted emotion label and the prepared actual emotion label, and trains based on the emotion loss to obtain a multimodal adaptive emotion Transformer model for multimodal emotion recognition.

[0012] Secondly, embodiments of the present invention provide a training system for a multimodal emotion recognition model, comprising:

[0013] The multimodal data acquisition module is used to acquire multimodal emotion datasets of subjects during the process of watching the first emotion material and the second emotion material. The multimodal emotion datasets include: electroencephalogram (EEG) signals and eye movement signals.

[0014] The feature input module is used to input the EEG differential entropy features and eye movement features extracted from the EEG signals and eye movement signals into the multimodal adaptive emotion Transformer model for emotion recognition. The multimodal adaptive emotion Transformer model includes: a multi-view embedding module, an adaptive Transformer module, a hybrid Transformer module, and a classifier.

[0015] The multimodal embedding determination module is used to convert the EEG differential entropy features and eye movement features into EEG embedding sequences and eye movement embedding sequences under multiple views using multiple parallel linear layers in the multiview embedding module, and input the EEG embedding sequences and eye movement embedding sequences into the batch normalization layer in the multiview embedding module to obtain multimodal embeddings.

[0016] A multimodal feature determination module is used to embed the multimodal features into the adaptive Transformer module and the hybrid Transformer module, and output the multimodal features through the EEG feedforward network or eye-tracking feedforward network in the adaptive Transformer module and the hybrid feedforward network in the hybrid Transformer module.

[0017] The training module is used to determine at least the predicted emotion label of the multimodal features through the classifier, determine the emotion loss based on the predicted emotion label and the pre-prepared actual emotion label, and train based on the emotion loss to obtain a multimodal adaptive emotion Transformer model for multimodal emotion recognition.

[0018] Thirdly, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of a multimodal emotion recognition model training method according to any embodiment of the present invention.

[0019] Fourthly, embodiments of the present invention provide a storage medium storing a computer program thereon, characterized in that, when the program is executed by a processor, it implements the steps of the training method for a multimodal emotion recognition model according to any embodiment of the present invention.

[0020] The beneficial effects of this invention are as follows: This method constructs a multimodal emotion dataset, including electroencephalograms (EEGs) and eye-tracking signals for seven basic emotions (happiness, sadness, fear, disgust, surprise, anger, and neutrality). Furthermore, this dataset has continuous labels indicating the intensity of emotions experienced by subjects while watching videos, which can be effectively applied to emotion recognition model training. In training the emotion recognition model, to enable training using this multimodal emotion dataset, a multimodal adaptive emotion Transformer model is constructed. The model systematically evaluates the performance of different methods in both single-modal and multimodal scenarios and also addresses cross-domain issues. The accuracy of emotion recognition is also significantly improved. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart of a method for training a multimodal emotion recognition model according to an embodiment of the present invention;

[0023] Figure 2 This is a flowchart illustrating the process of viewing emotional material in a method for training a multimodal emotion recognition model, as provided in an embodiment of the present invention.

[0024] Figure 3 This is a diagram of a multimodal adaptive emotion Transformer model architecture for a multimodal emotion recognition model training method provided by an embodiment of the present invention.

[0025] Figure 4 This is a performance diagram comparing a method for training a multimodal emotion recognition model according to an embodiment of the present invention with existing technologies;

[0026] Figure 5 This is a schematic diagram of the confusion matrix of EEG, eye tracking, or MAET (a combination of both) for training a multimodal emotion recognition model, provided by an embodiment of the present invention.

[0027] Figure 6 This is a schematic diagram illustrating the accuracy of different multimodal methods used in a training method for a multimodal emotion recognition model provided by an embodiment of the present invention.

[0028] Figure 7 This is a schematic diagram illustrating the accuracy of different methods across subjects in a multimodal emotion recognition model training method provided by an embodiment of the present invention.

[0029] Figure 8 This is a schematic diagram illustrating the accuracy of different methods with / without filtered data for training a multimodal emotion recognition model according to an embodiment of the present invention.

[0030] Figure 9 This is a schematic diagram of the MAET confusion matrix of data with and without filtering, provided in an embodiment of the present invention for a method for training a multimodal emotion recognition model.

[0031] Figure 10 This is a schematic diagram of the structure of a multimodal emotion recognition model training system provided in an embodiment of the present invention;

[0032] Figure 11 This is a schematic diagram of an embodiment of an electronic device for training a multimodal emotion recognition model, provided by an embodiment of the present invention. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0034] like Figure 1 The diagram shown is a flowchart of a method for training a multimodal emotion recognition model according to an embodiment of the present invention, including the following steps:

[0035] S11: Obtain a multimodal emotion dataset of the subjects during the viewing of the first and second emotional materials, wherein the multimodal emotion dataset includes: electroencephalogram (EEG) signals and eye movement signals;

[0036] S12: The EEG differential entropy features and eye movement features extracted from the EEG signals and eye movement signals are input into a multimodal adaptive emotion Transformer model for emotion recognition. The multimodal adaptive emotion Transformer model includes: a multi-view embedding module, an adaptive Transformer module, a hybrid Transformer module, and a classifier.

[0037] S13: Using multiple parallel linear layers in the multi-view embedding module, the EEG differential entropy features and eye movement features are converted into EEG embedding sequences and eye movement embedding sequences under multi-view conditions. The EEG embedding sequences and eye movement embedding sequences are input into the batch normalization layer in the multi-view embedding module to obtain multimodal embedding.

[0038] S14: Input the multimodal embedding into the adaptive Transformer module and the hybrid Transformer module, and output the multimodal features through the EEG feedforward network or eye-tracking feedforward network in the adaptive Transformer module and the hybrid feedforward network in the hybrid Transformer module;

[0039] S15: Determine at least the predicted emotion label of the multimodal features through the classifier, determine the emotion loss based on the predicted emotion label and the prepared actual emotion label, and train based on the emotion loss to obtain a multimodal adaptive emotion Transformer model for multimodal emotion recognition.

[0040] In this embodiment, to address the limitations of existing EEG emotion datasets—such as limited emotion categories, short durations, and a lack of multimodal signals—which result in limited effectiveness of emotion recognition models trained using existing EEG emotion datasets, this method constructs a multimodal emotion dataset. This dataset focuses on seven basic emotions (happiness, sadness, fear, disgust, surprise, anger, and neutrality) and records EEG and eye-tracking signals. This multimodal emotion dataset is then used to train a multimodal emotion recognition model across participants.

[0041] In step S11, the collected first and second emotional materials are shown to the subjects to obtain multimodal data on their viewing of these materials. It should be noted that although both the first and second emotional materials are video clips, they are materials used to evoke different emotions in the subjects. Since this method aims to expand the three commonly used emotion categories in existing technologies to seven (happiness, sadness, fear, disgust, surprise, anger, and neutrality), and happiness is easily misclassified as surprise, the selection of emotional materials is crucial to further highlight the differences between each emotion category. Specifically, this includes:

[0042] Acquire first emotion material for eliciting a first emotion category and second emotion material for eliciting a second emotion category, wherein the first emotion category includes: happiness, sadness, fear, disgust, anger, and neutrality, the first emotion material includes: video clips, the second emotion category includes: surprise, and the second emotion material includes: magic clips.

[0043] N video clips and / or magic clips of different emotion categories are selected from the first and second emotional materials, wherein N is less than a preset emotion conversion threshold. This is used to limit the breadth of emotion categories evoked by the subjects when watching the emotional materials, so as to reduce the impact of emotion conversion on the subjects' EEG signals and eye movement signals.

[0044] Based on the preset number of consecutive displays of materials of the same emotion category and the emotion category, the display order of the first emotion material and the second emotion material is adjusted, and video material clips or magic material clips are displayed to the subject in sequence according to the display order, in order to slow down the switching amplitude of the subject's emotions.

[0045] The subjects were charged with electroencephalogram (EEG) and eye movement signals while watching the video clips or magic clips.

[0046] The system receives actual emotion labels corresponding to the electroencephalogram (EEG) signals and eye movement signals input by the subject, as well as actual domain categories corresponding to the subject.

[0047] In this embodiment, the method collects electroencephalogram (EEG) and eye movement (EMT) signals to construct a multimodal emotion dataset. There are two main methods for representing emotions: dimensional models and discrete models. The dimensional model uses a two-dimensional spatial Russell model, in which all valid concepts reside at a single point with both emotional and arousal dimensions. Valence indicates whether the emotion is positive or negative, while arousal describes the activation or energy level associated with the emotional experience. The multimodal emotion dataset focuses on seven basic emotions (happiness, sadness, fear, disgust, surprise, anger, and neutrality) and records EEG and EMT signals. Furthermore, continuous labels representing the corresponding emotional intensity are collected. Emotional stimuli are selected to stimulate subjects, thereby obtaining a multimodal emotion dataset including EEG and EMT signals.

[0048] Regarding the primary and secondary emotional stimuli, to select the most effective video clips for evoking emotions, this method prepared primary emotional stimuli consisting of short video clips to evoke six emotions (excluding surprise). For example, video clips from the "Entertainment," "Emotion," and "Horror" sections of video websites were downloaded using a web scraping algorithm. To select the most effective video clips, a strategy was employed where 20 volunteers evaluated all video clips, rating each clip on a scale of 1 to 5. High-scoring clips were selected for each emotion. For the emotion of "surprise," magic videos—a secondary emotional stimulus—were specifically selected to evoke the emotion, as magic performances are effective in evoking surprise. Therefore, for each emotion (except neutral emotions), 12 clips with an average score of 3 or higher were selected. Neutral emotions consisted of eight clips, resulting in a total of 80 clips. Each clip lasted 2 to 5 minutes, with a total clip time of approximately 14,097.86 seconds. This method carefully divides these 80 segments into multiple stages, requiring subjects to complete the entire experiment within these stages, with an interval of 24 hours or longer between each stage.

[0049] About the participants: The 20 participants (10 men and 10 women) were aged between 19 and 26 years (mean: 22.5, standard deviation: 1.80). All participants were right-handed students with normal or corrected vision and normal hearing. Volunteers with high E scores on the EPQ (Eysenck Personality Questionnaire) were selected for the experiment. This method was used to ensure that participants possessed the characteristics required for accurate emotion recognition.

[0050] To ensure data quality, experiments were conducted in a controlled laboratory environment to minimize noise and other environmental interference. Furthermore, experiments were scheduled for the early morning or early afternoon to avoid confounding effects from fatigue. EEG and eye movement signals were acquired simultaneously using a 62-channel electrode cap with the International 10-20 system and a Tobii Pro Fusion eye tracker. EEG signals were acquired using the ESI NeuroScan system at a sampling rate of 1000 Hz, and eye movement signals were sampled at a sampling rate of 250 Hz.

[0051] During the presentation of emotional materials to participants, 20 trials can be conducted. Each trial consists of two parts: the first part involves participants watching the emotional material, and the second part involves self-assessment, where participants rate their emotional intensity from 1 to 10. In each segment, only five of the seven emotions are activated (N is set to 5), which reduces the impact of excessive emotional state transitions (N can also be set to 3 or 4 to consider the participants' emotional states). A 3-second countdown is displayed before and after each video clip of the emotional material to remind participants that the video is about to start or end. The playback order of the video clips is carefully arranged to avoid frequent emotional shifts. They can also be sorted according to similar emotion categories, such as playing emotional materials related to "happiness and surprise" or "disgust and anger" consecutively, thus slowing down the magnitude of emotional transitions. This carefully arranged playback order avoids sudden changes in participants' emotional values, as human emotions tend to change gradually. The eighty video clips were divided into four folds, each containing five segments from each segment, with an equal number of emotional videos in each fold.

[0052] Through the above methods, EEG and eye movement signals were obtained from the subjects' viewing of emotional material (happiness, sadness, fear, disgust, surprise, anger, and neutral).

[0053] After obtaining the EEG signals, they were preprocessed. The raw EEG signals collected during the experiment were contaminated by environmental and physiological artifacts, containing non-negligible noise, which hindered the accurate analysis of brain activity. To mitigate the impact of noise, the EEG signals were first visually inspected, and any undesirable channels were interpolated using the MNE Python toolbox. Then, bandpass filters with cutoff frequencies of 0.1Hz and 70Hz were applied to remove low-frequency noise. Additionally, a notch filter with a cutoff frequency of 50Hz was used to eliminate power frequency interference. To reduce computational complexity, the raw EEG signals were downsampled from the original sampling rate of 1000Hz to 200Hz. For EEG features, differential entropy (DE) has the ability to distinguish EEG modes between low-frequency and high-frequency energies. This method uses a 256-segment STFT (Short-Time Fourier Transform) and a 4-second non-overlapping Hanning-Hamming window to compute frequency domain features. DE features are extracted from five frequency bands (Δ: 1-4Hz, θ: 4-8Hz, α: 8-14Hz, β: 14-31Hz, γ: 31-49Hz), and defined as:

[0054]

[0055] Here, the random variable X follows a Gaussian distribution N(μ,σ). DE is equivalent to the logarithmic energy spectrum of a fixed-length sequence in a specific band. For a 62-channel EEG signal, the DE features in the five frequency bands are 310-dimensional. Based on the assumption that emotional states are defined in a continuous space and change gradually over time, this method uses LDS (linear dynamic system) to filter out components unrelated to emotional states.

[0056] Similarly, eye-tracking signals are preprocessed. Eye trackers can capture various parameters, such as pupil diameter, fixation detail, saccade detail, and fixation point detail. Among these, pupil diameter plays a crucial role in emotion recognition. However, pupil diameter is highly sensitive to ambient brightness. This method further uses linear interpolation to replace pupil diameter samples lost due to blinking. Based on the observation that subjects' responses to the same video in a controlled lighting environment have similar modalities, PCA (principal component analysis) is used to eliminate the influence of brightness on pupil diameter. The raw data is subtracted from the light reflection, which is estimated by the first principal component of the observation matrix containing pupil diameter data from the same video clips from all subjects. The remaining part then contains only the pupil response related to emotion. The DE features of the left and right pupil diameters are then calculated using STFT in four frequency bands with a 4-second non-overlapping Hanning window. In addition to the DE features, the mean and standard deviation of the pupil diameter are also calculated. The total number of features obtained from the eye movement signals is 33. Overall, the process of presenting the material to the subjects is as follows: Figure 2 As shown.

[0057] For step S12, the EEG differential entropy features and eye movement features from the EEG and eye movement signals in step S11 are extracted. These features are then input into the MAET (Multimodal Adaptive Emotion Transformer) constructed using this method. The overall architecture of the MAET model is as follows: Figure 3 As shown, it includes a multi-view embedding module, an adaptive Transformer module, a hybrid Transformer module, and a classifier. It is trained using EEG signals and eye-tracking features to give MAET the ability to process multimodal inputs. MAET can accept EEG or eye-tracking signals, or both EEG and eye movements, as input.

[0058] For step S13, in the multimodal adaptive emotion Transformer model, EEG differential entropy features and eye movement features are used as input features. Here, d is the feature size, and x is first passed to the multi-view embedding module to map a single feature to multiple labels from different views. Then, it is fed into an adaptive Transformer and a hybrid Transformer, and finally, a classifier is used to predict emotion. In this method, multimodality refers to the features of the subject's EEG modality and eye-tracking modality, while multi-view refers to different representations of the same object; for example, different forms of EEG features exhibited by the subject in response to the same emotional material.

[0059] The multi-view embedding module acquires input features and transforms them into multiple embeddings to encourage the model to focus on different views of the features. Input feature x is transformed into v embeddings by v parallel linear layers:

[0060] e i =Linear i (x), i = 1, ..., υ

[0061] in, `de` is the dimension of the embedding. The linear layer of the activation function is used to control the embedding with useful information for emotion recognition.

[0062]

[0063] in σ and σ are sigmoid functions that constrain the output value between 0 and 1. i and Element-wise multiplication, then superimposed on v embeddings, yields... The final output can be calculated as follows:

[0064]

[0065] Here, ⊙ represents the Hadamard product, and BN represents batch normalization. In this way, the input feature x is transformed into a multimodal embedding sequence from different views, which can be further processed by subsequent Transformer layers. It should be noted that the multi-view embedding module is optional for EEG, since EEG features are sequences formed by multiple channels or frequency bands and can be directly applied through multi-head self-attention. However, this method still uses this module for EEG because its inclusion has been observed to improve performance.

[0066] For step S14, the adaptive Transformer module and the hybrid Transformer module are flexible modules. Due to the flexibility of multi-head self-attention, these two modules can cover any scenario, such as inputting only EEG, inputting only eye-tracking data, or inputting both EEG and eye-tracking data simultaneously. Processing is flexibly performed from the adaptive Transformer module and the hybrid Transformer module based on the input content. Before the multimodal embedding sequence is passed to the adaptive Transformer, the embedding E is labeled with a learnable class. Its function is to aggregate information from the entire sequence and use it for later sentiment classification. To combine location and modality information, learnable location embeddings are used. and modal embedding When added to the input embedding, it can be formulated as follows:

[0067]

[0068] in,

[0069] The core component in both Adaptive Transformer and Hybrid Transformer is the same: MHSA (multi-head self-attention). Embedding The query Qi, key Ki, and value Vi are transformed through three linear layers. Self-attention can be computed as follows:

[0070]

[0071] This method uses an h-head for self-attention calculation, and each head can be represented as H. i =Attention(Q) i K i V i The output of multi-head attention is Concat(H1, H2, ..., H...). h Let W be the weight. The Adaptive Transformer introduces two modalities to replace the standard FFN (feed-forward network), namely EEG-FFN and EYE-FFN, and adaptively selects one to capture modality-specific information based on the input modality. For example, if the input is only eye-tracking data, EEG-FFN is used to encode features. Conversely, if the input contains multiple modalities, the other is used. Through the above self-attention calculation, the multimodal features of the Adaptive Transformer module and / or the Hybrid Transformer module are determined.

[0072] As one implementation, after inputting the multimodal embedding into the adaptive Transformer module and the hybrid Transformer module, the method further includes:

[0073] The emotion cue parameters are input into the adaptive Transformer module and the hybrid Transformer module;

[0074] Multimodal features are determined from the multimodal embedding and the emotion cue parameters through the EEG feedforward network or eye-tracking feedforward network in the adaptive Transformer module and the hybrid feedforward network in the hybrid Transformer module.

[0075] In this implementation, EPT (emotional prompt tuning) is introduced to adjust the model trained on multimodal inputs to adapt to a single modality. This method prepares a small set of learnable embeddings for the feature embeddings in each Transformer layer. This is called an emotional cue. Emotional cue regulation can be formalized as follows:

[0076]

[0077] Among them, TL i This represents the i-th Transformer layer. Represents the feature embedding of the i-th layer This represents the output and input of the (i+1)th Transformer layer. After all adaptive and hybrid Transformer layers, this method also applies mean pooling to all EEG or eye-tracking embeddings, and then applies C to the EEG. eeg Classifier or C for eye movements eye The classifier is handled this way so that in subsequent training steps, only the sentiment cues are adjusted along with the classifier, and it is not trained when processing multimodal signals. Therefore, while learning to predict sentiment using a single modality, the ability to process multimodal inputs is preserved.

[0078] For step S15, regarding training, the loss function is still used to train the data based on the error between the prediction and the actual data. Since this method deals with multimodal emotion data, it is necessary to classify each modality separately. At the same time, the combination of EEG and eye movement should also be considered, and the features of multimodal fusion need to be classified.

[0079] Specifically, the classifiers include: an EEG classifier, an eye-tracking classifier, a fusion classifier of EEG and eye-tracking, and a domain classifier with a gradient inversion layer;

[0080] The step of determining the predicted sentiment label of the multimodal features through the classifier, and determining the sentiment loss based on the predicted sentiment label and the pre-prepared actual sentiment label, includes:

[0081] The predicted emotion label of the multimodal features is determined based on the EEG classifier, eye-tracking classifier, and EEG and eye-tracking fusion classifier.

[0082] The predicted domain category of the multimodal features is determined based on the domain classifier.

[0083] Emotional loss is determined using pre-prepared actual emotion labels, actual domain categories, predicted emotion labels, and predicted domain categories.

[0084] Based on the emotion loss, the domain classifier, EEG classifier, eye-tracking classifier, and EEG-eye-tracking fusion classifier are subjected to domain adversarial training to obtain a multimodal adaptive emotion Transformer model across subjects.

[0085] In this embodiment, it is set that This represents the class label of the hybrid transformer output. This method introduces an attention-based fusion approach to adaptively fuse features from multiple modalities. First, the attention weight μ is calculated using the following formula. eeg and μ eye .

[0086]

[0087] in "<,>" denotes the dot product. Therefore, the fused features are extracted as follows:

[0088]

[0089] Finally, a classifier consisting of linear layers is applied to the fused features to obtain the final prediction y. The entire process can be formulated as follows:

[0090]

[0091] Where F represents the feature extractor of MAET, i.e., the component excluding the classifier, and C f This represents a fusion classifier. The objective function is the cross-entropy loss:

[0092]

[0093] in, These are real data labels.

[0094] The significant differences in EEG signals among different subjects reduce the generalizability of deep learning models and make cross-subject emotion recognition challenging. To mitigate the negative impact of individual differences, this method utilizes an adversarial domain generalization approach to make the model more robust. The core idea is to encourage the model to learn domain-invariant representations. Assume that for an input feature x, its corresponding domain label is d from K domains. A domain classifier C is designed. dIt consists of two linear layers and a GELU (Gaussian error linear units) function between them. The domain classifier is jointly trained with other components in MAET to distinguish which domain the input belongs to. However, overfitting of the domain classifier and domain label noise can lead to instability in domain adversarial training. To overcome this problem, this method employs ELS (environment label smoothing), which encourages the domain classifier to output soft probabilities. For domain labels d∈[0,1]... K Transform it into As shown below:

[0095]

[0096] Where i is from 1 to k and r controls the convergence and adversarial divergence minimization of the algorithm. During training, r is gradually decreased to... Where t is the current training step and T is the total number of steps. Therefore, the loss of the domain classifier is:

[0097]

[0098] To obfuscate the domain classifier and enable the feature extractor to learn a domain-invariant representation, a GRL (gradient reverse layer) is introduced. This layer can be ignored during forward propagation and will be transferred from C... d The gradient propagated backwards is fed back to F. Therefore, the total loss for cross-subject (i.e., cross-participant) emotion recognition based on EEG is:

[0099]

[0100] Among them, L eeg λ is the cross-entropy loss of the EEG classifier, and λ is a scaling factor that gradually changes from 0 to 1. Furthermore, this strategy makes the domain classifier insensitive to noise in the early stages of the training process. By continuously training multiple models in the above manner until the final total emotion loss reaches the preset condition, a multimodal adaptive emotion Transformer model across subjects is obtained.

[0101] As demonstrated by this implementation, our method constructs a multimodal emotion dataset, including electroencephalogram (EEG) and eye-tracking signals for seven basic emotions (happiness, sadness, fear, disgust, surprise, anger, and neutrality). Furthermore, this dataset features continuous labels indicating the intensity of emotions experienced by participants while watching videos, making it effective for training emotion recognition models. To train the emotion recognition model using this multimodal emotion dataset, a multimodal adaptive emotion Transformer model was constructed. The model systematically evaluated the performance of different methods in both single-modal and multimodal scenarios, and also addressed cross-domain issues. The accuracy of emotion recognition was also significantly improved.

[0102] Experiments illustrate this method. For the MAET hyperparameters, the number of adaptive Transformer modules (La) and hybrid Transformer modules (Lm) are set to 2 and 1, respectively. In the multi-view embedding module, the number of views v = 5 is empirically set. The number of heads h in MHSA is 4. The embedding dimension de is adjusted from {32, 64}. In the subjects' experiments, the batch size is 64, and in the cross-subject experiments, it is 256. AdamW is used as the optimizer, with the learning rate adjusted from {0.00003, 0.0001, 0.0003}. Furthermore, the weight decay is adjusted from {0.0001, 0.01, 0.1}. The hint length p is 1 or 2.

[0103] To evaluate the efficacy of EEG and eye movement in emotion recognition, a dependency model was constructed for each participant. Specifically, all four data points from a participant were merged and then divided into four segments for cross-validation. It should be noted that the input EEG and eye movement features were transformed using z-score normalization.

[0104] Regarding the classification performance of EEG, this method compares the classification performance of six existing baseline classifiers with that of MAET, namely KNN (K nearest neighbor) (k set to 1), HCNN (hierarchical convolutional neural network), RGNN (regularized graph neural network), Transformer, and Graph Convolutional Network with Channel Attention (GCNCA)

[26] . All methods were rigorously performed under the same conditions and were compared fairly with each other. Figure 4 The average accuracy and weighted F1 score for each method are shown.

[0105] MAET achieved the highest accuracy in classifying seven emotions. Specifically, the MAET machine achieved a top prediction accuracy of 58.11% while utilizing the total frequency band, highlighting MAET's effectiveness. Figure 5 The MAET (electroencephalography) analysis described the confusion matrix of MAET using only EEG signals. It can be observed that MAET distinguishes surprise and fear with higher accuracy than other emotions. Furthermore, happiness is easily misclassified as surprise, while neutral emotions are more easily confused with sadness. Moreover, compared to other emotions, sadness and anger show a greater likelihood of misclassification in EEG signal-based classification, suggesting a similarity between the neural patterns of sadness and anger.

[0106] For eye-tracking classification performance, this method compares MAET with KNN, as other baseline methods cannot handle eye-tracking features. Results are as follows: Figure 4 As shown. It is worth noting that MAET achieved the highest prediction accuracy of 50.31% and the highest F1 score of 47.10%. Figure 5 The MAET (eye-tracking) results show the confusion matrix of MAET using only eye-tracking signals. It is evident that eye-tracking signals demonstrate significant performance in distinguishing between neutral and fearful emotions. However, eye-tracking signals alone perform relatively poorly in classifying happiness and aversion, with an accuracy of less than 40%. Figure 5 In MAET (electroencephalography) and MAET (eye movement) studies, EEG signals were observed to be better at distinguishing between happiness, surprise, disgust, and anger, while eye movement signals were more accurate in distinguishing between neutral, sadness, and fear. Notably, since 27.53% of happy emotions were identified as anger, there are some similar eye movement patterns between happy and angry emotions.

[0107] Regarding the classification performance of multimodal signals, such as Figure 6 Results using different models employing EEG and eye-tracking are shown. A systematic comparison is presented of KNN, BDAE (Bimodal Deep AutoEncoder), ETF (Emotion Transformer Fusion), VigilanceNet, and MAET. For KNN, EEG and eye-tracking features are directly concatenated into a 343-dimensional feature set. MAET outperforms other methods with a best accuracy of 71.28% and an F1 score of 69.16%, demonstrating its effectiveness. The confusion matrix of MAET using multimodal signals is shown below. Figure 5The MAET (EEG + eye movement) results are shown. It can be seen that MAET achieved superior accuracy in the task of classifying emotions such as surprise, neutrality, fear, and anger. Among all emotions, fear had the highest accuracy, at 82.41%. Clearly, using multimodal methods can classify most emotions more accurately compared to using EEG or eye movement signals alone. These results indicate that multimodal methods can significantly improve classification performance, suggesting a complementary relationship between EEG and eye movement.

[0108] Regarding cross-subject (cross-participant) emotion recognition, this method employs LOSO (leave-one-subject-out) cross-validation as a strategy for measuring cross-participant performance. The results are as follows: Figure 7 As shown, due to variability among different subjects, all methods performed very poorly compared to methods under subject-related conditions, with a performance degradation of nearly 20%. MAET achieved the highest accuracy of 40.90% and an F1 score of 38.85%, demonstrating its robustness. Notably, MAET achieved the second-highest accuracy of 40.69% without AT (adversarial training), suggesting that adversarial training is somewhat helpful across subject conditions.

[0109] Regarding continuous label analysis, neutral sentiment was excluded from the experimental analysis due to the unclear definition of neutral intensity score. To further investigate the correlation between classification performance and intensity level, the classification performance of each method was compared in unfiltered and filtered cases, and low-evoking data was filtered in the filtered case. The criterion for judging whether an EEG signal is highly evoking is a score greater than 50% of the corresponding video clip. To distinguish the quality between highly evoking and low-evoking data, the results are as follows: Figure 8 As shown, MAET achieved a maximum accuracy of 58.24%. In the filtered case, the accuracy improvement exceeded 5%, with Transformer achieving a maximum increment of 7.13%.

[0110] Figure 9The confusion matrices of MAET under unfiltered and filtered conditions are described. It was observed that filtering highly elicitable data improved the accuracy for anger, disgust, and fear by 7.52%, 8.42%, and 4.7%, respectively, highlighting the importance of filtering in distinguishing these three emotions. However, the classification accuracy for happiness, surprise, and sadness decreased slightly. This may be because participants require a longer time to be aroused by stimuli for happiness, surprise, and sadness, resulting in a lack of physiological data after filtering. Since deep models require a large amount of data, insufficient data after filtering may be the reason for the decreased accuracy for these three emotions. This observation further demonstrates the effectiveness of filtering happiness, surprise, and sadness, indicating that filtered highly elicitable data is significant for classifying easily elicited emotions.

[0111] In summary, this method develops a novel multimodal emotion dataset, including EEG (encephalogram) and eye-tracking signals for seven basic emotions (happiness, sadness, fear, disgust, surprise, anger, and neutrality). Furthermore, the dataset features continuous labels indicating the intensity of emotions experienced by participants while watching videos. In the experiments, 80 video clips were selected as stimuli, and 20 participants were recruited using the EPQ questionnaire. A novel method, MAET, is also proposed, which flexibly handles multimodal inputs. The performance of different methods was systematically evaluated in both unimodal and multimodal scenarios. Cross-domain experiments were also conducted, using LOSO cross-validation to examine the performance of each method. A comparison was also made between unfiltered and filtered cases to explore the effect of filtering data based on continuous labels. Experimental results show that accuracy is significantly improved with increasing filtered data.

[0112] like Figure 10 The diagram shown is a structural schematic of a multimodal emotion recognition model training system provided in an embodiment of the present invention. The system can execute the multimodal emotion recognition model training method described in any of the above embodiments and is configured in a terminal.

[0113] This embodiment provides a training system 10 for a multimodal emotion recognition model, which includes: a multimodal data acquisition module 11, a feature input module 12, a multimodal embedding determination module 13, a multimodal feature determination module 14, and a training module 15.

[0114] The multimodal data acquisition module 11 is used to acquire multimodal data during the process of the subject viewing the first and second emotional materials. The multimodal data includes electroencephalogram (EEG) signals and eye-tracking signals. The feature input module 12 is used to input the EEG differential entropy features and eye-tracking features extracted from the EEG signals and eye-tracking signals into a multimodal adaptive emotion Transformer model for emotion recognition. The multimodal adaptive emotion Transformer model includes a multi-view embedding module, an adaptive Transformer module, a hybrid Transformer module, and a classifier. The multimodal embedding determination module 13 is used to convert the EEG differential entropy features and eye-tracking features into EEG embedding sequences and eye-tracking embedding sequences under multiple views using multiple parallel linear layers within the multi-view embedding module. The EEG embedding sequence and eye-tracking embedding sequence are input into the batch normalization layer in the multi-view embedding module to obtain multimodal embeddings. The multimodal feature determination module 14 is used to input the multimodal embeddings into the adaptive Transformer module and the hybrid Transformer module, and output multimodal features through the EEG feedforward network or eye-tracking feedforward network in the adaptive Transformer module and the hybrid feedforward network in the hybrid Transformer module. The training module 15 is used to determine at least the predicted emotion label of the multimodal features through the classifier, determine the emotion loss based on the predicted emotion label and the pre-prepared actual emotion label, and train based on the emotion loss to obtain a multimodal adaptive emotion Transformer model for multimodal emotion recognition.

[0115] This invention also provides a non-volatile computer storage medium storing computer-executable instructions that can execute the multimodal emotion recognition model training method in any of the above method embodiments.

[0116] In one embodiment, the non-volatile computer storage medium of the present invention stores computer-executable instructions, which are configured as follows:

[0117] Multimodal data were acquired from the subjects during the viewing of first and second emotional materials, wherein the multimodal data included: electroencephalogram (EEG) signals and eye movement signals;

[0118] The EEG differential entropy features and eye movement features extracted from the EEG signals and eye movement signals are input into a multimodal adaptive emotion Transformer model for emotion recognition. The multimodal adaptive emotion Transformer model includes: a multi-view embedding module, an adaptive Transformer module, a hybrid Transformer module, and a classifier.

[0119] Using multiple parallel linear layers within the multi-view embedding module, the EEG differential entropy features and eye movement features are converted into EEG embedding sequences and eye movement embedding sequences under multi-view conditions. The EEG embedding sequences and eye movement embedding sequences are then input into the batch normalization layer within the multi-view embedding module to obtain multimodal embeddings.

[0120] The multimodal embedding is input to the adaptive Transformer module and the hybrid Transformer module, and the multimodal features are output through the EEG feedforward network or eye-tracking feedforward network in the adaptive Transformer module and the hybrid feedforward network in the hybrid Transformer module.

[0121] The classifier determines at least the predicted emotion label of the multimodal features, determines the emotion loss based on the predicted emotion label and the prepared actual emotion label, and trains based on the emotion loss to obtain a multimodal adaptive emotion Transformer model for multimodal emotion recognition.

[0122] As a non-volatile computer-readable storage medium, it can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods in the embodiments of the present invention. One or more program instructions are stored in the non-volatile computer-readable storage medium, and when executed by a processor, they perform the multimodal emotion recognition model training method in any of the above method embodiments.

[0123] Figure 11 This is a schematic diagram of the hardware structure of an electronic device for a multimodal emotion recognition model training method provided in another embodiment of this application, as shown below. Figure 11 As shown, the device includes:

[0124] One or more processors 1110 and memory 1120, Figure 11 Taking a processor 1110 as an example, the device for training a multimodal emotion recognition model may further include an input device 1130 and an output device 1140.

[0125] The processor 1110, memory 1120, input device 1130, and output device 1140 can be connected via a bus or other means. Figure 11 Taking the example of a connection between China and Israel via a bus.

[0126] The memory 1120, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the multimodal emotion recognition model training method in the embodiments of this application. The processor 1110 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 1120, thereby implementing the multimodal emotion recognition model training method described in the above embodiments.

[0127] The memory 1120 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; the data storage area may store data, etc. Furthermore, the memory 1120 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 1120 may optionally include memory remotely located relative to the processor 1110, and these remote memories may be connected to the mobile device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0128] Input device 1130 can receive input numerical or character information. Output device 1140 may include display devices such as a display screen.

[0129] The one or more modules are stored in the memory 1120, and when executed by the one or more processors 1110, they execute the multimodal emotion recognition model training method in any of the above method embodiments.

[0130] The above-described product can perform the methods provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for performing the methods. Technical details not described in detail in this embodiment can be found in the methods provided in the embodiments of this application.

[0131] Non-volatile computer-readable storage media may include a stored program area and a stored data area, wherein the stored program area may store an operating system and an application program required for at least one function; the stored data area may store data created based on the use of the device, etc. Furthermore, the non-volatile computer-readable storage medium may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the non-volatile computer-readable storage medium may optionally include memory remotely located relative to the processor, and these remote memories may be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0132] This invention also provides an electronic device comprising: at least one processor and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the multimodal emotion recognition model training method of any embodiment of this invention.

[0133] The electronic devices described in this application exist in various forms, including but not limited to:

[0134] (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and primarily aim to provide voice and data communication. These terminals include smartphones, multimedia phones, feature phones, and low-end phones.

[0135] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers, possessing computing and processing capabilities, and generally also have mobile internet access features. These terminals include PDAs, MIDs, and UMPCs, such as tablet computers.

[0136] (3) Portable entertainment devices: These devices can display and play multimedia content. This category includes audio and video players, handheld game consoles, e-book readers, as well as smart toys and portable car navigation devices.

[0137] (4) Other electronic devices with data processing functions.

[0138] In this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, without necessarily requiring or implying any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising" or "including" include not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.

[0139] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0140] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for training a multimodal emotion recognition model, comprising: Multimodal data were acquired from the subjects during the viewing of first and second emotional materials, wherein the multimodal data included: electroencephalogram (EEG) signals and eye movement signals; The EEG differential entropy features and eye movement features extracted from the EEG signals and eye movement signals are input into a multimodal adaptive emotion Transformer model for emotion recognition. The multimodal adaptive emotion Transformer model includes a multi-view embedding module, an adaptive Transformer module, a hybrid Transformer module, and a classifier. The classifier includes an EEG classifier, an eye movement classifier, a fusion classifier of EEG and eye movement, and a domain classifier with a gradient inversion layer. Using multiple parallel linear layers within the multi-view embedding module, the EEG differential entropy features and eye movement features are converted into EEG embedding sequences and eye movement embedding sequences under multi-view conditions. The EEG embedding sequences and eye movement embedding sequences are then input into the batch normalization layer within the multi-view embedding module to obtain multimodal embeddings. The multimodal embedding is input to the adaptive Transformer module and the hybrid Transformer module, and the multimodal features are output through the EEG feedforward network or eye-tracking feedforward network in the adaptive Transformer module and the hybrid feedforward network in the hybrid Transformer module. The classifier determines at least the predicted emotion label of the multimodal features, and the emotion loss is determined based on the predicted emotion label and the pre-prepared actual emotion label. This includes: determining the predicted emotion label of the multimodal features based on the EEG classifier, eye-tracking classifier, and EEG-eye-tracking fusion classifier; determining the predicted domain category of the multimodal features based on the domain classifier; determining the emotion loss using the pre-prepared actual emotion label, actual domain category, predicted emotion label, and predicted domain category; and performing domain adversarial training on the domain classifier, EEG classifier, eye-tracking classifier, and EEG-eye-tracking fusion classifier based on the emotion loss to obtain a cross-subject multimodal adaptive emotion Transformer model.

2. The method according to claim 1, wherein, After inputting the multimodal embedding into the adaptive Transformer module and the hybrid Transformer module, the method further includes: The emotion cue parameters are input into the adaptive Transformer module and the hybrid Transformer module; Multimodal features are determined from the multimodal embedding and the emotion cue parameters through the EEG feedforward network or eye-tracking feedforward network in the adaptive Transformer module and the hybrid feedforward network in the hybrid Transformer module. The training based on the emotional loss includes: The emotion cue parameters and classifier are trained based on the emotion loss.

3. The method according to claim 1, wherein, The multimodal data obtained during the process of subjects viewing the first and second emotional materials includes: Acquire first emotion material for eliciting a first emotion category and second emotion material for eliciting a second emotion category, wherein the first emotion category includes: happiness, sadness, fear, disgust, anger, and neutrality, the first emotion material includes: video clips, the second emotion category includes: surprise, and the second emotion material includes: magic clips. N video clips and / or magic clips of different emotion categories are selected from the first and second emotional materials, wherein N is less than a preset emotion conversion threshold. This is used to limit the breadth of emotion categories evoked by the subjects when watching the emotional materials, so as to reduce the impact of emotion conversion on the subjects' EEG signals and eye movement signals. Based on the preset number of consecutive displays of materials of the same emotion category and the emotion category, the display order of the first emotion material and the second emotion material is adjusted, and video material clips or magic material clips are displayed to the subject in sequence according to the display order, in order to slow down the switching amplitude of the subject's emotions. The subjects were charged with electroencephalogram (EEG) and eye movement signals while watching the video clips or magic clips. The system receives actual emotion labels corresponding to the electroencephalogram (EEG) signals and eye movement signals input by the subject, as well as actual domain categories corresponding to the subject.

4. A training system for a multimodal emotion recognition model, comprising: The multimodal data acquisition module is used to acquire multimodal data of the subjects during the process of watching the first emotional material and the second emotional material, wherein the multimodal data includes: electroencephalogram (EEG) signals and eye movement signals; The feature input module is used to input the EEG differential entropy features and eye movement features extracted from the EEG signals and eye movement signals into a multimodal adaptive emotion Transformer model for emotion recognition. The multimodal adaptive emotion Transformer model includes a multi-view embedding module, an adaptive Transformer module, a hybrid Transformer module, and a classifier. The classifier includes an EEG classifier, an eye movement classifier, a fusion classifier of EEG and eye movement, and a domain classifier with a gradient inversion layer. The multimodal embedding determination module is used to convert the EEG differential entropy features and eye movement features into EEG embedding sequences and eye movement embedding sequences under multiple views using multiple parallel linear layers in the multiview embedding module, and input the EEG embedding sequences and eye movement embedding sequences into the batch normalization layer in the multiview embedding module to obtain multimodal embeddings. A multimodal feature determination module is used to embed the multimodal features into the adaptive Transformer module and the hybrid Transformer module, and output the multimodal features through the EEG feedforward network or eye-tracking feedforward network in the adaptive Transformer module and the hybrid feedforward network in the hybrid Transformer module. The training module is used to determine at least the predicted emotion label of the multimodal features through the classifier, and to determine the emotion loss based on the predicted emotion label and the pre-prepared actual emotion label. This includes: determining the predicted emotion label of the multimodal features based on the EEG classifier, eye-tracking classifier, and EEG-eye-tracking fusion classifier; determining the predicted domain category of the multimodal features based on the domain classifier; determining the emotion loss using the pre-prepared actual emotion label, actual domain category, the predicted emotion label, and the predicted domain category; and performing domain adversarial training on the domain classifier, the EEG classifier, the eye-tracking classifier, and the EEG-eye-tracking fusion classifier based on the emotion loss to obtain a multimodal adaptive emotion Transformer model across subjects.

5. The system according to claim 4, wherein, The multimodal embedding determination module is used for The emotion cue parameters are input into the adaptive Transformer module and the hybrid Transformer module; Multimodal features are determined from the multimodal embedding and the emotion cue parameters through the EEG feedforward network or eye-tracking feedforward network in the adaptive Transformer module and the hybrid feedforward network in the hybrid Transformer module. The training module is used for: The emotion cue parameters and classifier are trained based on the emotion loss.

6. The system according to claim 4, wherein, The multimodal data acquisition module is used for: Acquire first emotion material for eliciting a first emotion category and second emotion material for eliciting a second emotion category, wherein the first emotion category includes: happiness, sadness, fear, disgust, anger, and neutrality, the first emotion material includes: video clips, the second emotion category includes: surprise, and the second emotion material includes: magic clips. N video clips and / or magic clips of different emotion categories are selected from the first and second emotional materials, wherein N is less than a preset emotion conversion threshold. This is used to limit the breadth of emotion categories evoked by the subjects when watching the emotional materials, so as to reduce the impact of emotion conversion on the subjects' EEG signals and eye movement signals. Based on the preset number of consecutive displays of materials of the same emotion category and the emotion category, the display order of the first emotion material and the second emotion material is adjusted, and video material clips or magic material clips are displayed to the subject in sequence according to the display order, in order to slow down the switching amplitude of the subject's emotions. The subjects were charged with electroencephalogram (EEG) and eye movement signals while watching the video clips or magic clips. The system receives actual emotion labels corresponding to the electroencephalogram (EEG) signals and eye movement signals input by the subject, as well as actual domain categories corresponding to the subject.

7. An electronic device comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the steps of the method according to any one of claims 1-3.

8. A storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1-3.

Citation Information

Patent Citations

  • Emotional electroencephalogram feature representation method and system, electronic equipment and storage medium

    CN116671917A

  • Multi-modal emotion recognition method and system based on video-smell combined induction

    CN116894206A