A multimodal cross-subject emotion recognition method, system, electronic device and medium
By constructing a multimodal cross-participant emotion recognition network, using the self-attention encoding and cross-attention layers of EEG and eye movement data, the problem of difference across subjects is solved, and the accuracy and robustness of emotion recognition are improved.
Patent Information
- Application Number
- CN202311248898.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-25
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2043-09-25
AI Technical Summary
Existing emotions recognition methods are difficult to effectively alleviate the cross-subject problems in multimodal emotions recognition, resulting in low recognition accuracy.
A multimodal cross-participant emotion recognition network is constructed, including a cross-participant alignment module, a cross-modal alignment module and a multimodal fusion module. The trained emotion recognition model can effectively alleviate the differences between subjects through self-attention encoding and cross-attention layers of EEG and eye movement data.
Improve the accuracy of emotion recognition, and enhance the robustness and accuracy of recognition by aligning and fusion of multimodal information.
Smart Images

Figure CN117272227B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of emotion recognition technology, and in particular to a multimodal cross-subject emotion recognition method, system, electronic device and medium. Background Art
[0002] Emotion recognition plays a crucial role in understanding human intentions and behaviors and is a core research topic in the field of human-computer interaction. Because emotions are influenced by a variety of complex factors, including physiological, psychological, and environmental factors, machines often struggle to accurately understand human emotions. In recent years, intelligent emotion recognition methods have garnered widespread attention. However, uncertainties in emotion data, such as modality heterogeneity and cross-subject distribution variability, have severely limited their practical application.
[0003] Existing emotion recognition methods are primarily based on behavioral data or physiological signals. Behavioral data includes facial expressions, vocalizations, and eye movements, while physiological signals include galvanic skin response (GSR), electromyography (EMG), electrocardiography (ECG), and electroencephalography (EEG). The former is more consistent with the mechanisms of human emotion recognition, while the latter has certain anti-spoofing capabilities. Among various physiological features, EEG is currently the most widely used emotion recognition method due to its high temporal resolution and anti-spoofing properties. Because multimodal fusion techniques can provide complementary semantic information for emotion recognition, many studies have combined EEG with other modalities such as eye movements and facial expressions to improve the robustness and accuracy of emotion recognition. However, the heterogeneity of inter-modal structures poses greater challenges to the recognition task. Furthermore, in practical applications, the differences in data distribution between different subjects are far greater than the differences in emotion categories, further exacerbating the complexity of multimodal emotion recognition. In recent years, the problem of cross-subject distribution differences in emotion recognition has received increasing attention. However, existing methods are all designed for cross-subject alignment of a single modality and cannot effectively alleviate the cross-subject problem in multimodal emotion recognition. Summary of the Invention
[0004] The purpose of the present invention is to provide a multimodal cross-subject emotion recognition method, system, electronic device and medium, which can effectively alleviate the cross-subject problem in multimodal emotion recognition and improve the accuracy of emotion recognition results.
[0005] To achieve the above object, the present invention provides the following solutions:
[0006] A multimodal cross-subject emotion recognition method, comprising:
[0007] Construct a multimodal cross-subject emotion recognition network; the multimodal cross-subject emotion recognition network includes: a cross-subject alignment module, a cross-modal alignment module and a multimodal fusion module connected in sequence; the cross-subject alignment module includes an EEG subject alignment submodule and an eye movement subject alignment submodule; the EEG subject alignment submodule includes a first shallow encoder, a first deep encoder and a first domain predictor connected to the first shallow encoder, and a first emotion label predictor connected to the first shallow encoder; the eye movement subject alignment submodule includes a second shallow encoder, a second deep encoder and a second domain predictor connected to the second shallow encoder, and a second emotion label predictor connected to the second shallow encoder; the cross-modal The alignment module includes an EEG modality alignment submodule and an eye movement modality alignment submodule; the EEG modality alignment submodule includes a first self-attention encoder and a first momentum encoder both connected to the first shallow encoder; the eye movement modality alignment submodule includes a second self-attention encoder and a second momentum encoder both connected to the second shallow encoder; the multimodal fusion module includes a cross-attention layer connected to the first self-attention encoder and the second self-attention encoder, and a first fusion branch and a second fusion branch both connected to the cross-attention layer; the first fusion branch and the second fusion branch both include a fully connected layer and a SoftMax function connected in sequence; the fully connected layer is connected to the cross-attention layer;
[0008] Acquire multiple source domain data, one target domain data, true emotions corresponding to each feature data set of each source domain data, and true emotions corresponding to each feature data set of the target domain data; each of the source domain data and the target domain data includes multiple feature data sets; one feature data set includes EEG modal feature data and eye movement modal feature data of a subject at the same time;
[0009] Using all the source domain data and the target domain data to train the multimodal cross-subject emotion recognition network to obtain a trained multimodal cross-subject emotion recognition network;
[0010] An emotion recognition model is constructed based on the trained multimodal cross-subject emotion recognition network, and the emotion recognition model is used for emotion recognition; the emotion recognition model includes a first shallow encoder, a second shallow encoder, a first self-attention encoder, a second self-attention encoder, a cross-attention layer, a fully connected layer and a SoftMax function; the first self-attention encoder is connected to the first shallow encoder and the cross-attention layer, the second self-attention encoder is connected to the second shallow encoder and the cross-attention layer, and the cross-attention layer, the fully connected layer and the SoftMax function are connected in sequence.
[0011] A multimodal cross-subject emotion recognition system, comprising:
[0012] A construction module for constructing a multimodal cross-subject emotion recognition network; the multimodal cross-subject emotion recognition network includes: a cross-subject alignment module, a cross-modal alignment module and a multimodal fusion module connected in sequence; the cross-subject alignment module includes an EEG subject alignment submodule and an eye movement subject alignment submodule; the EEG subject alignment submodule includes a first shallow encoder, a first deep encoder and a first domain predictor connected to the first shallow encoder, and a first emotion label predictor connected to the first shallow encoder; the eye movement subject alignment submodule includes a second shallow encoder, a second deep encoder and a second domain predictor connected to the second shallow encoder, and a second emotion label predictor connected to the second shallow encoder; the The cross-modal alignment module includes an EEG modality alignment submodule and an eye movement modality alignment submodule; the EEG modality alignment submodule includes a first self-attention encoder and a first momentum encoder, both of which are connected to the first shallow encoder; the eye movement modality alignment submodule includes a second self-attention encoder and a second momentum encoder, both of which are connected to the second shallow encoder; the multimodal fusion module includes a cross-attention layer connected to the first self-attention encoder and the second self-attention encoder, and a first fusion branch and a second fusion branch, both of which are connected to the cross-attention layer; the first fusion branch and the second fusion branch each include a fully connected layer and a SoftMax function connected in sequence; the fully connected layer is connected to the cross-attention layer;
[0013] An acquisition module is configured to acquire multiple source domain data, one target domain data, the true emotions corresponding to each feature data set of each source domain data, and the true emotions corresponding to each feature data set of the target domain data; each source domain data and the target domain data includes multiple feature data sets; a feature data set includes EEG modal feature data and eye movement modal feature data of a subject at the same time;
[0014] a training module, configured to train the multimodal cross-subject emotion recognition network using all the source domain data and the target domain data to obtain a trained multimodal cross-subject emotion recognition network;
[0015] A recognition module is used to construct an emotion recognition model based on a trained multimodal cross-subject emotion recognition network, and the emotion recognition model is used to perform emotion recognition; the emotion recognition model includes a first shallow encoder, a second shallow encoder, a first self-attention encoder, a second self-attention encoder, a cross-attention layer, a fully connected layer and a SoftMax function; the first self-attention encoder is connected to the first shallow encoder and the cross-attention layer, the second self-attention encoder is connected to the second shallow encoder and the cross-attention layer, and the cross-attention layer, the fully connected layer and the SoftMax function are connected in sequence.
[0016] An electronic device, comprising:
[0017] A memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform the multimodal cross-subject emotion recognition method described above.
[0018] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the multimodal cross-subject emotion recognition method as described above.
[0019] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0020] The multimodal cross-subject emotion recognition network of the present invention, by setting a cross-modal alignment module and a multimodal fusion module, can fuse the information of the two modalities during model training and extract complementary information, so that the trained model can effectively alleviate the cross-subject problem in multimodal emotion recognition and improve the accuracy of emotion recognition results. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 A flowchart of a multimodal cross-subject emotion recognition method provided by an embodiment of the present invention;
[0023] Figure 2 Schematic diagram of the multimodal cross-subject emotion recognition network training process provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0025] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0026] like Figure 1 As shown, an embodiment of the present invention provides a multimodal cross-subject emotion recognition method, comprising:
[0027] Step 101: Construct a multimodal cross-subject emotion recognition network. Figure 2 As shown, the multimodal cross-subject emotion recognition network includes: a cross-subject alignment module, a cross-modal alignment module and a multimodal fusion module connected in sequence; the cross-subject alignment module includes an EEG subject alignment submodule and an eye movement subject alignment submodule; the EEG subject alignment submodule includes a first shallow encoder, a first deep encoder and a first domain predictor connected to the first shallow encoder, and a first emotion label predictor connected to the first shallow encoder; the eye movement subject alignment submodule includes a second shallow encoder, a second deep encoder and a second domain predictor connected to the second shallow encoder, and a second emotion label predictor connected to the second shallow encoder; the cross-modal alignment module includes an EEG subject alignment submodule and an eye movement subject alignment submodule; the EEG subject alignment submodule includes a first shallow encoder, a first deep encoder and a first domain predictor connected to the first shallow encoder, and a first emotion label predictor connected to the first shallow encoder; the eye movement subject alignment submodule includes a second shallow encoder, a second deep encoder and a second domain predictor connected to the second shallow encoder, and a second emotion label predictor connected to the second shallow encoder EEG modality alignment submodule and eye movement modality alignment submodule; the EEG modality alignment submodule includes a first self-attention encoder and a first momentum encoder both connected to the first shallow encoder; the eye movement modality alignment submodule includes a second self-attention encoder and a second momentum encoder both connected to the second shallow encoder; the multimodal fusion module includes a cross-attention layer connected to the first self-attention encoder and the second self-attention encoder, and a first fusion branch and a second fusion branch both connected to the cross-attention layer; the first fusion branch and the second fusion branch both include a fully connected layer and a SoftMax function connected in sequence; the fully connected layer is connected to the cross-attention layer. Considering that the shallow layer of the feature encoder is easy to generate domain-invariant features, while the deep layer is easy to generate downstream task-specific features, a shallow encoder E is set. shallow (.;θ1) and the deep encoder E deep (.;θ2).
[0028] In practical applications, the domain predictor P domain (.;θ3) connected to E shallow(.; θ1) and then the emotion label predictor P label (.;θ3) connected to E deep (.;θ2). In E shallow (.;θ1) and P domain A gradient reversal layer (GRL) is added between (.; θ3) to achieve the adversarial effect of the game, forcing the shallow encoder to generate domain-independent representations that are aligned across subjects.
[0029] Step 102: Acquire multiple source domain data, one target domain data, the true emotions corresponding to each feature dataset of each source domain data, and the true emotions corresponding to each feature dataset of the target domain data; each source domain data and target domain data includes multiple feature datasets; a feature dataset includes EEG modal feature data and eye movement modal feature data of a subject at the same time. Eye movement modal data is obtained by preprocessing and feature extraction of data derived from an eye tracker. EEG modal data is obtained by a series of processing of data derived from an EEG cap.
[0030] Step 103: using all the source domain data and the target domain data to train the multimodal cross-subject emotion recognition network to obtain a trained multimodal cross-subject emotion recognition network.
[0031] Step 104: Construct an emotion recognition model based on the trained multimodal cross-subject emotion recognition network, wherein the emotion recognition model is used to perform emotion recognition; the emotion recognition model includes a first shallow encoder, a second shallow encoder, a first self-attention encoder, a second self-attention encoder, a cross-attention layer, a fully connected layer, and a SoftMax function; the first self-attention encoder is connected to the first shallow encoder and the cross-attention layer, the second self-attention encoder is connected to the second shallow encoder and the cross-attention layer, and the cross-attention layer, the fully connected layer, and the SoftMax function are connected in sequence.
[0032] In practical applications, all the source domain data and the target domain data are used to train the multimodal cross-subject emotion recognition network to obtain a trained multimodal cross-subject emotion recognition network, specifically including:
[0033] At the current number of iterations, a self-paced learning strategy is adopted to obtain a training set at the current number of iterations based on the target domain data and all the source domain data at the previous number of iterations. The training set at the current number of iterations is used as input, and the true emotions corresponding to each feature data set in the training set at the current number of iterations and the true domain labels of each feature data set in the training set at the current number of iterations are used as output. The cross-subject alignment module is trained with the minimum cross-loss function as the goal to obtain a cross-subject alignment module trained at the current number of iterations, and the training set at the current number of iterations is deleted from all the source domain data at the previous number of iterations to obtain all the source domain data at the current number of iterations. The iteration number is then updated to enter the next iteration until there is no source domain data at the current number of iterations. The cross-subject alignment module trained at the last number of iterations is determined to be the optimal cross-subject alignment module; the domain label is the source domain or the target domain; and the cross-loss function is obtained according to the cross-subject alignment module.
[0034] All the source domain data and the target domain data are input into the first shallow encoder and the second shallow encoder in the optimal cross-subject alignment module to obtain a feature vector set; the feature vector set includes: the cross-subject aligned EEG modal feature vector corresponding to the EEG modal feature data in each feature data set of the source domain data, the cross-subject aligned EEG modal feature vector corresponding to the EEG modal feature data in each feature data set of the target domain data, the cross-subject aligned eye movement modal feature vector corresponding to the eye movement modal feature data in each feature data set of the source domain data, and the cross-subject aligned eye movement modal feature vector corresponding to the eye movement modal feature data in each feature data set of the target domain data.
[0035] Taking the feature vector set as input, taking the true emotions corresponding to each feature data set of each source domain data, the true emotions corresponding to each feature data set of the target domain data and the matching probability of each sample pair as output, and taking the minimum total loss function as the goal, the cross-modal alignment module and the multimodal fusion module are trained to obtain the optimal cross-modal alignment module and the optimal multimodal fusion module; a sample pair includes an EEG modality encoding feature vector and an eye movement modality encoding feature vector; the EEG modality encoding feature vector in the sample pair is the EEG modality encoding feature vector corresponding to the EEG modality feature data in any one of the feature data in the source domain data or the EEG modality encoding feature vector corresponding to any one of the feature data in the target domain data The EEG modality encoding feature vector corresponds to the EEG modality feature data in the data; the eye movement modality encoding feature vector in the sample pair includes the eye movement modality encoding feature vector corresponding to the eye movement modality feature data in any one of the feature data in the source domain data or the eye movement modality encoding feature vector corresponding to the eye movement modality feature data in any one of the feature data in the target domain data; the EEG modality encoding feature vector and the eye movement modality encoding feature vector are obtained by inputting the feature vector set into the first self-attention encoder and the second self-attention encoder; specifically, the EEG modality encoding feature vector corresponding to the EEG modality feature data in each feature data set of the source domain data and the The EEG modality encoding feature vector corresponding to the EEG modality feature data in each feature data set of the target domain data is obtained by inputting the cross-subject aligned EEG modality feature vector corresponding to the EEG modality feature data in each feature data set of the source domain data and the cross-subject aligned EEG modality feature vector corresponding to the EEG modality feature data in each feature data set of the target domain data into the first self-attention encoder; the eye movement modality encoding feature vector corresponding to the eye movement modality feature data in each feature data set of the source domain data or the eye movement modality encoding feature vector corresponding to the eye movement modality feature data in each feature data set of the target domain data is obtained by inputting the cross-subject aligned EEG modality feature vector corresponding to the EEG modality feature data in each feature data set of the source domain data into the first self-attention encoder. The eye movement modality feature vector after cross-subject alignment and the eye movement modality feature vector corresponding to the eye movement modality feature data in each feature data set of the target domain data are input into the second self-attention encoder; if the EEG modality encoding feature vector and the eye movement modality encoding feature vector in a sample pair correspond to the same feature data, then the sample pair matching probability of the sample pair is 1, otherwise the sample pair matching probability of the sample pair is 0; the total loss function is the sum of the contrast loss function, the matching loss function and the classification loss function; the contrast loss function is obtained according to the cross-modal alignment module; the matching loss function is obtained according to the cross-attention layer and the first fusion branch;The classification loss function is obtained according to the cross attention layer, the second fusion branch, the first shallow encoder, the first deep encoder, the first emotion label predictor, the second shallow encoder, the second deep encoder, and the second emotion label predictor.
[0036] In practical applications, at the current iteration number, a self-paced learning strategy is adopted to obtain a training set at the current iteration number based on the target domain data and all the source domain data at the previous iteration number. The training set at the current iteration number is used as input, and the true emotions corresponding to each feature data set in the training set at the current iteration number and the true domain labels of each feature data set in the training set at the current iteration number are used as output. With the minimum cross loss function as the goal, the cross-subject alignment module is trained to obtain a cross-subject alignment module trained at the current iteration number, and the training set at the current iteration number is deleted from all the source domain data at the previous iteration number to obtain all the source domain data at the current iteration number. Then, the iteration number is updated to enter the next iteration until there is no remaining source domain data. Then, the cross-subject alignment module trained at the last iteration number is determined to be the optimal cross-subject alignment module, which specifically includes:
[0037] At the current number of iterations, a self-paced learning strategy is adopted to obtain a training set at the current number of iterations based on the target domain data and all the source domain data at the previous number of iterations.
[0038] Taking the EEG modal feature data in each feature data set of the training set at the current number of iterations as input, taking the real emotions corresponding to the EEG modal feature data in each feature data set of the training set at the current number of iterations and the real domain labels of the EEG modal feature data in each feature data set of the training set at the current number of iterations as output, and taking the minimization of the first cross loss function as the goal, the parameters of the first shallow encoder and the first deep encoder in the EEG subject alignment submodule are adjusted to obtain the EEG subject alignment submodule trained at the current number of iterations; the first cross loss function is obtained based on the EEG subject alignment submodule.
[0039] Taking the eye movement modal feature data in each feature data set of the training set at the current number of iterations as input, taking the real emotions corresponding to the eye movement modal feature data in each feature data set of the training set at the current number of iterations and the real domain labels of the eye movement modal feature data in each feature data set of the training set at the current number of iterations as output, and taking the minimization of the second cross loss function as the goal, the parameters of the second shallow encoder and the second deep encoder in the eye movement subject alignment submodule are adjusted to obtain the eye movement subject alignment submodule trained at the current number of iterations; the value of the second cross loss function is obtained according to the eye movement subject alignment submodule.
[0040] Determine whether the number of all source domain data in the last iteration is greater than 1.
[0041] If so, the training set at the current iteration number is deleted from all the source domain data at the previous iteration number, to obtain all the source domain data at the current iteration number, and then the iteration number is updated to enter the next iteration.
[0042] If not, the optimal cross-subject alignment module is obtained based on the EEG subject alignment submodule trained at the current iteration number and the eye movement subject alignment submodule trained at the current iteration number.
[0043] In practical applications, the first cross loss function and the second cross loss function are L cross-subject =L label +α×L domain .
[0044] where α is the control L domain The weight parameter. L label is the classification loss on the source domain data, which is defined as:
[0045]
[0046] Where H is the cross entropy loss, M represents the state of emotion, and y m is the true value probability distribution of emotion class m, is the probability distribution of the predicted value of emotion class m. When training the first shallow encoder and the first deep encoder in the EEG subject alignment submodule, The EEG modality feature data is input into the first shallow encoder and the first emotion label prediction submodule. When training the second shallow encoder and the second deep encoder in the eye movement subject alignment submodule, This is obtained by inputting the eye movement modality feature data into the second shallow encoder and the second emotion label prediction submodule. Therefore, the update method of parameters θ1 and θ2 is:
[0047] And L domain is the domain classification loss on the source domain and the target domain, defined as: H represents the cross entropy loss, p is the true domain label, To predict domain labels, when training the first shallow encoder and the first deep encoder in the EEG subject alignment submodule, The EEG modality feature data is input into the first shallow encoder and the first domain predictor. When training the second shallow encoder and the second deep encoder in the eye movement subject alignment submodule, The eye movement modality feature data is input into the second shallow encoder and the second domain predictor, and the first domain predictor and the second domain predictor predict whether the input EEG modality feature data and eye movement modality feature data are from the target domain or the source domain. Represents that the input data comes from the source domain, and p or Denotes that it comes from the target domain. The update method of parameters θ1 and θ3 is:
[0048] In practical applications, all the source domain data and the target domain data are input into the first shallow encoder and the second shallow encoder in the optimal cross-subject alignment module to obtain a feature vector set, specifically including:
[0049] The EEG modal feature data in each feature data set of all source domain data and the EEG modal feature data in each feature data set of the target domain data are input into the first shallow encoder in the optimal cross-subject alignment module to obtain the cross-subject aligned EEG modal feature vectors corresponding to the EEG modal feature data in each feature data set of the source domain data and the cross-subject aligned EEG modal feature vectors corresponding to the EEG modal feature data in each feature data set of the target domain data.
[0050] The eye movement modal feature data in each feature data set of all source domain data and the eye movement modal feature data in each feature data set of the target domain data are input into the second shallow encoder in the optimal cross-subject alignment module to obtain the cross-subject aligned eye movement modal feature vectors corresponding to the eye movement modal feature data in each feature data set of the source domain data and the cross-subject aligned eye movement modal feature vectors corresponding to the eye movement modal feature data in each feature data set of the target domain data.
[0051] In practical applications, the feature vector set is used as input, the true emotions corresponding to each feature data set of the source domain data, the true emotions corresponding to each feature data set of the target domain data, and the matching probability of each sample pair are output, and the cross-modal alignment module and the multimodal fusion module are trained with the minimum total loss function as the goal to obtain the optimal cross-modal alignment module and the optimal multimodal fusion module, which specifically includes:
[0052] At the current number of iterations, the EEG modal feature vectors after cross-subject alignment corresponding to the EEG modal feature data in each feature data set of each source domain data and the EEG modal feature vectors after cross-subject alignment corresponding to the EEG modal feature data in each feature data set of the target domain data are input into the cross-attention layer of the cross-modal alignment module at the previous number of iterations to obtain the EEG modal coding feature vectors corresponding to the EEG modal feature data in each feature data set of each source domain data and the EEG modal coding feature vectors corresponding to the EEG modal feature data in each feature data set of the target domain data.
[0053] The EEG modal feature vectors after cross-subject alignment corresponding to the EEG modal feature data in each feature data set of each source domain data and the EEG modal feature vectors after cross-subject alignment corresponding to the EEG modal feature data in each feature data set of the target domain data are input into the first momentum encoder to obtain the EEG modal momentum feature vectors corresponding to the EEG modal feature data in each feature data set of each source domain data and the EEG modal momentum feature vectors corresponding to the EEG modal feature data in each feature data set of the target domain data.
[0054] The eye movement modality feature vectors after cross-subject alignment corresponding to the eye movement modality feature data in each feature data set of each source domain data and the eye movement modality feature vectors after cross-subject alignment corresponding to the eye movement modality feature data in each feature data set of the target domain data are input into the second self-attention encoder of the cross-modality alignment module under the last iteration number to obtain the eye movement modality encoding feature vectors corresponding to the eye movement modality feature data in each feature data set of each source domain data and the eye movement modality encoding feature vectors corresponding to the eye movement modality feature data in each feature data set of the target domain data.
[0055] The eye movement modal feature vectors after cross-subject alignment corresponding to the eye movement modal feature data in each feature data set of the source domain data and the eye movement modal feature vectors after cross-subject alignment corresponding to the eye movement modal feature data in each feature data set of the target domain data are input into the second momentum encoder to obtain the eye movement modal momentum feature vectors corresponding to the eye movement modal feature data in each feature data set of the source domain data and the eye movement modal momentum feature vectors corresponding to the eye movement modal feature data in each feature data set of the target domain data.
[0056] The value of the contrast loss function is obtained according to the EEG modal coding feature vectors corresponding to the EEG modal feature data in each feature data set of each source domain data, the EEG modal coding feature vectors corresponding to the EEG modal feature data in each feature data set of the target domain data, the EEG modal momentum feature vectors corresponding to the EEG modal feature data in each feature data set of the source domain data, the EEG modal momentum feature vectors corresponding to the EEG modal feature data in each feature data set of the target domain data, the eye movement modal coding feature vectors corresponding to the eye movement modal feature data in each feature data set of the source domain data, the eye movement modal coding feature vectors corresponding to the eye movement modal feature data in each feature data set of the target domain data, the eye movement modal momentum feature vectors corresponding to the eye movement modal feature data in each feature data set of the source domain data, and the eye movement modal momentum feature vectors corresponding to the eye movement modal feature data in each feature data set of the target domain data.
[0057] Multiple sample pairs are constructed based on the EEG modality coding feature vectors corresponding to the EEG modality feature data in each feature data set of the source domain data, the EEG modality coding feature vectors corresponding to the EEG modality feature data in each feature data set of the target domain data, the eye movement modality coding feature vectors corresponding to the eye movement modality feature data in each feature data set of the source domain data, and the eye movement modality coding feature vectors corresponding to the eye movement modality feature data in each feature data set of the target domain data.
[0058] Each of the sample pairs is input into the cross attention layer of the cross-modal alignment module at the last iteration and the first fusion branch of the multimodal fusion module at the last iteration to obtain the first emotion prediction probability corresponding to each sample pair.
[0059] According to the first emotion prediction probability corresponding to each sample pair, the value of the matching loss function is obtained.
[0060] Each of the sample pairs is input into the cross attention layer of the cross-modal alignment module at the previous iteration and the second fusion branch of the multimodal fusion module at the previous iteration to obtain the second emotion prediction probability corresponding to each sample pair.
[0061] The EEG modal feature data in each feature data set of the source domain data and the EEG modal feature data in each feature data set of the target domain data are input into the first shallow encoder, the first deep encoder and the first emotion label predictor of the optimal cross-subject alignment module to obtain the predicted emotions corresponding to the EEG modal feature data in each feature data set of the source domain data and the predicted emotions corresponding to the EEG modal feature data in each feature data set of the target domain data.
[0062] The eye movement modal feature data in each feature data set of the source domain data and the eye movement modal feature data in each feature data set of the target domain data are input into the second shallow encoder, the second deep encoder and the second emotion label predictor of the optimal cross-subject alignment module to obtain the predicted emotions corresponding to the eye movement modal feature data in each feature data set of the source domain data and the predicted emotions corresponding to the eye movement modal feature data in each feature data set of the target domain data.
[0063] The value of the classification loss function is obtained based on the predicted probability of the second emotion corresponding to each sample pair, the predicted emotion corresponding to the EEG modal feature data in each feature data set of the source domain data, the predicted emotion corresponding to the EEG modal feature data in each feature data set of the target domain data, the predicted emotion corresponding to the eye movement modal feature data in each feature data set of the source domain data, and the predicted emotion corresponding to the eye movement modal feature data in each feature data set of the target domain data.
[0064] The value of the total loss function is obtained according to the value of the contrast loss function, the value of the matching loss function, and the value of the classification loss function.
[0065] With the goal of minimizing the value of the total loss function, the parameters of the cross-modal alignment module and the parameters of the multimodal fusion module are adjusted to obtain the cross-modal alignment module under the current training times and the multimodal fusion module under the current training times, and the training times are updated to enter the next training until the training stop condition is reached, and the cross-modal alignment module under the last iteration number and the multimodal fusion module under the last iteration number are determined to be the optimal cross-modal alignment module and the optimal multimodal fusion module.
[0066] In practical applications, in emotion recognition, multimodal data can provide richer discriminative features than unimodal data, but at the same time, due to the heterogeneity between modalities, it also brings greater challenges to recognition. In order to capture potential effective intra-modal information and alleviate the structural heterogeneity between modalities, the EEG modal feature vector after cross-subject alignment and the eye movement modal feature vector after cross-subject alignment obtained by the trained first shallow encoder and the second shallow encoder are input into the designed cross-modal alignment module. The cross-modal alignment module is based on the deep encoder and sets the inter-modal contrast learning loss with a self-attention mechanism. For the sake of convenience, define are respectively aligned across subjects in the shallow encoder E shallow The projected EEG modal feature vector after cross-subject alignment and the eye movement modal feature vector after cross-subject alignment, where n is the number of samples, d x and d yare the dimensions of the feature vectors for the two modalities, respectively. The self-attention encoder module consists of three self-attention layers. Furthermore, a residual connection layer is added to mitigate the issues of vanishing gradients and network degradation. The three self-attention layers are connected sequentially, with one residual connection layer between every two self-attention layers.
[0067] In the self-attention layer, the query Q, key K and value V of the two modalities are calculated through different linear methods:
[0068] The K value of the EEG modality feature vector after cross-subject alignment is defined as K X The parameter matrix of .
[0069] The Q value of the EEG modality feature vector after cross-subject alignment is defined as Q X The parameter matrix of .
[0070] The V value of the EEG modality feature vector after cross-subject alignment is defined as V X The parameter matrix of .
[0071] The K value of the eye movement modality feature vector after alignment across subjects is defined as K Y The parameter matrix of .
[0072] The Q value of the eye movement modality feature vector after alignment across subjects is defined as Q Y The parameter matrix of .
[0073] The V value of the eye movement modality feature vector after alignment across subjects is defined as V Y The parameter matrix of .
[0074] The parameter matrix and Updated by back propagation during the training phase, it is used to generate Q, K and V. The EEG modality encoding feature vector f X According to the formula Calculate the eye movement modality encoding feature vector g Y According to the formula calculate.
[0075] Since the heterogeneity between modalities limits the performance of the model, it is necessary to adjust the modal data before modal fusion. Inspired by contrastive representation learning, a contrastive loss function L between synchronous features and asynchronous features is constructed. contrast, to reduce the structural differences between modalities. Emotional stimuli can be reflected in human physiological activities, such as EEG, facial expressions, and eye movements, so the similarity between synchronous modalities is higher than that between asynchronous modalities. Therefore, a momentum encoder with the same structure as the self-attention module is set up, and two queues are set up to store the most recent M EEG-eye movement modality sample pairs. The EEG-eye movement modality sample pair includes an EEG modality encoding feature vector and an eye movement encoding momentum feature vector. If these two vectors correspond to the same feature dataset, then the true similarity probability of this modality sample pair is set to 1, otherwise it is set to 0.
[0076] The EEG modal momentum eigenvector and eye movement modal momentum eigenvector after projection by the momentum encoder are denoted as f′ X and g′ Y .
[0077] Modal similarity is defined as:
[0078] The modal similarity between X and Y is: The modal similarity between Y and X is:
[0079] Therefore, the similarity probability of different sample modalities is calculated as follows:
[0080]
[0081]
[0082] in, and Both are used to represent the similarity between paired modalities. The bold X and Y represent the EEG and eye movement modality data of n samples, respectively. represents the modal similarity between the EEG modal data of n samples and the eye movement modal data of the mth sample, represents the modal similarity between the eye movement modal data of n samples and the EEG modal data of the mth sample, τ is the temperature parameter. The true sample modal similarity is defined as q x (X) and q y (Y).q x (X) represents the modal similarity between the EEG modal data and the eye movement modal data of n samples, q y (Y) represents the modal similarity between the eye movement modality data and the EEG modality data of n samples, and the contrast loss function is defined as the cross entropy between p and q:
[0083]
[0084] It is p xan element of (X), and p y (Y) is the same. E is the expectation; (X, Y) ~ D indicates that the data of X and Y are distributed on the domain. Although EEG and eye movement data are different in perceptual mode, they are semantically consistent because the brain's neural patterns and human behavior interact when responding to the same emotional stimuli. Therefore, in order to obtain high-order semantic information between modalities to improve the reliability of emotion recognition, the present invention provides a modality matching loss L based on contrastive learning. match To constrain the cross-attention encoder to extract potential inter-modal information. Specifically, the outputs A and B of the first and second self-attention encoders will be concatenated and fed into the cross-attention encoder as EEG-eye movement modality sample pairs Z = [AB] for further modality fusion at the feature level.
[0085] In this crisscross attention encoder, the key K Z , query Q Z , value V Z The expressions are:
[0086]
[0087]
[0088]
[0089] The emotion prediction probability u obtained by the SoftMax function xy Specifically in, K Z The parameter matrix of Represents Q Z The parameter matrix of Indicates V Z The parameter matrix of K A The parameter matrix of K B The parameter matrix of Represents Q A The parameter matrix of Represents Q B The parameter matrix of Indicates V A The parameter matrix of Indicates V B The parameter matrix, K A represents the K value of the EEG modality feature vector after cross-modality alignment, K B represents the K value of the eye movement modality feature vector after cross-modal alignment, Q Arepresents the Q value of the EEG modality feature vector after cross-modality alignment, Q B represents the Q value of the eye movement modality feature vector after cross-modal alignment, V A represents the V value of the EEG modality feature vector after cross-modality alignment, V B represents the V value of the eye movement modality feature vector after cross-modal alignment,
[0090] These representations Q obtained by the cross attention encoder Z ,K Z ,V will be input into the fully connected layer and softmax function in the first fusion branch for binary classification to obtain the emotion prediction probability u xy In the multimodal fusion module, contrastive learning is used to capture the semantic similarity between modalities. The synchronously collected EEG signals and eye movement data, i.e., modality sample pairs, constitute positive sample pairs. Asynchronous modal data constitutes negative sample pairs For any sample pair C(X i ,Y j ), expecting it to be closer to the positive sample composed of synchronous modalities and away from the negative sample composed of asynchronous modalities, and matching the modality label v of the positive sample (the two data in the sample pair correspond to the same feature dataset) xy Set to 1 if yes, otherwise, set to 0.
[0091] The modality matching loss function is: L match =E (X,Y)~D H(v xy ,u xy ).
[0092] In order to ensure the discriminability of the model in emotion classification, the present invention also uses classification loss on positive sample pairs:
[0093] L class =E (X,Y)~D H(y class ,p class )
[0094] where y class It's a real emotion, class The predicted value is the predicted emotion output by the softmax function, the first emotion label predictor, and the second emotion label predictor in the second fusion branch.
[0095] Therefore, the overall loss of the cross-modal alignment module and the multimodal fusion module, that is, the total loss function, is:
[0096] L cross-modal =L contrast +L match +Lclass .
[0097] Considering the problem that modality loss may cause the model to fail in the real world, the present invention fully utilizes the discriminant information in the cross-subject alignment module at the decision layer. In the cross-subject alignment module, the deep encoder tends to further generate category discriminant information based on the topic-independent representation, and the emotion label predictor can be generalized from the source domain to the target domain. The output P of the softmax function in the multimodal fusion module in the last layer of the model of the present invention is fusion It can be expressed as:
[0098] P fusion =softmax(p eeg +p eye +p cross-modal )
[0099] Among them, p eeg and p eye They are the normalized vectors of EEG modality data and eye movement modality data before the softmax function in the two emotion label predictors in the cross-subject alignment module. cross-modal Represents the normalized vector output from the multimodal fusion module without the softmax function. Output P fusion The index of the maximum value is the predicted value of the final emotion classification.
[0100] To reduce the negative impact of distribution shifts between subjects, an embodiment of the present invention provides a multimodal cross-subject emotion recognition network to extract domain-invariant features. Previous adversarial domain adaptation networks were mostly designed specifically for single-source domain adaptation problems. When dealing with multi-source domain adaptation problems, these methods all treat multiple source domains as a single source domain. This inconsiderate approach ignores the semantic differences between source domains and affects the stability of the model during training. To alleviate the above problems, the present invention introduces a self-paced learning strategy to gradually transform the multi-source domain adaptation problem into a single-source domain adaptation problem.
[0101] In practical applications, at the current number of iterations, a self-paced learning strategy is used to obtain a training set at the current number of iterations based on the target domain data and all the source domain data at the previous number of iterations, specifically including:
[0102] At the current number of iterations, the EEG modal feature data in each feature data set of all the source domain data at the previous number of iterations and the EEG modal feature data in each feature data set of the target domain data are input into the first shallow encoder and the first domain predictor to obtain the predicted domain labels of the EEG modal feature data in each feature data set of each source domain data at the previous number of iterations and the predicted domain labels of the EEG modal feature data in each feature data set of the target domain data.
[0103] The eye movement modality feature data in each feature data set of all the source domain data at the last iteration and the eye movement modality feature data in each feature data set of the target domain data are input into the second shallow encoder and the second domain predictor to obtain the predicted domain labels of the eye movement modality feature data in each feature data set of each source domain data at the last iteration and the predicted domain labels of the eye movement modality feature data in each feature data set of the target domain data.
[0104] According to the predicted domain labels of the EEG modal feature data in each feature data set of each source domain data at the last iteration number, the predicted domain labels of the EEG modal feature data in each feature data set of the target domain data, the predicted domain labels of the eye movement modal feature data in each feature data set of each source domain data at the last iteration number, and the predicted domain labels of the eye movement modal feature data in each feature data set of the target domain data, the distance between each source domain data and the target domain data at the last iteration number is obtained.
[0105] Determine the source domain data and target domain data at the previous iteration corresponding to the minimum distance to form the training set at the current iteration.
[0106] Taking into account the class differences between multiple source domains, this paper adopts a self-paced learning strategy to dynamically select the source domain data closest to the target domain in the projection space to participate in the training process. This strategy gradually transforms the multi-source domain adaptation problem into multiple single-source domain adaptation problems, rather than simply treating the multi-source domain as a complete single-source domain. In particular, the entire dataset can be defined as {S1, S2, ..., S N ,T}, where S i is the i-th labeled source domain, and T is the unlabeled target domain. At the beginning of the training phase, the selected domain set D selected Contains only the target domain T, that is, D selected = T. For the rest of the source domains, the shallow encoder is mapped to the target domain distance dis(S i ,T)The smallest source domain S i Add to the training set. Similarly, repeat the above process until all source domains are included in the training process. The distance between the source domain and the target domain dis(S i ,T) is defined as:
[0107] dis(S i ,T)=-H(P domain (S i ∪T,θ3),p)
[0108] Among them, P domain (S i ∪T,θ3) is the sum of Si The predicted domain labels obtained by inputting them into the domain predictor are t and t, respectively, and p is the true domain label of both. Compared with random sample selection, this method of dynamically selecting training samples can further alleviate the model instability problem caused by class differences between multiple source domains in adversarial training.
[0109] This paper proposes an emotion recognition method that unifies cross-subject alignment and heterogeneous data fusion, which can extract subject-independent emotion representations in a multimodal common feature space. First, a cognitively driven cross-subject alignment module is designed for cross-subject distribution alignment, and a self-paced learning strategy is introduced to dynamically select source domain data during training. Secondly, self-attention and cross-attention mechanisms are adopted to simultaneously capture intra-modal and inter-modal emotion-related features. Then, two contrast loss functions are applied to the cross-modal alignment module and the multimodal fusion module to further reduce the heterogeneity between modalities and explore the high-order semantic similarity between the synchronously collected EEG and eye movement data. Finally, the output of the softmax function is used as the prediction value.
[0110] An embodiment of the present invention further provides a multimodal cross-subject emotion recognition system corresponding to the above method, the system comprising:
[0111] A construction module for constructing a multimodal cross-subject emotion recognition network; the multimodal cross-subject emotion recognition network includes: a cross-subject alignment module, a cross-modal alignment module and a multimodal fusion module connected in sequence; the cross-subject alignment module includes an EEG subject alignment submodule and an eye movement subject alignment submodule; the EEG subject alignment submodule includes a first shallow encoder, a first deep encoder and a first domain predictor connected to the first shallow encoder, and a first emotion label predictor connected to the first shallow encoder; the eye movement subject alignment submodule includes a second shallow encoder, a second deep encoder and a second domain predictor connected to the second shallow encoder, and a second emotion label predictor connected to the second shallow encoder; the The cross-modal alignment module includes an EEG modality alignment submodule and an eye movement modality alignment submodule; the EEG modality alignment submodule includes a first self-attention encoder and a first momentum encoder, both of which are connected to the first shallow encoder; the eye movement modality alignment submodule includes a second self-attention encoder and a second momentum encoder, both of which are connected to the second shallow encoder; the multimodal fusion module includes a cross-attention layer connected to the first self-attention encoder and the second self-attention encoder, and a first fusion branch and a second fusion branch, both of which are connected to the cross-attention layer; the first fusion branch and the second fusion branch both include sequentially connected fully connected layers and SoftMax functions; the fully connected layers are connected to the cross-attention layer.
[0112] An acquisition module is used to acquire multiple source domain data, one target domain data, the real emotions corresponding to each feature data set of each source domain data, and the real emotions corresponding to each feature data set of the target domain data; each of the source domain data and the target domain data includes multiple feature data sets; a feature data set includes EEG modal feature data and eye movement modal feature data of an object at the same time.
[0113] A training module is used to train the multimodal cross-subject emotion recognition network using all the source domain data and the target domain data to obtain a trained multimodal cross-subject emotion recognition network.
[0114] A recognition module is used to construct an emotion recognition model based on a trained multimodal cross-subject emotion recognition network, and the emotion recognition model is used to perform emotion recognition; the emotion recognition model includes a first shallow encoder, a second shallow encoder, a first self-attention encoder, a second self-attention encoder, a cross-attention layer, a fully connected layer and a SoftMax function; the first self-attention encoder is connected to the first shallow encoder and the cross-attention layer, the second self-attention encoder is connected to the second shallow encoder and the cross-attention layer, and the cross-attention layer, the fully connected layer and the SoftMax function are connected in sequence.
[0115] In practical applications, the training module specifically includes:
[0116] The cross-subject alignment module training unit is used to adopt a self-paced learning strategy at the current iteration number to obtain a training set at the current iteration number based on the target domain data and all the source domain data at the previous iteration number, take the training set at the current iteration number as input, take the true emotions corresponding to each feature data set in the training set at the current iteration number and the true domain labels of each feature data set in the training set at the current iteration number as output, and train the cross-subject alignment module with the minimum cross loss function as the goal to obtain a cross-subject alignment module trained at the current iteration number, and delete the training set at the current iteration number from all the source domain data at the previous iteration number to obtain all the source domain data at the current iteration number, and then update the iteration number to enter the next iteration until there is no source domain data at the current iteration number, then determine that the cross-subject alignment module trained at the last iteration number is the optimal cross-subject alignment module; the domain label is the source domain or the target domain; the cross loss function is obtained according to the cross-subject alignment module.
[0117] A feature vector set determination unit is used to input all the source domain data and the target domain data into the first shallow encoder and the second shallow encoder in the optimal cross-subject alignment module to obtain a feature vector set; the feature vector set includes: the cross-subject aligned EEG modal feature vector corresponding to the EEG modal feature data in each feature data set of the source domain data, the cross-subject aligned EEG modal feature vector corresponding to the EEG modal feature data in each feature data set of the target domain data, the cross-subject aligned eye movement modal feature vector corresponding to the eye movement modal feature data in each feature data set of the source domain data, and the cross-subject aligned eye movement modal feature vector corresponding to the eye movement modal feature data in each feature data set of the target domain data.
[0118] The cross-modal alignment module and the multimodal fusion module training unit are used to take the feature vector set as input, take the true emotions corresponding to each feature data set of each source domain data, the true emotions corresponding to each feature data set of the target domain data and the matching probability of each sample pair as output, and take the minimum total loss function as the goal to train the cross-modal alignment module and the multimodal fusion module to obtain the optimal cross-modal alignment module and the optimal multimodal fusion module; a sample pair includes an EEG modality encoding feature vector and an eye movement modality encoding feature vector; the EEG modality encoding feature vector and the eye movement modality encoding feature vector are obtained by inputting the feature vector set into the first self-attention encoder and the second self-attention encoder; if a If the EEG modality encoding feature vector and the eye movement modality encoding feature vector in a sample pair correspond to the same feature data, the sample pair matching probability of the sample pair is 1, otherwise the sample pair matching probability of the sample pair is 0; the total loss function is the sum of the contrast loss function, the matching loss function and the classification loss function; the contrast loss function is obtained according to the cross-modal alignment module; the matching loss function is obtained according to the cross-attention layer and the first fusion branch; the classification loss function is obtained according to the cross-attention layer, the second fusion branch, the first shallow encoder, the first deep encoder, the first emotion label predictor, the second shallow encoder, the second deep encoder and the second emotion label predictor.
[0119] An embodiment of the present invention provides an electronic device, including:
[0120] A memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform the multimodal cross-subject emotion recognition method described in the above embodiment.
[0121] An embodiment of the present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the multimodal cross-subject emotion recognition method described in the above embodiment is implemented.
[0122] Compared with existing technologies, the present invention takes into account both the heterogeneity between multiple modalities and the differences in data distribution between different subjects. The effectiveness and stability of the performance of the cross-subject multimodal emotion recognition method proposed in the present invention are superior to other algorithms.
[0123] Advantage 1: It has wider applicability among different subjects, because the present invention performs cross-subject alignment in the cross-subject alignment module, and the model can learn effective and essential features that are not related to individual subjects but are related to emotions.
[0124] Advantage 2: Because the present invention designs a dynamic source domain selection algorithm based on a self-paced learning strategy during the training process, the emotion recognition model of the present invention is more stable and more capable of mining emotion-related features.
[0125] Advantage 3: The multimodal fusion module of the present invention mines high-order semantic information between modalities, and has better performance than the emotion recognition method based on a single modality.
[0126] Advantage 4: By setting up a cross-subject alignment module, the model can alleviate the problem of large data distribution differences between different subjects in emotion recognition during self-paced adversarial training, thereby improving the generalization and robustness of emotion recognition.
[0127] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0128] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.
Claims
1. A multimodal cross-subject emotion recognition method, characterized in that: include: Construct a multimodal cross-subject emotion recognition network; The multimodal cross-subject emotion recognition network includes: a cross-subject alignment module, a cross-modal alignment module and a multimodal fusion module connected in sequence; the cross-subject alignment module includes an EEG subject alignment submodule and an eye movement subject alignment submodule; the EEG subject alignment submodule includes a first shallow encoder, a first deep encoder and a first domain predictor connected to the first shallow encoder, and a first emotion label predictor connected to the first shallow encoder; the eye movement subject alignment submodule includes a second shallow encoder, a second deep encoder and a second domain predictor connected to the second shallow encoder, and a second emotion label predictor connected to the second shallow encoder; the cross-modal alignment module includes an EEG subject alignment submodule and an eye movement subject alignment submodule; the EEG subject alignment submodule includes a first shallow encoder, a first deep encoder and a first domain predictor connected to the first shallow encoder, and a second emotion label predictor connected to the second shallow encoder G modality alignment submodule and eye movement modality alignment submodule; the EEG modality alignment submodule includes a first self-attention encoder and a first momentum encoder both connected to the first shallow encoder; the eye movement modality alignment submodule includes a second self-attention encoder and a second momentum encoder both connected to the second shallow encoder; the multimodal fusion module includes a cross-attention layer connected to the first self-attention encoder and the second self-attention encoder and a first fusion branch and a second fusion branch both connected to the cross-attention layer; the first fusion branch and the second fusion branch both include a fully connected layer and a SoftMax function connected in sequence; the fully connected layer is connected to the cross-attention layer; Acquire multiple source domain data, one target domain data, true emotions corresponding to each feature data set of each source domain data, and true emotions corresponding to each feature data set of the target domain data; each of the source domain data and the target domain data includes multiple feature data sets; one feature data set includes EEG modal feature data and eye movement modal feature data of a subject at the same time; Using all the source domain data and the target domain data to train the multimodal cross-subject emotion recognition network to obtain a trained multimodal cross-subject emotion recognition network; An emotion recognition model is constructed based on the trained multimodal cross-subject emotion recognition network, and the emotion recognition model is used for emotion recognition; the emotion recognition model includes a first shallow encoder, a second shallow encoder, a first self-attention encoder, a second self-attention encoder, a cross-attention layer, a fully connected layer and a SoftMax function; the first self-attention encoder is connected to the first shallow encoder and the cross-attention layer, the second self-attention encoder is connected to the second shallow encoder and the cross-attention layer, and the cross-attention layer, the fully connected layer and the SoftMax function are connected in sequence.
2. The multimodal cross-subject emotion recognition method according to claim 1, characterized in that: Using all the source domain data and the target domain data to train the multimodal cross-subject emotion recognition network to obtain a trained multimodal cross-subject emotion recognition network specifically includes: At the current iteration number, a self-paced learning strategy is adopted to obtain a training set at the current iteration number based on the target domain data and all the source domain data at the previous iteration number, the training set at the current iteration number is used as input, the real emotions corresponding to each feature data set in the training set at the current iteration number and the real domain labels of each feature data set in the training set at the current iteration number are used as output, and the cross-subject alignment module is trained with the minimum cross-loss function as the goal to obtain a cross-subject alignment module trained at the current iteration number, and the training set at the current iteration number is deleted from all the source domain data at the previous iteration number to obtain all the source domain data at the current iteration number, and then the iteration number is updated to enter the next iteration until there is no source domain data at the current iteration number, then the cross-subject alignment module trained at the last iteration number is determined to be the optimal cross-subject alignment module; the domain label is the source domain or the target domain; the cross-loss function is obtained according to the cross-subject alignment module; Inputting all the source domain data and the target domain data into the first shallow encoder and the second shallow encoder in the optimal cross-subject alignment module to obtain a feature vector set; the feature vector set includes: the cross-subject aligned EEG modality feature vector corresponding to the EEG modality feature data in each feature data set of the source domain data, the cross-subject aligned EEG modality feature vector corresponding to the EEG modality feature data in each feature data set of the target domain data, the cross-subject aligned eye movement modality feature vector corresponding to the eye movement modality feature data in each feature data set of the source domain data, and the cross-subject aligned eye movement modality feature vector corresponding to the eye movement modality feature data in each feature data set of the target domain data; Taking the feature vector set as input, taking the true emotions corresponding to each feature data set of each source domain data, the true emotions corresponding to each feature data set of the target domain data and the matching probability of each sample pair as output, and taking the minimum total loss function as the goal, the cross-modal alignment module and the multimodal fusion module are trained to obtain the optimal cross-modal alignment module and the optimal multimodal fusion module; a sample pair includes an EEG modality encoding feature vector and an eye movement modality encoding feature vector; the EEG modality encoding feature vector and the eye movement modality encoding feature vector are obtained by inputting the feature vector set into the first self-attention encoder and the second self-attention encoder; if the EEG module in a sample pair If the state encoding feature vector and the eye movement modality encoding feature vector correspond to the same feature data, the sample pair matching probability of the sample pair is 1, otherwise the sample pair matching probability of the sample pair is 0; the total loss function is the sum of the contrast loss function, the matching loss function and the classification loss function; the contrast loss function is obtained according to the cross-modal alignment module; the matching loss function is obtained according to the cross-attention layer and the first fusion branch; the classification loss function is obtained according to the cross-attention layer, the second fusion branch, the first shallow encoder, the first deep encoder, the first emotion label predictor, the second shallow encoder, the second deep encoder and the second emotion label predictor.
3. The multimodal cross-subject emotion recognition method according to claim 2, characterized in that: At the current iteration number, a self-paced learning strategy is adopted to obtain a training set at the current iteration number based on the target domain data and all the source domain data at the previous iteration number. The training set at the current iteration number is used as input, and the true emotions corresponding to each feature data set in the training set at the current iteration number and the true domain labels of each feature data set in the training set at the current iteration number are used as output. With the minimum cross loss function as the goal, the cross-subject alignment module is trained to obtain a cross-subject alignment module trained at the current iteration number, and the training set at the current iteration number is deleted from all the source domain data at the previous iteration number to obtain all the source domain data at the current iteration number. Then, the iteration number is updated to enter the next iteration until there is no remaining source domain data. Then, the cross-subject alignment module trained at the last iteration number is determined to be the optimal cross-subject alignment module, specifically including: At the current number of iterations, a self-paced learning strategy is used to obtain a training set at the current number of iterations based on the target domain data and all the source domain data at the previous number of iterations; Taking the EEG modal feature data in each feature data set of the training set at the current iteration number as input, taking the true emotion corresponding to the EEG modal feature data in each feature data set of the training set at the current iteration number and the true domain label of the EEG modal feature data in each feature data set of the training set at the current iteration number as output, and minimizing the first cross loss function as the goal, adjusting the parameters of the first shallow encoder and the first deep encoder in the EEG subject alignment submodule to obtain the EEG subject alignment submodule trained at the current iteration number; the first cross loss function is obtained based on the EEG subject alignment submodule; Taking the eye movement modal feature data in each feature data set of the training set at the current iteration number as input, taking the real emotions corresponding to the eye movement modal feature data in each feature data set of the training set at the current iteration number and the real domain labels of the eye movement modal feature data in each feature data set of the training set at the current iteration number as output, and taking the minimization of the second cross loss function as the goal, adjusting the parameters of the second shallow encoder and the second deep encoder in the eye movement subject alignment submodule to obtain the eye movement subject alignment submodule trained at the current iteration number; the value of the second cross loss function is obtained according to the eye movement subject alignment submodule; Determine whether the number of all source domain data in the last iteration is greater than 1; If so, the training set at the current iteration number is deleted from all the source domain data at the previous iteration number, to obtain all the source domain data at the current iteration number, and then the iteration number is updated to enter the next iteration; If not, the optimal cross-subject alignment module is obtained based on the EEG subject alignment submodule trained at the current iteration number and the eye movement subject alignment submodule trained at the current iteration number.
4. The multimodal cross-subject emotion recognition method according to claim 2, characterized in that: Inputting all the source domain data and the target domain data into the first shallow encoder and the second shallow encoder in the optimal cross-subject alignment module to obtain a feature vector set, specifically including: Inputting the EEG modal feature data in each feature data set of all source domain data and the EEG modal feature data in each feature data set of the target domain data into the first shallow encoder in the optimal cross-subject alignment module, obtaining the cross-subject aligned EEG modal feature vectors corresponding to the EEG modal feature data in each feature data set of the source domain data and the cross-subject aligned EEG modal feature vectors corresponding to the EEG modal feature data in each feature data set of the target domain data; The eye movement modal feature data in each feature data set of all source domain data and the eye movement modal feature data in each feature data set of the target domain data are input into the second shallow encoder in the optimal cross-subject alignment module to obtain the cross-subject aligned eye movement modal feature vectors corresponding to the eye movement modal feature data in each feature data set of the source domain data and the cross-subject aligned eye movement modal feature vectors corresponding to the eye movement modal feature data in each feature data set of the target domain data.
5. The multimodal cross-subject emotion recognition method according to claim 2, characterized in that: Taking the feature vector set as input, taking the true emotion corresponding to each feature data set of each source domain data, the true emotion corresponding to each feature data set of the target domain data, and the matching probability of each sample pair as output, and minimizing the total loss function as the goal, the cross-modal alignment module and the multimodal fusion module are trained to obtain the optimal cross-modal alignment module and the optimal multimodal fusion module, specifically including: At the current number of iterations, the EEG modal feature vectors after cross-subject alignment corresponding to the EEG modal feature data in each feature data set of each source domain data and the EEG modal feature vectors after cross-subject alignment corresponding to the EEG modal feature data in each feature data set of the target domain data are input into the cross-attention layer of the cross-modal alignment module at the previous number of iterations to obtain the EEG modal coding feature vectors corresponding to the EEG modal feature data in each feature data set of each source domain data and the EEG modal coding feature vectors corresponding to the EEG modal feature data in each feature data set of the target domain data; Inputting the EEG modal feature vectors corresponding to the EEG modal feature data in each feature data set of each source domain data and the EEG modal feature vectors corresponding to the EEG modal feature data in each feature data set of the target domain data into a first momentum encoder to obtain the EEG modal momentum feature vectors corresponding to the EEG modal feature data in each feature data set of each source domain data and the EEG modal momentum feature vectors corresponding to the EEG modal feature data in each feature data set of the target domain data; Inputting the eye movement modality feature vectors after cross-subject alignment corresponding to the eye movement modality feature data in each feature data set of each source domain data and the eye movement modality feature vectors after cross-subject alignment corresponding to the eye movement modality feature data in each feature data set of the target domain data into the second self-attention encoder of the cross-modality alignment module at the last iteration number, to obtain the eye movement modality encoding feature vectors corresponding to the eye movement modality feature data in each feature data set of each source domain data and the eye movement modality encoding feature vectors corresponding to the eye movement modality feature data in each feature data set of the target domain data; Inputting the eye movement modality feature vectors corresponding to the eye movement modality feature data in each feature data set of each source domain data and the eye movement modality feature vectors corresponding to the eye movement modality feature data in each feature data set of the target domain data after cross-subject alignment into a second momentum encoder, thereby obtaining the eye movement modality momentum feature vectors corresponding to the eye movement modality feature data in each feature data set of each source domain data and the eye movement modality momentum feature vectors corresponding to the eye movement modality feature data in each feature data set of the target domain data; Obtaining a value of a contrast loss function according to the EEG modal coding feature vector corresponding to the EEG modal feature data in each feature data set of each source domain data, the EEG modal coding feature vector corresponding to the EEG modal feature data in each feature data set of the target domain data, the EEG modal momentum feature vector corresponding to the EEG modal feature data in each feature data set of the source domain data, the EEG modal momentum feature vector corresponding to the EEG modal feature data in each feature data set of the target domain data, the eye movement modal coding feature vector corresponding to the eye movement modal feature data in each feature data set of the source domain data, the eye movement modal coding feature vector corresponding to the eye movement modal feature data in each feature data set of the target domain data, the eye movement modal momentum feature vector corresponding to the eye movement modal feature data in each feature data set of the source domain data, and the eye movement modal momentum feature vector corresponding to the eye movement modal feature data in each feature data set of the target domain data; constructing a plurality of sample pairs according to the EEG modality encoding feature vectors corresponding to the EEG modality feature data in each feature data set of each source domain data, the EEG modality encoding feature vectors corresponding to the EEG modality feature data in each feature data set of the target domain data, the eye movement modality encoding feature vectors corresponding to the eye movement modality feature data in each feature data set of the source domain data, and the eye movement modality encoding feature vectors corresponding to the eye movement modality feature data in each feature data set of the target domain data; Inputting each of the sample pairs into the cross attention layer of the cross-modal alignment module at the previous iteration and the first fusion branch of the multimodal fusion module at the previous iteration to obtain a first emotion prediction probability corresponding to each sample pair; According to the first emotion prediction probability corresponding to each sample pair, the value of the matching loss function is obtained; Inputting each of the sample pairs into the cross attention layer of the cross-modal alignment module at the previous iteration and the second fusion branch of the multimodal fusion module at the previous iteration to obtain a second emotion prediction probability corresponding to each sample pair; Inputting the EEG modal feature data in each feature data set of each source domain data and the EEG modal feature data in each feature data set of the target domain data into the first shallow encoder, the first deep encoder, and the first emotion label predictor of the optimal cross-subject alignment module, obtaining the predicted emotions corresponding to the EEG modal feature data in each feature data set of each source domain data and the predicted emotions corresponding to the EEG modal feature data in each feature data set of the target domain data; Inputting the eye movement modality feature data in each feature data set of each source domain data and the eye movement modality feature data in each feature data set of the target domain data into the second shallow encoder, the second deep encoder, and the second emotion label predictor of the optimal cross-subject alignment module, to obtain the predicted emotions corresponding to the eye movement modality feature data in each feature data set of each source domain data and the predicted emotions corresponding to the eye movement modality feature data in each feature data set of the target domain data; Obtaining a value of a classification loss function according to the predicted probability of the second emotion corresponding to each sample pair, the predicted emotion corresponding to the EEG modal feature data in each feature data set of the source domain data, the predicted emotion corresponding to the EEG modal feature data in each feature data set of the target domain data, the predicted emotion corresponding to the eye movement modal feature data in each feature data set of the source domain data, and the predicted emotion corresponding to the eye movement modal feature data in each feature data set of the target domain data; Obtaining a value of the total loss function according to the value of the contrast loss function, the value of the matching loss function, and the value of the classification loss function; With the goal of minimizing the value of the total loss function, the parameters of the cross-modal alignment module and the parameters of the multimodal fusion module are adjusted to obtain the cross-modal alignment module under the current training times and the multimodal fusion module under the current training times, and the training times are updated to enter the next training until the training stop condition is reached, and the cross-modal alignment module under the last iteration number and the multimodal fusion module under the last iteration number are determined to be the optimal cross-modal alignment module and the optimal multimodal fusion module.
6. The multimodal cross-subject emotion recognition method according to claim 3, characterized in that: At the current iteration number, a self-paced learning strategy is used to obtain a training set at the current iteration number based on the target domain data and all the source domain data at the previous iteration number, specifically including: At the current iteration number, inputting the EEG modal feature data in each feature data set of all the source domain data at the previous iteration number and the EEG modal feature data in each feature data set of the target domain data into the first shallow encoder and the first domain predictor, obtaining the predicted domain labels of the EEG modal feature data in each feature data set of each source domain data at the previous iteration number and the predicted domain labels of the EEG modal feature data in each feature data set of the target domain data; Inputting the eye movement modality feature data in each feature data set of all the source domain data at the last iteration and the eye movement modality feature data in each feature data set of the target domain data into a second shallow encoder and a second domain predictor, obtaining predicted domain labels for the eye movement modality feature data in each feature data set of each source domain data at the last iteration and predicted domain labels for the eye movement modality feature data in each feature data set of the target domain data; Obtaining the distance between each source domain data and the target domain data at the last iteration number according to the predicted domain labels of the EEG modality feature data in each feature data set of each source domain data at the last iteration number, the predicted domain labels of the EEG modality feature data in each feature data set of the target domain data, the predicted domain labels of the eye movement modality feature data in each feature data set of each source domain data at the last iteration number, and the predicted domain labels of the eye movement modality feature data in each feature data set of the target domain data; Determine the source domain data and target domain data at the previous iteration corresponding to the minimum distance to form the training set at the current iteration.
7. A multimodal cross-subject emotion recognition system, characterized by: include: Building module for constructing a multimodal cross-subject emotion recognition network; The multimodal cross-subject emotion recognition network includes: a cross-subject alignment module, a cross-modal alignment module and a multimodal fusion module connected in sequence; the cross-subject alignment module includes an EEG subject alignment submodule and an eye movement subject alignment submodule; the EEG subject alignment submodule includes a first shallow encoder, a first deep encoder and a first domain predictor connected to the first shallow encoder, and a first emotion label predictor connected to the first shallow encoder; the eye movement subject alignment submodule includes a second shallow encoder, a second deep encoder and a second domain predictor connected to the second shallow encoder, and a second emotion label predictor connected to the second shallow encoder; the cross-modal alignment module includes an EEG subject alignment submodule and an eye movement subject alignment submodule; the EEG subject alignment submodule includes a first shallow encoder, a first deep encoder and a first domain predictor connected to the first shallow encoder, and a second emotion label predictor connected to the second shallow encoder G modality alignment submodule and eye movement modality alignment submodule; the EEG modality alignment submodule includes a first self-attention encoder and a first momentum encoder both connected to the first shallow encoder; the eye movement modality alignment submodule includes a second self-attention encoder and a second momentum encoder both connected to the second shallow encoder; the multimodal fusion module includes a cross-attention layer connected to the first self-attention encoder and the second self-attention encoder and a first fusion branch and a second fusion branch both connected to the cross-attention layer; the first fusion branch and the second fusion branch both include a fully connected layer and a SoftMax function connected in sequence; the fully connected layer is connected to the cross-attention layer; An acquisition module is configured to acquire multiple source domain data, one target domain data, the true emotions corresponding to each feature data set of each source domain data, and the true emotions corresponding to each feature data set of the target domain data; each source domain data and the target domain data includes multiple feature data sets; a feature data set includes EEG modal feature data and eye movement modal feature data of a subject at the same time; a training module, configured to train the multimodal cross-subject emotion recognition network using all the source domain data and the target domain data to obtain a trained multimodal cross-subject emotion recognition network; A recognition module is used to construct an emotion recognition model based on a trained multimodal cross-subject emotion recognition network, and the emotion recognition model is used to perform emotion recognition; the emotion recognition model includes a first shallow encoder, a second shallow encoder, a first self-attention encoder, a second self-attention encoder, a cross-attention layer, a fully connected layer and a SoftMax function; the first self-attention encoder is connected to the first shallow encoder and the cross-attention layer, the second self-attention encoder is connected to the second shallow encoder and the cross-attention layer, and the cross-attention layer, the fully connected layer and the SoftMax function are connected in sequence.
8. The multimodal cross-subject emotion recognition system according to claim 7, characterized in that: The training module specifically includes: A cross-subject alignment module training unit is configured to, at the current iteration number, adopt a self-paced learning strategy to obtain a training set at the current iteration number based on the target domain data and all the source domain data at the previous iteration number, take the training set at the current iteration number as input, take the true emotions corresponding to each feature data set in the training set at the current iteration number and the true domain labels of each feature data set in the training set at the current iteration number as output, train the cross-subject alignment module with the minimum cross-loss function as the goal, obtain a cross-subject alignment module trained at the current iteration number, delete the training set at the current iteration number from all the source domain data at the previous iteration number, obtain all the source domain data at the current iteration number, and then update the iteration number to enter the next iteration until there is no source domain data at the current iteration number, then determine that the cross-subject alignment module trained at the last iteration number is the optimal cross-subject alignment module; the domain label is the source domain or the target domain; and the cross-loss function is obtained according to the cross-subject alignment module; a feature vector set determining unit, configured to input all the source domain data and the target domain data into the first shallow encoder and the second shallow encoder in the optimal cross-subject alignment module to obtain a feature vector set; the feature vector set comprising: the cross-subject aligned EEG modality feature vector corresponding to the EEG modality feature data in each feature data set of the source domain data, the cross-subject aligned EEG modality feature vector corresponding to the EEG modality feature data in each feature data set of the target domain data, the cross-subject aligned eye movement modality feature vector corresponding to the eye movement modality feature data in each feature data set of the source domain data, and the cross-subject aligned eye movement modality feature vector corresponding to the eye movement modality feature data in each feature data set of the target domain data; The cross-modal alignment module and the multimodal fusion module training unit are used to take the feature vector set as input, take the true emotions corresponding to each feature data set of each source domain data, the true emotions corresponding to each feature data set of the target domain data and the matching probability of each sample pair as output, and take the minimum total loss function as the goal to train the cross-modal alignment module and the multimodal fusion module to obtain the optimal cross-modal alignment module and the optimal multimodal fusion module; a sample pair includes an EEG modality encoding feature vector and an eye movement modality encoding feature vector; the EEG modality encoding feature vector and the eye movement modality encoding feature vector are obtained by inputting the feature vector set into the first self-attention encoder and the second self-attention encoder; if a If the EEG modality encoding feature vector and the eye movement modality encoding feature vector in a sample pair correspond to the same feature data, the sample pair matching probability of the sample pair is 1, otherwise the sample pair matching probability of the sample pair is 0; the total loss function is the sum of the contrast loss function, the matching loss function and the classification loss function; the contrast loss function is obtained according to the cross-modal alignment module; the matching loss function is obtained according to the cross-attention layer and the first fusion branch; the classification loss function is obtained according to the cross-attention layer, the second fusion branch, the first shallow encoder, the first deep encoder, the first emotion label predictor, the second shallow encoder, the second deep encoder and the second emotion label predictor.
9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform the multimodal cross-subject emotion recognition method according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that It stores a computer program, which, when executed by a processor, implements the multimodal cross-subject emotion recognition method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Cross-subject and cross-modal multi-modal nervous emotion recognition method and cross-subject and cross-modal multi-modal nervous emotion recognition system
CN114305415A
Electroencephalogram emotion recognition method and system based on dynamic convolution residual multi-source migration
CN115105076A