Multi-mode emotion recognition brain-computer interface instrument based on comparative learning gating network

By adopting a comparative learning gated network in multimodal emotion recognition technology, the problem of inflexible feature fusion and modal selection switching in the existing technology is solved, more efficient emotion recognition performance is achieved, and the user experience of human-computer interaction is improved.

CN120046000AActive Publication Date: 2025-05-27TIANJIN UNIV

Patent Information

Application Number
CN202510190843.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-27
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The existing multimodal emotion recognition technology is not flexible enough in feature fusion and modal selection switching, resulting in the failure to fully utilize all the valuable information of the modal in some samples, affecting the accuracy of emotion recognition.

Method used

A multimodal emotion recognition instrument based on a contrast learning gating network is adopted. Through the dual optimization of hardware and algorithms, a multimodal contrast learning gating network (MCGNet) is designed. This network includes a single-modal feature encoder pre-training module, a multimodal contrast learning training module and a gated structure training module. The comparison characterization learning and gated structure are used to enhance the representation and fusion of multimodal information.

Benefits of technology

Effectively integrate multi-scale features to improve the accuracy and robustness of emotion recognition, improve the performance of the model under different samples, and achieve a more accurate and natural human-computer interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046000A_ABST
    Figure CN120046000A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal emotion recognition instrument based on a comparative learning gating network, which comprises the following steps: firstly, designing hardware acquisition equipment through an electroencephalogram and an eye signal to ensure efficient and accurate acquisition of data; secondly, a multi-modal contrast learning gating network is designed, so that the emotion recognition accuracy is improved;
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of brain-computer interface, and in particular relates to a multimodal emotion recognition brain-computer interface instrument based on contrastive learning gated network. Background Art

[0002] With the development of computer and information technology, machines with emotion recognition capabilities can significantly improve the user experience of human-computer interaction and provide a smoother and more natural interface for human-computer interaction. In this context, emotion recognition has received widespread attention from academia and industry, and has been widely used in applications such as intelligent driving, medical care, and intelligent robots. The input methods used in existing emotion recognition research are generally divided into physiological signals and non-physiological signals. Among them, physiological signals have the advantage of not being affected by human subjective consciousness and can more truly and reliably reflect human emotional changes.

[0003] Emotion recognition research based on EEG and peripheral physiological signals has a long history. However, traditional deep learning feature extractors usually only focus on specific feature scales, such as single spatial, temporal or frequency domain features. This limitation makes the range of extracted features not comprehensive enough and may ignore the key multi-dimensional information in emotion recognition. Therefore, comprehensive features at multiple scales are crucial to improving the accuracy of emotion recognition. Effective integration of these multi-scale features can help capture more subtle emotional changes, thereby improving the performance and robustness of the model.

[0004] Traditional emotion recognition methods mainly rely on single-modal data, but human emotion expression is diverse, and different emotional information is contained in different forms of physiological signals, so it is necessary to combine multiple physiological signals to express emotions. By utilizing the differences and complementarities of different modalities, more comprehensive emotional information can be obtained.

[0005] In existing multimodal emotion recognition, a common method is to simply concatenate or fuse features from different modalities. This simple feature fusion method may not be able to fully capture the complex relationship between the modalities, thus affecting the performance of the model. In order to effectively fuse multi-source data features, some studies use contrastive learning methods to learn the similarities and differences between different modalities to extract and align semantic information in different modal data.

[0006] In addition, in multimodal tasks, each modality has unique advantages and information. How to flexibly select the optimal modality according to the specific situation and effectively switch between the dominant modality and multimodality is an important challenge. Existing multimodal emotion recognition methods are not flexible enough in modality selection and switching, resulting in failure to fully utilize all valuable information of the modality in some samples.

[0007] It can be seen that in multimodal emotion recognition, more attention should be paid to the diversity and complementarity of features, and more efficient dynamic decision-making mechanisms should be explored to achieve a more accurate and natural human-computer interaction experience. Summary of the invention

[0008] In order to overcome the problem of inaccurate multimodal emotion recognition in the prior art, the present invention provides a multimodal emotion recognition instrument based on a contrastive learning gated network, which comprehensively improves the performance of emotion recognition through dual optimization of hardware and algorithm. The present invention first designs hardware acquisition equipment in combination with electroencephalogram (EEG) and eye signals to ensure efficient and accurate data acquisition. Secondly, a multimodal contrastive learning gated network (MCGNet) is designed to improve the accuracy of emotion recognition.

[0009] The present invention relates to a multimodal emotion recognition brain-computer interface instrument based on a contrastive learning gated network, which is composed of an EEG data acquisition device, an eye movement data acquisition device and a multimodal data processing device. The multimodal data processing device includes an MCU, a neural network accelerator, a data bus, a data storage device and an output interface. The EEG data acquisition device is used to collect EEG data, and the eye movement data acquisition device is used to collect eye movement data. After the collection, the EEG data and eye movement data are stored in the data storage device under the control of the MCU. Furthermore, the MCU controls the neural network accelerator to take out the EEG data and eye movement data from the data storage device for training the multimodal contrastive learning gated network.

[0010] Furthermore, the EEG data acquisition device is composed of a power supply, a main control chip and an analog-to-digital converter.

[0011] Furthermore, the eye movement data acquisition device is composed of a binocular infrared camera, which can capture clear eye images in low-light or no-light environments to avoid interference from visible light.

[0012] In order to verify the actual effect of the multimodal acquisition helmet in emotion recognition, the multimodal emotion recognition method MCGNet combining EEG signals and eye movement signals is implemented in the multimodal emotion recognition brain-computer interface instrument proposed in the present invention. Specifically, the MCU controls the neural network accelerator to take out the EEG data and eye movement data from the data storage device for training the multimodal contrast learning gating network. The training includes the following three modules: single-modal feature encoder pre-training module, multimodal contrast learning training module and gating structure training module.

[0013] The unimodal feature encoder pre-training module inputs the EEG primary features and the eye movement primary features into the EEG feature encoder and the eye movement feature encoder, respectively, and inputs the encoding results into the classifier to predict the unimodal emotion category respectively.

[0014] The multimodal contrastive learning training module inputs the EEG primary features and the eye movement primary features into the feature encoder of the unimodal feature encoder pre-training module respectively to obtain the multimodal features EEG and Eye. Then, the EEG and Eye signals are decomposed by using the EEG similarity mapper and EEG dissimilarity mapper, the eye similarity mapper and eye dissimilarity mapper respectively, and the decomposition results are fused and subjected to intra-sample contrastive learning and inter-sample contrastive learning.

[0015] The gated structure training module fuses the EEG primary features and the eye movement primary features and inputs them into the gated network. The gated network generates a two-dimensional vector as output to select and activate different expert network branches to complete the emotion classification task.

[0016] The present invention uses contrastive representation learning and contrastive feature decomposition to enhance the representation of multimodal information. In addition, in order to solve the problem that the fusion of multimodal information may sometimes introduce noise interference and cause the accuracy to be lower than that of a single modality, a gating structure is introduced to select a single modality or multimodality and its corresponding expert network according to sample characteristics to improve the accuracy of emotion recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0018] Figure 1 Diagram of the multimodal emotion recognition brain-computer interface instrument system based on contrastive learning gating network.

[0019] Figure 2 Diagram of the pre-training module for a unimodal feature encoder.

[0020] Figure 3 Diagram of the multimodal contrastive learning training module.

[0021] Figure 4 This is a diagram of the gated structure training module.

[0022] Figure 5 These are the experimental results of the present invention on the TJU-Emotion dataset.

[0023] Figure 6These are the experimental results of the present invention on the public dataset SEED.

[0024] Figure 7 Graph showing the performance of the present invention in ablation experiments.

[0025] Figure 8 This is a confusion diagram of the EEG modality and the eye modality of the present invention.

[0026] Fig. 9 It is the experimental result diagram of single-modality EEG, eye and two multi-modality fusion methods of the present invention.

[0027] Fig.10 It is the two-dimensional projection result of four decomposition features of the test sample of the present invention. DETAILED DESCRIPTION

[0028] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present invention. However, it should be clear to those skilled in the art that the present invention may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present invention.

[0029] like Figure 1 As shown, the instrument consists of an EEG data acquisition device, an Eye data acquisition device and a multimodal data processing device. Among them, the multimodal data processing device includes an MCU, a neural network accelerator, a data bus, a data storage device and an output interface. The EEG data acquisition device is used to collect EEG data, and the eye movement data acquisition device is used to collect eye movement data. After the collection, the EEG data and eye movement data are stored in the data storage device under the control of the MCU. Furthermore, the MCU controls the neural network accelerator to take out the EEG data and eye movement data from the data storage device for training the multimodal contrast learning gated network.

[0030] Furthermore, the EEG data acquisition device is composed of a power supply, a main control chip and an analog-to-digital converter.

[0031] Furthermore, the eye movement data acquisition device is composed of a binocular infrared camera, which can capture clear eye images in low-light or no-light environments to avoid interference from visible light.

[0032] In order to verify the actual effect of the multimodal acquisition helmet in emotion recognition, the multimodal emotion recognition method MCGNet combining EEG signals and eye movement signals is implemented in the multimodal emotion recognition brain-computer interface instrument proposed in the present invention. Specifically, the MCU controls the neural network accelerator to take out the EEG data and eye movement data from the data storage device for training the multimodal contrast learning gating network. The training includes the following three modules: single-modal feature encoder pre-training module, multimodal contrast learning training module and gating structure training module.

[0033] The unimodal feature encoder pre-training module inputs the EEG primary features and the eye movement primary features into the EEG feature encoder and the eye movement feature encoder, respectively, and inputs the encoding results into the classifier to predict the unimodal emotion category respectively.

[0034] The multimodal contrastive learning training module inputs the EEG primary features and the eye movement primary features into the feature encoder of the unimodal feature encoder pre-training module respectively to obtain the multimodal features EEG and Eye. Then, the EEG and Eye signals are decomposed by using the EEG similarity mapper and EEG dissimilarity mapper, the eye similarity mapper and eye dissimilarity mapper respectively, and the decomposition results are fused and subjected to intra-sample contrastive learning and inter-sample contrastive learning.

[0035] The gated structure training module fuses the EEG primary features and the eye movement primary features and inputs them into the gated network. The gated network generates a two-dimensional vector as output to select and activate different expert network branches to complete the emotion classification task.

[0036] The single-mode feature encoder pre-training module constructed by the present invention is as follows: Figure 2 As shown. The pre-training module includes an EEG modality pre-training module and an eye movement modality pre-training module. The EEG modality pre-training module consists of an EEG primary feature extraction module, an EEG feature encoder DSSTNet module, and a classifier module connected in sequence; the eye movement modality pre-training module consists of an eye movement primary feature extraction module, an eye movement feature encoder module, and a classifier module connected in sequence.

[0037] Specifically, in the EEG modality pre-training module, in the pre-training stage, given a sample, the EEG primary feature extraction module is used to extract the EEG features of the sample data, and the features are input into the EEG feature encoder DSSTNet module, and the output of the EEG feature encoder DSSTNet module is input into the classifier module for predicting the category of unimodal emotions. The EEG feature encoder DSSTNet module includes a spatial feature extractor, a frequency band attention module and a time domain feature extractor connected in sequence, the spatial feature extractor is used to extract the spatial features of the EEG signal, the frequency band attention module is used to extract the frequency domain features, and the time domain feature extractor is used to extract the time domain features.

[0038] Specifically, the original EEG signal is preprocessed and then input into the spatial feature extractor, which consists of a graph converter and an adaptive graph convolution layer. In the graph converter, a convolution kernel with a size of 64×1 and a step size of 1×1 is first used to convolve along the electrode channel dimension to obtain the spatial features of the EEG signal. The purpose of setting the convolution kernel size larger is to obtain a wider receptive field, so as to fully extract the spatial features. This layer uses a padding operation to maintain the input and output feature dimensions unchanged. Then the obtained spatial features are represented as graph structure data.

[0039] Since the distribution of EEG channels is irregular and discrete, an adaptive graph convolution layer is used to extract the spatial features of the EEG signal channel dimension. In graph theory, a graph can be defined as G = (v, e), where v and e represent nodes and edges in the graph, respectively. The connection relationship between N different nodes in the graph can be represented by the adjacency matrix A∈R N×N Let the filter function g θ =diag(θ), for a given spatial signal x∈R N The graph convolution operation can be defined as:

[0040] g θ * G x=Udiag(θ)U T x (1)

[0041] Among them, θ is a variable parameter, * G Represents the graph convolution operation, U∈R N×N is the Laplace matrix L∈R in graph G N×N In the actual calculation process, g θ It is difficult to calculate directly, so the K-order Chebyshev polynomial is used to approximate g θ The adaptive graph convolution layer can automatically update the adjacency matrix during the back-propagation process without manual construction. First, the adjacency matrix is ​​randomly initialized, and the optimal adjacency matrix A is adaptively learned when the back-propagation operation is applied. *When the loss function is calculated, A * The partial derivative of A * Update according to the following rules:

[0042]

[0043] Wherein ρ is a parameter used to control the updating speed of the adjacency matrix, and ρ=0.001 is set in the present invention.

[0044] This spatial feature extractor solves the problem of difficulty in effectively extracting spatial features of EEG signals, avoids intervening in the adjacency matrix, and can adaptively update the adjacency matrix for specific subjects.

[0045] Frequency band attention module: This module is located between the spatial feature extractor and the temporal feature extractor, and includes three parts: semi-global pooling, local cross-band interaction, and adaptive weighting. First, semi-global pooling is used to perform global average pooling of the spatial dimension while retaining the temporal dimension information. Then, a 7×1 convolution kernel is used to extract frequency domain features, and a padding operation is used to maintain the temporal features and electrode channel dimensions of the feature map. Compared with the fully connected layer, the convolution layer reduces the number of parameters and promotes local interaction. Finally, the filtered features are adaptively weighted, and the frequency domain information is probabilized using the softmax function, and the softmax output is Hadamard-producted with the input feature.

[0046] Time domain feature extractor: The time domain feature extraction module is constructed by stacking multiple one-dimensional convolutional layers. The first layer consists of D convolution kernels of size 1×2 and step size 1×1, which are convolved along the time domain dimension. The second layer of convolution consists of D kernels of size 32×1 and step size 1×1, the purpose of which is to change the spatial dimension of the feature to 1 to better extract the time domain information. Subsequently, the batch normalization method is used to alleviate the problem of gradient vanishing and speed up the model training process. The ELU activation function is then used to nonlinearize the data. The third layer is average pooled along the time domain dimension, with a convolution kernel size of 1×1 and a step size of 1×2. This pooling layer achieves the effect of preventing overfitting and reducing computational complexity while smoothing the time features. The last layer is composed of D kernels of size 1×2 and step size 1×1 to further extract the time domain features. In the present invention, the hyperparameter D is set to 40.

[0047] Finally, the EEG feature map is passed through a classifier module consisting of two fully connected layers, and its output is calculated using the softmax function to obtain the predicted probability of each of the three emotions.

[0048] Specifically, in the eye movement modality pre-training module, in the pre-training stage, given a sample, the eye movement primary feature extraction module is used to extract the eye feature of the sample data, and input the feature into the eye movement feature encoder module. The output of the eye movement feature encoder module is input into the classifier module for predicting the category of unimodal emotions.

[0049] After preprocessing the original eye signal, the primary feature differential entropy features are extracted in the frequency bands of [0, 0.2] Hz, [0.2, 0.4] Hz, [0.4, 0.6] Hz, and [0.6, 1] Hz respectively; the features extracted from the left and right eyes include the mean, standard deviation, and differential entropy features of the four frequency bands, which are used as the input of the eye movement feature encoder module, denoted as F Eye ∈R T×K ; The sliding window length is T and the feature dimension is K.

[0050] The eye movement feature encoder module adopts the structural design of the encoder in the transformer architecture, including input embedding layer, multi-head attention, layer normalization, and feedforward layer. The advantage of transformer is that it can simultaneously process long-distance dependencies and capture global context information, and effectively learn the intrinsic structure and important features in the input sequence through the self-attention mechanism.

[0051] The multimodal contrastive learning training module constructed by the present invention is as follows: Figure 3 In this module, the EEG primary features and eye movement primary features are input into the encoder of the unimodal feature encoder pre-training module respectively to obtain the multimodal features EEG and Eye. Then, the multimodal feature EEG is decomposed into similar feature EEG using the EEG similarity mapper and EEG dissimilarity mapper. s and dissimilar features EEG d , use the eye similarity mapper and eye dissimilarity mapper to decompose the multimodal feature Eye into similar features Eye s and dissimilar features Eye d ; Then, the four decomposed features are fused in different splicing methods and intra-sample contrast learning and inter-sample contrast learning are performed; the present invention provides two splicing methods, the first method is to splice according to the feature dimension, and output the splicing result to the MLP classifier to obtain the emotion prediction label (i.e. Figure 3 The y in 1 ); The second method is to concatenate them according to the batch dimension and input them into the weight-sharing MLP classifier to obtain 2 Represents the four unimodal prediction labels.

[0052] In the first approach, the dataset is denoted as M. For a given sample i∈M, the cross entropy (CE) loss is used to calculate the multimodal sentiment classification loss L: pred , calculated as follows:

[0053]

[0054] Where [;] means splicing according to feature dimension, y i is the true label of the emotion category.

[0055] In the second approach, the unimodal sentiment classification loss L uni for:

[0056]

[0057]

[0058] Where [,] means concatenation according to batch dimension, y i Represents the true label of the unimodal prediction, and the four decomposed features are predicted separately.

[0059] In the intra-sample contrastive learning and inter-sample contrastive learning, the four decomposed features are used to construct positive and negative sample pairs, such as Figure 3 As shown. First, for the sample pair (i, j) given by the dataset M, the feature encoder extracts features and calculates the sample [EEG i ;Eye i ] and [EEG j ;Eye j ]:

[0060] C i,j =sim([EEG i ;Eye i ],[EEG j ;Eye j ]) (7)

[0061] Next, select similar samples and dissimilar samples for sample i. i The same samples are sorted from low to high according to the cosine similarity scores calculated in the previous step to construct a set of candidate similar samples. At the same time, the label non-y i The samples are divided into candidate dissimilar sample sets From the collection Randomly select two samples with higher cosine similarity scores to form a positive pair with sample i, denoted as Neighbor i ; From the collection Randomly select two samples with lower cosine similarity scores, denoted as Randomly select two samples with higher cosine similarity scores, denoted as and Together with sample i, it forms an inter-sample negative pair Outlier i .

[0062] First, construct the in-sample positive and negative pair

[0063]

[0064] Where j∈Neighbor i ∪Outlier i , Neighbor i and Outlier i They represent similar samples and dissimilar samples of sample i respectively.

[0065] Then, construct the positive alignment between samples and negative pair The construction is as follows:

[0066]

[0067] Where j∈Neighbor i , k∈Outlier i

[0068] The loss function of intra-sample contrastive learning and inter-sample contrastive learning is represented by the joint contrastive loss, which includes two aspects: the contrast between similar samples and dissimilar samples between samples, and the contrast between similar features and dissimilar features within samples. Given a sample i, the contrastive learning loss L c for:

[0069]

[0070] In the NT-Xent contrastive learning loss framework, the contrastive learning between samples and samples is combined to perform modal feature decomposition and modal representation learning. The loss of sample i is expressed as:

[0071]

[0072] Where (a, p) and (a, n) represent a pair of decomposed eigenvectors in the sample, for example Or a pair of decomposed feature vectors between samples, such as P i is a positive set, expressed as Including the sample Aligned with the sample N i is a negative pair set, expressed as Including in-sample negative pairs and negative pairs between samples (a,p) is P i The positive pair in (a,n) is N i The negative pair in .

[0073] The total loss function of this module is:

[0074] L=L pred +λ uni L uni +λ c L c (14)

[0075] Where L pred is the multimodal sentiment classification loss, L uni is the unimodal sentiment classification loss, L c is the contrastive learning loss. uni and λ c Determines the contribution of each task to the update of model parameters during training.

[0076] The gated structure training module constructed by the present invention is as follows Figure 4 As shown in this module, the EEG primary feature x 1 and eye primary feature x 2 The concatenation and subsequent input into a gated network G(x) consisting of an MLP is shown. The gated network generates a two-dimensional vector g as output to select and activate different network branches to complete the emotion classification task. A group of expert networks are selected based on the mixed expert model. Each expert specializes in a subset of all modalities. In a specific task, there can be up to three expert networks for the two modalities, denoted as E 1 (x 1 ), E 2 (x 1 ,x 2 ), E 3 (x 2 ). EEG can provide useful clues when combined with eye, but in practical applications, single-modal EEG is sometimes better at recognizing emotions. The present invention uses the single-modal EEG emotion classification network model pre-trained by the single-modal feature encoder pre-training module as the expert network-E 1 (x 1 ), the multimodal contrastive learning network model trained in the multimodal contrastive learning training module is used as the expert network II 2 (x 1 ,x 2 ). The final output prediction label is y3 It is expressed as:

[0077] y 3 =g 1 E 1 (x 1 )+g 2 E 2 (x 1 ,x 2 ) (15)

[0078] The two expert networks selected in the present invention have different model complexities. Generally, the expert network with a high model complexity has a stronger representation ability. If the network is trained only by minimizing the loss of the emotion classification task, the gating network will always select the branch with a high model complexity. Therefore, an additional loss function β[g 1 C(E 1 )+g 2 C(E 2 )]. Where g 1 and g 2 represents the decision vector output by the gating structure, C(E) represents the cost of executing the expert network (e.g., MAdds), and β is a hyperparameter used to adjust the relative importance between the two loss terms.

[0079] The present invention jointly optimizes the expert network and the gating network in an end-to-end manner. For the logic vector output by the gating structure MLP, we use a softmax function with a temperature parameter to process it and adjust the sensitivity of the decision boundary. In addition, the present invention involves selecting different expert networks for subsequent training so that each network focuses on a specific mode of the data.

[0080] The present invention uses a self-built dataset (TJU-Emotion) and a publicly available SEED dataset to evaluate the performance of the proposed method. Experimental results show that the performance of the proposed method in emotion recognition tasks is better than a series of state-of-the-art comparative methods.

[0081] First, the present invention constructs a three-category multimodal emotion dataset, named TJU-Emotion dataset, which contains EEG and eye multimodal data. The experiment was conducted at Tianjin University, and a total of 10 subjects were invited to participate in the experiment. The EEG modality uses 32 electrodes for the experiment, and the eye data collects the pupil diameter data of the left and right eyes of the subjects. The experiment uses 15 movie clips (which can cause positive, neutral and negative emotions) as stimulus sources, and the experiment contains 15 trials.

[0082] The preprocessing steps of the EEG part include filtering, downsampling, re-reference conversion, segmentation and artifact removal. The preprocessed EEG data will be divided into samples with a sliding window length of T = 15s and a step size of ΔT = 1s. The preprocessing of the eye part uses principal component analysis to remove the first principal component of light reflection.

[0083] The publicly available SEED dataset contains EEG data of 15 subjects watching 15 Chinese movie clips with negative, positive and neutral emotions. The EEG electrode distribution adopts the international 10-20 electrode positioning scheme, with a total of 62 electrode channels.

[0084] A total of four quantitative and qualitative effect analyses were conducted on the present invention: comparative experiments with the most advanced methods, the impact of various structures on model performance, multimodal information complementarity analysis, and comparative learning visualization analysis.

[0085] The present invention implements a subject-dependent experiment and calculates the average classification accuracy, standard deviation and F1-score of all subjects. The experiment sets the learning rate of the Adam optimizer to 1e-5 and 1e-3 respectively in two data sets, and the parameter λ uni , c and β are both set to 0.1.

[0086] The present invention selects six baseline models to compare the performance of the present invention, including simple decision-level fusion methods, maximum rule MAX and sum rule SUM, a model fusion method based on fuzzy integral, a deep canonical correlation analysis method DCCA-AM with attention mechanism, a method BAT using bidirectional adapters to transfer complementary information between different modalities, and a method CARAT for contrastive learning and feature reconstruction.

[0087] First, we perform unimodal emotion classification experiments on 10 subjects in the TJU-Emotion dataset using a unimodal feature encoder, and compare the performance of the proposed method using the six baseline models in the previous section. Figure 5 .

[0088] From the first two columns in the figure, we found that the performance of unimodal emotion classification depends on the type of data used, and the EEG modality performs best in the unimodality. After combining multimodal information, the model of the present invention significantly improves the performance of emotion classification. Then, a comparative experiment of 6 multimodal fusion methods was conducted on the data set, and the t-test significance p-values ​​of several models relative to the present invention were calculated. The classification performance of the present invention is significantly different from that of other methods. The experimental results show that the present invention not only has significant advantages in classification accuracy, but also shows great advantages in model stability. In addition, the average F1-score of the present invention is improved compared with other methods.

[0089] We also validate the performance of our proposed method on the widely used public dataset SEED. The experimental results are shown in Figure 6 ,The results show that our proposed method is also competitive on public datasets.

[0090] In order to explore the impact of various structures on the performance of the present invention, we set up several ablation experiments, and the results are as follows: Figure 7 As shown. The following ablation experiment models are included: a model without contrastive learning method, a model using only intra-sample contrastive learning method, a model using only inter-sample contrastive learning method, and a model without gating structure.

[0091] For the TJU-Emotion dataset, Figure 7 It can be seen that the performance of the four ablation experimental models is not as good as the model of the present invention. Compared with the model that does not use the contrastive learning method, the model that uses the intra-sample contrastive learning method or the inter-sample contrastive learning method has obvious performance improvement. In addition, the model of the present invention has a more flexible structure and can dynamically select different modalities and expert networks according to the characteristics and needs of each sample. Similar conclusions were obtained for the SEED dataset.

[0092] Figure 8 Confusion plots for the EEG modality and the eye modality are shown. In the eye modality, the recognition accuracy of the positive state is higher than in the EEG modality. The EEG modality is better at recognizing neutral and negative states. Therefore, it is expected that the two modalities can complement each other and improve the performance of recognizing each emotional state.

[0093] Fig. 9 The confusion matrices of single-modality EEG, eye, and two multimodal fusion methods: the model without gating structure and the model of the present invention are shown, showing the advantages and disadvantages of different modalities and different modal fusion methods. We found that EEG and eye modalities have important complementary features, and the multimodal fusion method significantly improves the classification performance. Overall, the effective combination of multiple modalities can utilize the characteristics and complementarity of each modality to improve the ability to recognize three types of emotions.

[0094] A subject is randomly selected in TJU-Emotion, and the two-dimensional projections of the four decomposition features of all test samples are as follows: Fig.10 As shown, Fig.10 (a) No contrastive learning method is used, Fig.10 (b) The contrastive learning method is used. As can be seen from the figure, after using the contrastive learning method, similar features (EEG s and Eye s ) become closer, while dissimilar features (EEG d and Eye d) are getting further and further away from their corresponding similar features. This demonstrates the effectiveness of using contrastive learning methods to learn the consistency and inconsistency between modalities.

[0095] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0096] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0097] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0098] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0099] In the embodiments provided by the present invention, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0100] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0101] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0102] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. . Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0103] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A multimodal emotion recognition instrument based on contrastive learning gated network, characterized in that: The instrument comprises an EEG data acquisition device for acquiring EEG data, an eye movement data acquisition device for acquiring eye movement data, and a multimodal data processing device; The multimodal data processing device includes an MCU, a neural network accelerator, a data bus, a data storage device, and an output interface; the MCU controls the neural network accelerator to retrieve EEG data and eye movement data from the data storage device for training the multimodal contrastive learning gating network, specifically including: The unimodal feature encoder pre-training module inputs the EEG primary features and the eye movement primary features into the EEG feature encoder and the eye movement feature encoder, respectively, and inputs the encoding results into the classifier to predict the unimodal emotion category respectively; The multimodal contrastive learning training module inputs the EEG primary features and the eye movement primary features into the feature encoder of the unimodal feature encoder pre-training module respectively to obtain the multimodal features EEG and Eye. Then, the EEG similarity mapper and the EEG dissimilarity mapper, the eye similarity mapper and the eye dissimilarity mapper are used to decompose the EEG and eye movement signals respectively, and the decomposition results are fused and subjected to intra-sample contrastive learning and inter-sample contrastive learning. The gated structure training module fuses the EEG primary features and the eye movement primary features and inputs them into the gated network. The gated network generates a two-dimensional vector as output to select and activate different expert network branches to complete the emotion classification task.

2. The multimodal emotion recognition instrument based on contrastive learning gated network as claimed in claim 1, wherein the unimodal feature encoder pre-training module is composed of an EEG modality pre-training module and an eye movement modality pre-training module, the EEG modality pre-training module is composed of an EEG primary feature extraction module, an EEG feature encoder DSSTNet module and a classifier module connected in sequence; the eye movement modality pre-training module is composed of an eye movement primary feature extraction module, an eye movement feature encoder module and a classifier module connected in sequence; The EEG feature encoder DSSTNet module consists of a spatial feature extractor, a frequency band attention module, and a temporal feature extractor connected sequentially.

3. The multimodal emotion recognition apparatus based on contrastive learning gating network according to claim 1, wherein the multimodal contrastive learning training module specifically comprises: Decomposing multimodal feature EEG into similar feature EEG using EEG similarity mapper and EEG dissimilarity mapper s and dissimilar features EEG d ; The multimodal feature Eye is decomposed into similar features Eye using the eye similarity mapper and eye dissimilarity mapper. s and dissimilar features Eye d ; The four decomposed features are fused in the feature dimension or batch dimension, and intra-sample contrastive learning and inter-sample contrastive learning are performed.

4. The multimodal emotion recognition instrument based on contrastive learning gated network as claimed in claim 1, wherein in the gated structure training module, the unimodal EEG emotion classification network model pre-trained by the unimodal feature encoder pre-training module is used as the expert network 1 E1(x1), and the multimodal contrastive learning network model trained in the multimodal contrastive learning training module is used as the expert network 2 E2(x1, x2), and the final output prediction label is represented by y3 as: y3=g1E1(x1)+g2E2(x1,x2) (1) Where g1 and g2 represent the decision vectors output by the gating structure.

5. In the multimodal emotion recognition instrument based on contrastive learning gated network as described in claim 2, the spatial domain feature extractor is used to extract the spatial domain features of the EEG signal, the frequency band attention module is used to extract the frequency domain features, and the time domain feature extractor is used to extract the time domain features.

6. The multimodal emotion recognition instrument based on contrastive learning gated network as described in claim 2, wherein the eye movement feature encoder module adopts the structural design of the encoder in the transformer architecture, including an input embedding layer, multi-head attention, layer normalization, and a feedforward layer.

7. The multimodal emotion recognition apparatus based on contrastive learning gated network as claimed in claim 3, wherein when fusion is performed according to feature dimension, the data set is recorded as M, and for a given sample i∈M, the multimodal emotion classification loss L is calculated using cross entropy loss pred , calculated as follows: Where [;] means splicing according to feature dimension, y i is the true label of the emotion category; MLP is a multi-layer perceptron.

8. The multimodal emotion recognition apparatus based on contrastive learning gated network as claimed in claim 3, when fused according to the batch dimension, the single-modal emotion classification loss L uni for: Where [,] means concatenation according to batch dimension, y i Represents the true label of unimodal prediction; MLP is a multi-layer perceptron.

9. The multimodal emotion recognition apparatus based on contrastive learning gated network according to claim 3, wherein in the intra-sample contrastive learning and inter-sample contrastive learning, the four decomposed features are used to construct positive and negative sample pairs, specifically including: For a given sample pair (i, j) in the dataset M, the feature encoder extracts features and calculates the sample [EEG i ;Eye i ] and [EEG j ;Eye j ]: C i,j =sim([EEG i ;Eye i ],[EEG j ;Eye j ]) (6) Select similar samples and dissimilar samples for sample i; i The same samples are sorted from low to high according to the cosine similarity scores calculated by formula (6) to construct a set of candidate similar samples. At the same time, the label is not y i The samples are divided into candidate dissimilar sample sets From the collection Randomly select two samples with higher cosine similarity scores to form a positive pair with sample i, denoted as Neighbor i ; From the collection Randomly select two samples with lower cosine similarity scores, denoted as Randomly select two samples with higher cosine similarity scores, denoted as and Together with sample i, it forms an inter-sample negative pair Outlier i ; First, construct the in-sample positive and negative pair Where j∈Neighbor i ∪Outlier i , Neighbor i and Outlier i Respectively represent similar samples and dissimilar samples of sample i; Constructing positive alignments between samples and negative pair The construction is as follows: Where j∈Neighbor i , k∈Outlier i .

10. The multimodal emotion recognition instrument based on contrastive learning gated network as described in claim 1, wherein the EEG data acquisition device is composed of a power supply, a main control chip and an analog-to-digital converter; and the eye movement data acquisition device is composed of a binocular infrared camera.

Citation Information

Patent Citations

  • Dual-mode emotion identification method and system based on facial expression and eyeball movement

    CN105868694A

  • Multi-modal physiological signal emotion calculation method based on mixed Bi-LSTM

    CN117297604A

  • Multi-modal feature fusion emotion recognition method based on gating cross-attention mechanism

    CN117370828A

  • Multi-modal emotion recognition method and system based on wearable device

    CN117520826A

  • Emotion perception intelligent glasses based on physiological and non-physiological multi-modal data fusion

    CN117992832A

Cited By

  • Three-dimensional eye movement data emotion change analysis method based on deep learning

    CN121774522A