Multimodal Sentiment Recognition Method and Device Based on Hierarchical Interactive Alignment Network
Through the hierarchical interactive alignment network, the problems of individual differences and sample scarcity in multimodal emotion recognition are solved, and efficient emotion recognition under the condition of few samples is achieved.
Patent Information
- Application Number
- CN202510340474.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The existing multimodal emotion recognition methods rely on singlemodal signal sources, resulting in incomplete and accurate emotion analysis, large variability in individual sample distribution, scarce marker samples, and difficult to effectively identify complex emotional states.
A multimodal emotion recognition method based on a hierarchical interactive alignment network is adopted, and a multimodal emotion recognition model and a hierarchical representation distribution alignment layer is constructed, and EEG and eye movement data are used, combined with a hierarchical adaptive interactive attention module and a small sample learning module to achieve the fusion of cross-modal features and the elimination of inter-domain differences.
The performance and accuracy of multimodal emotion recognition are significantly improved under the condition of few samples, and can effectively deal with the problems of individual differences and limited sample size, dynamically capture multi-dimensional emotional characteristics, and improve the model's adaptability under different data distributions.
Smart Images

Figure CN119848794B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and particularly to a multi-modal emotion recognition method and device based on a hierarchical interaction alignment network. Background Art
[0002] With the in-depth research of neuroscience and the continuous improvement of the demand for understanding human emotions, multi-modal emotion recognition (MER) has gradually become a key research direction for promoting the development of human-computer interaction and emotion computing. Emotions play a crucial role in the human cognitive process and behavior regulation, involving individual subjective experiences and the complex regulation of the central and peripheral nervous systems. Therefore, evaluating emotional states by capturing and analyzing physiological signals can provide a more objective basis. However, most current emotion recognition methods rely heavily on single-modal signal sources, and this limitation poses challenges to the comprehensive analysis and accurate recognition of emotions. Therefore, multi-modal emotion recognition (MER) methods that fuse multiple physiological data have received increasing attention.
[0003] Electroencephalogram (EEG) signals, with the characteristics of non-invasiveness, economy, and high sampling rate, are widely used in detecting human emotions and can directly reflect neural activities during the emotion induction process. For this reason, many researchers have developed a series of emotion recognition models based on EEG data, such as multi-channel convolutional neural networks, bidirectional long short-term memory networks (BiLSTM), and graph-based networks. In addition, eye movement (EM) signals, as one of the physiological signals of the peripheral nervous system, can characterize the fixation behavior and pupil changes of individuals in different emotional states, thus providing supplementary information for the MER task.
[0004] Although significant progress has been made in MER, there are still many challenges, mainly in the following aspects: First, there is variability in the distribution of individual samples. Multi-modal physiological signals are non-stationary and show significant differences among different individuals. Second, the scarcity of labeled samples. Obtaining data with emotion labels is costly, and existing MER models have great difficulties in extracting emotion information using limited labeled samples. Summary of the Invention
[0005] The purpose of this application is to propose a multi-modal emotion recognition method and device based on a hierarchical interaction alignment network for the above-mentioned technical problems.
[0006] In the first aspect, the present invention provides a multi-modal emotion recognition method based on a hierarchical interaction alignment network, including the following steps:
[0007] Construct a multi-modal emotion recognition model and a hierarchical representation distribution alignment layer. The multi-modal emotion recognition model includes a feature extraction module, a hierarchical adaptive interaction attention module, and a few-shot learning module. The feature extraction module includes an electroencephalogram (EEG) sub-network and an eye movement sub-network. The hierarchical adaptive interaction attention module includes several dynamically connected interaction attention layers in sequence.
[0008] Obtain source domain data and target domain data. The source domain data includes pairs of EEG data and eye movement data of the source domain population and their corresponding true emotion category labels. The target domain data includes pairs of EEG data and eye movement data of the target domain population. Input the pairs of EEG data and eye movement data in the source domain data and the pairs of EEG data and eye movement data in the target domain data into the multi-modal emotion recognition model respectively. First, the EEG sub-network and the eye movement sub-network in the feature extraction module are used to extract the corresponding EEG depth features and eye movement depth features of the source domain data and the corresponding EEG depth features and eye movement depth features of the target domain data respectively. The corresponding EEG depth features and eye movement depth features of the source domain data and the corresponding EEG depth features and eye movement depth features of the target domain data are respectively input into the hierarchical adaptive interaction attention module. After passing through each layer of the dynamic interaction attention layer, the cross-modal features of each layer corresponding to the source domain data and the cross-modal features of each layer corresponding to the target domain data, as well as the final cross-modal features corresponding to the source domain data and the final cross-modal features corresponding to the target domain data, are output. The cross-modal features of each layer corresponding to the source domain data and the cross-modal features of each layer corresponding to the target domain data are input into the hierarchical representation distribution alignment layer to calculate the multi-layer maximum mean discrepancy loss function. Extract a support set and a query set from the source domain data. Determine the final cross-modal features corresponding to the support set and the final cross-modal features corresponding to the query set according to the final cross-modal features corresponding to the source domain data and input them into the few-shot learning module to obtain the probability values of each query sample in the query set belonging to each emotion category. Calculate the cross-entropy loss function according to the probability values of each query sample in the query set belonging to each emotion category and their corresponding true emotion category labels. Construct a total loss function based on the multi-layer maximum mean discrepancy loss function and the cross-entropy loss function. Use the total loss function to train the multi-modal emotion recognition model to obtain a trained multi-modal emotion recognition model.
[0009] Obtain a pair of EEG data and eye movement data of one of the persons to be recognized in the target domain population and input it into the trained multi-modal emotion recognition model. Pass through the feature extraction module and the hierarchical adaptive interaction attention module in sequence to obtain the final cross-modal features corresponding to the person to be recognized. The final cross-modal features corresponding to the person to be recognized and the final cross-modal features corresponding to the target domain data are input into the few-shot learning module to obtain the probability values of the person to be recognized belonging to each emotion category. Select the emotion category corresponding to the largest probability value as the predicted emotion category of the person to be recognized.
[0010] Preferably, the electroencephalogram (EEG) sub-network includes a residual block and a hybrid attention module connected in sequence, and the eye movement sub-network adopts a DenseNet structure;
[0011] The EEG data in the EEG data and eye movement data pair is input into the EEG sub-network to obtain the corresponding EEG depth features, as shown in the following formula:
[0012] ;
[0013] Wherein, represents the EEG data, represents the EEG depth features, represents the EEG sub-network;
[0014] The eye movement data in the EEG data and eye movement data pair is input into the eye movement sub-network to obtain the corresponding eye movement depth features, as shown in the following formula:
[0015] ;
[0016] Wherein, represents the eye movement data, represents the eye movement depth features, represents the eye movement sub-network.
[0017] Preferably, each dynamic interaction attention layer in the hierarchical adaptive interaction attention module includes an attention module, a gating module, a splicing layer, and a multi-layer perceptron;
[0018] The cross-modal features of the (i - 1)-th layer output by the dynamic interaction attention layer of the (i - 1)-th layer are input into the dynamic interaction attention layer of the i-th layer, and after linear transformation in the attention module, a query matrix is obtained, as shown in the following formula:
[0019] ;
[0020] Wherein, , represents the cross-modal features of the (i - 1)-th layer, represents the query matrix of the i-th layer, represents the query weight matrix of the i-th layer. When i = 1, the cross-modal features output by the dynamic interaction attention layer of the 0-th layer are the EEG depth features
[0021] The eye movement depth features are input into the dynamic interaction attention layer of the i-th layer, and after linear transformation in the attention module, a key matrix and a value matrix are obtained, as shown in the following formula:
[0022] ;
[0023] ;
[0024] Among them, and respectively represent the key weight matrix of the i-th layer and the value weight matrix of the i-th layer; and respectively represent the key matrix of the i-th layer and the value matrix of the i-th layer, represents the eye movement depth feature;
[0025] The attention feature of the i-th layer is calculated by the following formula:
[0026] ;
[0027] Among them, (⋅) represents function, d represents the feature dimension, T represents the transposed matrix, represents the attention feature of the i-th layer;
[0028] The attention feature of the i-th layer is input into the gating module, and the channel weight of the i-th layer is calculated through the gated attention mechanism, as shown in the following formula:
[0029] ;
[0030] Among them, σ(⋅) represents the sigmoid activation function, represents the gating weight, represents the channel weight of the i-th layer;
[0031] The gated selection of the attention feature of the i-th layer is performed through the channel weight of the i-th layer to obtain the selection feature of the i-th layer, as shown in the following formula:
[0032] ;
[0033] Among them, represents the selection feature of the i-th layer, and ⊙ represents element-wise multiplication;
[0034] The selection feature of the i-th layer is concatenated with the cross-modal feature of the i - 1-th layer and then input into the multi-layer perceptron to obtain the fusion feature of the i-th layer, as shown in the following formula:
[0035] ;
[0036] Among them, [⋅;⋅] represents the feature concatenation operation, represents the function corresponding to the multi-layer perceptron, represents the fusion feature of the i-th layer;
[0037] The fused features of the $i$-th layer and the cross-modal features of the $(i - 1)$-th layer are connected through a residual connection to obtain the cross-modal features of the $i$-th layer, as shown in the following formula:
[0038] ;
[0039] where, represents the cross-modal features of the $i$-th layer.
[0040] Preferably, in the hierarchical representation distribution alignment layer, the maximum mean discrepancy loss function is used to measure the inter-domain difference of the cross-modal features of each layer, as shown in the following formula:
[0041] ;
[0042] where, is the cross-modal feature of the $i$-th layer corresponding to one of the samples in the source domain data, is the set of, represents the cross-modal feature of the $i$-th layer corresponding to one of the samples in the target domain data, is the set of, the cross-modal features of each layer corresponding to the source domain data and the cross-modal features of each layer corresponding to the target domain data, represents the MMD loss of the $i$-th layer, represents the mapping function that maps the feature points to a reproducing kernel Hilbert space in, represents the reproducing kernel Hilbert space the square of the norm in;
[0043] Introduce a hierarchical weighting mechanism based on logarithmic decay:
[0044] ;
[0045] Secondly, represents the alignment weight of the $i$-th layer, and represent the weight reference parameter and the logarithmic growth parameter respectively;
[0046] The following formula is used to calculate the multi-layer maximum mean discrepancy loss function:
[0047] ;
[0048] where, $I$ represents the total number of layers of the dynamic interaction attention layer, represents the multi-layer maximum mean discrepancy loss function.
[0049] Preferably, the cross-modal features of the I-th layer output by the dynamic interaction attention layer of the last layer are used as the final cross-modal features; during the inference process of the trained multi-modal sentiment recognition model, the person to be recognized is used as a query sample in the query set, and the support set is extracted from the target domain data.
[0050] In the few-shot learning module, prototype vectors for each sentiment category are calculated based on the final cross-modal features corresponding to the support set, as shown in the following formula:
[0051] ;
[0052] where, represents one of the support samples in the support set that belongs to the c-th sentiment category, represents one of the support samples in the support set that belongs to the c-th sentiment category corresponding final cross-modal features, represents the set of all support samples in the support set that belong to the c-th sentiment category, represents the prototype vector of the c-th sentiment category, c = 1, 2, …, N;
[0053] The probability values for each query sample in the query set belonging to each sentiment category are calculated based on the final cross-modal features corresponding to the query set and the prototype vectors for each sentiment category, as shown in the following formula:
[0054] ;
[0055] where, represents one of the query samples in the query set, represents one of the query samples in the query set corresponding final cross-modal features, represents the prototype vector of the u-th sentiment category, represents the prototype vector of the v-th sentiment category, u = 1, 2, …, N, v = 1, 2, …, N, represents one of the query samples in the query set belonging to the probability value of the u-th sentiment category, represents the predicted sentiment category, represents the calculation function of the Euclidean distance, represents the exponential function with e as the base.
[0056] Preferably, the cross-entropy loss function is calculated using the following formula:
[0057] ;
[0058] where, represents the cross-entropy loss function, Denote one of the query samples in the query set The one-hot encoding of the true sentiment class label. If one of the query samples in the query set The true sentiment class label is u, then Otherwise ;
[0059] The total loss function is calculated using the following formula:
[0060] ;
[0061] Wherein Denotes the multi-layer maximum mean discrepancy loss function Denotes the total loss function Denotes the balancing weight
[0062] In a second aspect, the present invention provides a multi-modal sentiment recognition device based on a hierarchical interactive alignment network, comprising:
[0063] A model construction module configured to construct a multi-modal sentiment recognition model and a hierarchical representation distribution alignment layer. The multi-modal sentiment recognition model includes a feature extraction module, a hierarchical adaptive interactive attention module, and a few-shot learning module. The feature extraction module includes an electroencephalogram network and an eye movement network. The hierarchical adaptive interactive attention module includes several sequentially connected dynamic interactive attention layers;
[0064] A training module, configured to obtain source domain data and target domain data, where the source domain data includes pairs of electroencephalogram (EEG) data and eye movement data of the source domain population and their corresponding true emotion category labels, and the target domain data includes pairs of EEG data and eye movement data of the target domain population; input the pairs of EEG data and eye movement data in the source domain data and the pairs of EEG data and eye movement data in the target domain data into a multi-modal emotion recognition model respectively. First, the EEG sub-network and the eye movement sub-network in the feature extraction module are used to extract the corresponding EEG depth features and eye movement depth features of the source domain data and the corresponding EEG depth features and eye movement depth features of the target domain data respectively; the EEG depth features and eye movement depth features corresponding to the source domain data and the EEG depth features and eye movement depth features corresponding to the target domain data are respectively input into the hierarchical adaptive interaction attention module. After passing through each layer of dynamic interaction attention layer, the cross-modal features of each layer corresponding to the source domain data, the cross-modal features of each layer corresponding to the target domain data, the final cross-modal features corresponding to the source domain data, and the final cross-modal features corresponding to the target domain data are output; the cross-modal features of each layer corresponding to the source domain data and the cross-modal features of each layer corresponding to the target domain data are input into the hierarchical representation distribution alignment layer to calculate the multi-layer maximum mean discrepancy loss function; in the source domain data, a support set and a query set are extracted. According to the final cross-modal features corresponding to the source domain data, the final cross-modal features corresponding to the support set and the final cross-modal features corresponding to the query set are determined and input into the few-shot learning module to obtain the probability values of each query sample in the query set belonging to each emotion category. According to the probability values of each query sample in the query set belonging to each emotion category and their corresponding true emotion category labels, the cross-entropy loss function is calculated; based on the multi-layer maximum mean discrepancy loss function and the cross-entropy loss function, a total loss function is constructed, and the multi-modal emotion recognition model is trained using the total loss function to obtain a trained multi-modal emotion recognition model;
[0065] A prediction module, configured to obtain a pair of EEG data and eye movement data of one of the persons to be recognized in the target domain population and input it into the trained multi-modal emotion recognition model. Sequentially passing through the feature extraction module and the hierarchical adaptive interaction attention module, the final cross-modal features corresponding to the person to be recognized are obtained. The final cross-modal features corresponding to the person to be recognized and the final cross-modal features corresponding to the target domain data are input into the few-shot learning module to obtain the probability values of the person to be recognized belonging to each emotion category, and the emotion category corresponding to the largest probability value is selected as the predicted emotion category of the person to be recognized.
[0066] In a third aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any implementation manner of the first aspect.
[0067] Fourthly, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0068] Fifthly, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0069] Compared with the prior art, the present invention has the following beneficial effects:
[0070] (1) The multi-modal emotion recognition method based on the hierarchical interaction alignment network proposed by the present invention makes full use of the complementary features between electroencephalogram (EEG) data and eye movement (EM) data. Through the hierarchical adaptive interaction attention module and the cross-domain optimization strategy, the feature fusion of the two modal data is mapped into the common feature space to realize the effective recognition of complex emotional states, and can significantly improve the performance of multi-modal emotion recognition under the condition of few samples.
[0071] (2) The multi-modal emotion recognition method based on the hierarchical interaction alignment network proposed by the present invention effectively captures multi-dimensional emotion features through the dynamic interaction attention layer in the hierarchical adaptive interaction attention module, and fuses the cross-modal information of EEG data and eye movement data. The hierarchical representation distribution alignment layer is used to eliminate the inter-domain differences, and the complementary features are transformed into the common feature space, which significantly improves the accuracy of multi-modal emotion recognition under the condition of few samples, and can effectively cope with the problems of large individual differences and limited sample numbers in multi-modal emotion recognition.
[0072] (3) The multi-modal emotion recognition method based on the hierarchical interaction alignment network proposed by the present invention can dynamically capture the optimal interaction representation between modalities, adaptively reduce the modal feature differences, improve the adaptability of the model under different data distributions, and provide an efficient and robust solution for multi-modal few-shot emotion recognition. Description of the Drawings
[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0074] Figure 1 It is a schematic flowchart of the multi-modal emotion recognition method based on the hierarchical interaction alignment network for the embodiments of the present application;
[0075] Figure 2Schematic diagram of the multi-modal emotion recognition model and the hierarchical representation distribution alignment layer of the multi-modal emotion recognition method based on the hierarchical interaction alignment network according to the embodiment of the present application;
[0076] Figure 3 Schematic diagram of the electroencephalogram network of the multi-modal emotion recognition method based on the hierarchical interaction alignment network according to the embodiment of the present application;
[0077] Figure 4 Schematic diagram of the eye movement sub-network of the multi-modal emotion recognition method based on the hierarchical interaction alignment network according to the embodiment of the present application;
[0078] Figure 5 Schematic diagram of the hierarchical adaptive interaction attention module of the multi-modal emotion recognition method based on the hierarchical interaction alignment network according to the embodiment of the present application;
[0079] Figure 6 Schematic diagram of the multi-modal emotion recognition device based on the hierarchical interaction alignment network according to the embodiment of the present application;
[0080] Figure 7 Schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. Detailed implementation manners
[0081] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0082] Figure 1 The embodiment of the present application provides a multi-modal emotion recognition method based on a hierarchical interaction alignment network, including the following steps:
[0083] S1. Construct a multi-modal emotion recognition model and a hierarchical representation distribution alignment layer. The multi-modal emotion recognition model includes a feature extraction module, a hierarchical adaptive interaction attention module, and a few-shot learning module. The feature extraction module includes an electroencephalogram network and an eye movement sub-network. The hierarchical adaptive interaction attention module includes several dynamically connected interaction attention layers in sequence.
[0084] Specifically, refer to Figure 2, embodiments of the present application construct a multi-modal emotion recognition model and a hierarchical representation distribution alignment (HRDA) layer, where the hierarchical representation distribution alignment layer is used during the training process of the multi-modal emotion recognition model, and the two can form a hierarchical interactive alignment network, and the hierarchical representation distribution alignment layer is not used during the inference process. The multi-modal emotion recognition model includes a feature extraction module, a hierarchical adaptive interactive attention (HAIA) module, and a few-shot learning (FSL) module, where the feature extraction module includes an electroencephalogram sub-network for extracting electroencephalogram depth features and an eye movement sub-network for extracting eye movement depth features. Specifically, referring to Figure 3 and Figure 4 , the electroencephalogram data extracts electroencephalogram depth features through an electroencephalogram sub-network integrating a residual block and a convolutional block attention module (CBAM), and the eye movement data extracts eye movement depth features using an eye movement sub-network with a DenseNet network architecture. Among them, the residual block, the convolutional block attention module, and the DenseNet network architecture are all existing modules or networks, and their specific structures and calculation processes will not be elaborated here. The hierarchical adaptive interactive attention module includes several sequentially connected dynamic interactive attention (DIA) layers, which fuse the electroencephalogram depth features and the eye movement depth features through the dynamic interactive attention layers to extract cross-modal features. During the training process of the multi-modal emotion recognition model, the hierarchical representation distribution alignment layer is used to align the hierarchical cross-modal feature distributions of the source domain data and the target domain data to reduce the domain difference. Finally, through the few-shot learning module, prototype vector metric learning is performed for each emotion category to achieve emotion classification.
[0085] S2. Obtain source domain data and target domain data. The source domain data includes pairs of electroencephalogram (EEG) data and eye movement data of the source domain population and their corresponding true emotion category labels. The target domain data includes pairs of EEG data and eye movement data of the target domain population. Input the pairs of EEG data and eye movement data in the source domain data and the pairs of EEG data and eye movement data in the target domain data into the multi-modal emotion recognition model respectively. First, the EEG sub-network and the eye movement sub-network in the feature extraction module are used to extract the corresponding EEG depth features and eye movement depth features of the source domain data and the corresponding EEG depth features and eye movement depth features of the target domain data respectively. The corresponding EEG depth features and eye movement depth features of the source domain data and the corresponding EEG depth features and eye movement depth features of the target domain data are respectively input into the hierarchical adaptive interactive attention module. After passing through each layer of dynamic interactive attention layer, output the cross-modal features of each layer corresponding to the source domain data, the cross-modal features of each layer corresponding to the target domain data, the final cross-modal features corresponding to the source domain data, and the final cross-modal features corresponding to the target domain data. The cross-modal features of each layer corresponding to the source domain data and the cross-modal features of each layer corresponding to the target domain data are input into the hierarchical representation distribution alignment layer to calculate the multi-layer maximum mean discrepancy loss function. Extract the support set and query set from the source domain data. Determine the final cross-modal features corresponding to the support set and the final cross-modal features corresponding to the query set according to the final cross-modal features corresponding to the source domain data and input them into the few-shot learning module to obtain the probability values of each query sample in the query set belonging to each emotion category. Calculate the cross-entropy loss function according to the probability values of each query sample in the query set belonging to each emotion category and their corresponding true emotion category labels. Construct a total loss function based on the multi-layer maximum mean discrepancy loss function and the cross-entropy loss function, and use the total loss function to train the multi-modal emotion recognition model to obtain the trained multi-modal emotion recognition model.
[0086] In a specific embodiment, in the hierarchical representation distribution alignment layer, the maximum mean discrepancy loss function is used to measure the inter-domain difference of the cross-modal features of each layer, as shown in the following formula:
[0087] ;
[0088] Where, is the cross-modal feature of the i-th layer corresponding to one of the samples in the source domain data, is the set of, represents the cross-modal feature of the i-th layer corresponding to one of the samples in the target domain data, is the set of, the cross-modal features of each layer corresponding to the source domain data and the cross-modal features of each layer corresponding to the target domain data, represents the MMD loss of the i-th layer, Denote the mapping function that transforms feature points into a reproducing kernel Hilbert space in Denote the reproducing kernel Hilbert space the square of the norm in;
[0089] Introduce a hierarchical weighted mechanism based on logarithmic decay:
[0090] ;
[0091] Secondly, Denote the alignment weight of the i-th layer, and respectively denote the weight reference parameter and the logarithmic growth parameter;
[0092] Calculate the multi-layer maximum mean discrepancy loss function using the following formula:
[0093] ;
[0094] where I represents the total number of layers of the dynamic interaction attention layer, Denote the multi-layer maximum mean discrepancy loss function.
[0095] In a specific embodiment, calculate the cross-entropy loss function using the following formula:
[0096] ;
[0097] where, Denote the cross-entropy loss function, Denote one of the query samples in the query set the one-hot encoding of the true sentiment category label, if one of the query samples in the query set the true sentiment category label is u, then otherwise ;
[0098] Calculate the total loss function using the following formula:
[0099] ;
[0100] where, Denote the multi-layer maximum mean discrepancy loss function, Denote the total loss function, Denote the balance weight.
[0101] Specifically, the source domain data containing the pair of electroencephalogram data and eye movement data and the target domain data , where and respectively represent the The EEG data and eye movement data of the -th sample, which form a pair of EEG data and eye movement data in a set of source domain data, represent the true emotion category label of the -th sample in the source domain population, and respectively represent the EEG data and eye movement data of the -th sample of the target domain population, which form a pair of EEG data and eye movement data in a set of target domain data, and
[0102] are the total number of samples in the source domain population and the target domain population respectively. During the training process, the pair of EEG data and eye movement data in the source domain data and the pair of EEG data and eye movement data in the target domain data are respectively input into the multi-modal emotion recognition model. With the help of the feature extraction module in the hierarchical interaction alignment network, the dynamic interaction attention layer of the hierarchical adaptive interaction attention module, the hierarchical representation distribution alignment layer, and the few-shot learning module, the extraction of EEG deep features and eye movement deep features, the fusion of EEG deep features and eye movement deep features, the reduction of the inter-domain distribution difference, and the emotion recognition classification based on the prototype vector metric learning of emotion categories are completed in sequence. determines the reference size of the alignment weight, and the parameter affects the growth rate of the logarithmic function. Finally, the multi-layer maximum mean discrepancy loss function is calculated.
[0103] In the few-shot learning module, first calculate the prototype vector of each emotion category based on the cross-modal features corresponding to the support set extracted from the source domain data, and then classify the query samples in the query set extracted from the source domain data. Finally, the cross-entropy loss function is used as the classification loss function. Combining the cross-entropy loss function (emotion classification supervision) and the multi-layer maximum mean discrepancy loss function (domain alignment constraint) to construct the total loss function, and optimizing the parameters of the multi-modal emotion recognition model through gradient descent, minimizing the total loss function, to obtain a trained multi-modal emotion recognition model with optimized generalization performance.
[0104] S3. Obtain the electroencephalogram (EEG) data and eye movement data pair of one person to be recognized in the target domain population, and input it into the trained multi-modal emotion recognition model. Pass through the feature extraction module and the hierarchical adaptive interaction attention module in sequence to obtain the final cross-modal feature corresponding to the person to be recognized. The final cross-modal feature corresponding to the person to be recognized and the final cross-modal feature corresponding to the target domain data are input into the few-shot learning module to obtain the probability value of the person to be recognized belonging to each emotion category. Select the emotion category corresponding to the largest probability value as the predicted emotion category of the person to be recognized.
[0105] In a specific embodiment, the EEG sub-network includes a residual block and a hybrid attention module connected in sequence, and the eye movement sub-network adopts a DenseNet structure;
[0106] The EEG data in the EEG data and eye movement data pair is input into the EEG sub-network to obtain the corresponding EEG depth feature, as shown in the following formula:
[0107] ;
[0108] Where, represents the EEG data, represents the EEG depth feature, represents the EEG sub-network;
[0109] The eye movement data in the EEG data and eye movement data pair is input into the eye movement sub-network to obtain the corresponding eye movement depth feature, as shown in the following formula:
[0110] ;
[0111] Where, represents the eye movement data, represents the eye movement depth feature, represents the eye movement sub-network.
[0112] Specifically, deploy the trained multi-modal emotion recognition model. In the inference stage, input the EEG data and eye movement data pair of one person to be recognized in the target domain population into the trained multi-modal emotion recognition model. Among them, the EEG data passes through the EEG sub-network combined with a residual block and a hybrid attention module to extract the EEG depth feature; the eye movement data passes through the eye movement sub-network, using the DenseNet network architecture, to extract the eye movement depth feature.
[0113] In a specific embodiment, each dynamic interaction attention layer in the hierarchical adaptive interaction attention module includes an attention module, a gating module, a splicing layer, and a multi-layer perceptron;
[0114] The cross-modal features of the (i-1)-th layer output by the dynamic interaction attention layer of the (i-1)-th layer are input into the dynamic interaction attention layer of the i-th layer, and after a linear transformation in the attention module, a query matrix is obtained, as shown in the following formula:
[0115] ;
[0116] Among them, , represents the cross-modal features of the (i-1)-th layer, represents the query matrix of the i-th layer, represents the query weight matrix of the i-th layer. When i = 1, the cross-modal features output by the dynamic interaction attention layer of the 0-th layer are the electroencephalogram depth features
[0117] The eye movement depth features are input into the dynamic interaction attention layer of the i-th layer, and after a linear transformation in the attention module, a key matrix and a value matrix are obtained, as shown in the following formula:
[0118] ;
[0119] ;
[0120] Among them, and represent the key weight matrix of the i-th layer and the value weight matrix of the i-th layer respectively; and represent the key matrix of the i-th layer and the value matrix of the i-th layer respectively, represents the eye movement depth features;
[0121] The attention features of the i-th layer are calculated using the following formula:
[0122] ;
[0123] Among them, (⋅) represents function, d represents the feature dimension, T represents the transpose matrix, represents the attention features of the i-th layer;
[0124] The attention features of the i-th layer are input into the gating module, and the channel weights of the i-th layer are calculated through the gating attention mechanism, as shown in the following formula:
[0125] ;
[0126] Among them, σ(⋅) represents the sigmoid activation function, represents the gating weight, represents the channel weights of the i-th layer;
[0127] The channel weights of the i-th layer are used to gate-select the attention features of the i-th layer, obtaining the selected features of the i-th layer, as shown in the following formula:
[0128] ;
[0129] where, represents the selected features of the i-th layer, and ⊙ represents channel-wise multiplication;
[0130] The selected features of the i-th layer are concatenated with the cross-modal features of the (i - 1)-th layer and then input into a multi-layer perceptron to obtain the fused features of the i-th layer, as shown in the following formula:
[0131] ;
[0132] where, [⋅;⋅] represents the feature concatenation operation, represents the function corresponding to the multi-layer perceptron, represents the fused features of the i-th layer;
[0133] The fused features of the i-th layer and the cross-modal features of the (i - 1)-th layer are connected through a residual connection to obtain the cross-modal features of the i-th layer, as shown in the following formula:
[0134] ;
[0135] where, represents the cross-modal features of the i-th layer.
[0136] Specifically, through the collaborative processing of multi-modules for electroencephalogram depth features and eye movement depth features, emotion recognition and classification under few-shot conditions are realized, specifically including: using the dynamic interaction attention layer in the hierarchical adaptive interaction attention module to fuse electroencephalogram and eye movement data with interaction attention; referring to Figure 5, in the dynamic interaction attention layer of the first layer, the EEG depth features are input into the attention module. The linear layer is used to linearly project the EEG depth features to generate a query matrix. The linear layer is used to linearly project the eye movement depth features and map them to the feature space of the EEG depth features to generate a key matrix and a value matrix. The attention features of the first layer are calculated through the query matrix, the key matrix, and the value matrix. The channel weights of the first layer are calculated through the gated attention (GA) mechanism and the attention features of the first layer. The gated selection of the attention features of the first layer is performed using the channel weights of the first layer to obtain the selected features of the first layer. The selected features of the first layer and the EEG depth features are concatenated and then input into a multi-layer perceptron (MLP) to obtain the fused features of the first layer. The fused features of the first layer and the EEG depth features are then connected through a residual connection to obtain the cross-modal features of the first layer. Since the dynamic interaction attention layer retains the original EEG depth features through the residual connection, the dynamic interaction attention layer can be stacked into a multi-layer structure to enhance the multi-modal interaction ability. After stacking, the EEG depth features input to the next layer of the dynamic interaction attention layer are replaced by the cross-modal features output by the previous layer of the dynamic interaction attention layer. That is, the cross-modal features output by the previous layer of the dynamic interaction attention layer are input into the next layer of the dynamic interaction attention layer. First, the attention features of the next layer are calculated through the attention module. The channel weights of the next layer are calculated through the gated attention (GA) mechanism and the attention features of the next layer. The gated selection of the attention features of the next layer is performed using the channel weights of the next layer to obtain the selected features of the next layer. The selected features of the next layer and the cross-modal features output by the previous layer of the dynamic interaction attention layer are concatenated and then input into a multi-layer perceptron (MLP) to obtain the fused features of the next layer. The fused features of the next layer and the cross-modal features output by the previous layer of the dynamic interaction attention layer are then connected through a residual connection to obtain the cross-modal features of the next layer. Finally, the cross-modal features of the last layer are used as the final cross-modal features.
[0137] In a specific embodiment, the cross-modal features of the I-th layer output by the last layer of the dynamic interaction attention layer are used as the final cross-modal features; during the inference process of the trained multi-modal emotion recognition model, the person to be recognized serves as a query sample in the query set, and the support set is extracted from the target domain data;
[0138] In the few-shot learning module, the prototype vectors of each emotion category are calculated based on the final cross-modal features corresponding to the support set, as shown in the following formula:
[0139] ;
[0140] where represents one of the support samples in the support set belonging to the c-th emotion category, represents one of the support samples in the support set belonging to the c-th emotion category The corresponding final cross-modal features denotes the set of support samples in all support sets that belong to the c-th emotion category denotes the prototype vector of the c-th emotion category, c = 1, 2, …, N;
[0141] Based on the final cross-modal features corresponding to the query set and the prototype vectors of each emotion category, the probability values that each query sample in the query set belongs to each emotion category are calculated as follows:
[0142] ;
[0143] where denotes one of the query samples in the query set denotes one of the query samples in the query set The corresponding final cross-modal features denotes the prototype vector of the u-th emotion category denotes the prototype vector of the v-th emotion category, u = 1, 2, …, N, v = 1, 2, …, N denotes one of the query samples in the query set The probability value that it belongs to the u-th emotion category denotes the predicted emotion category denotes the calculation function of the Euclidean distance denotes the exponential function with base e
[0144] Specifically, during the inference process, the support set is extracted from the target domain data, and the person to be recognized is used as the query sample in the query set to calculate the probability values that the person to be recognized belongs to each emotion category in the few-shot learning module. Finally, the emotion category with the largest probability value is selected as the predicted emotion category corresponding to the person to be recognized, and the emotion recognition and classification are finally completed.
[0145] Data and processing: In the cross-subject experiment, the embodiments of the present application used two multi-modal emotion datasets with electroencephalogram (EEG) data and eye movement (EM) data: SEED and SEED-FRA.
[0146] SEED: The multi-modal version of this dataset was used in this study, which contains EEG signals from 12 participants and the corresponding emotion labels. After each participant watched a series of videos, they needed to annotate their emotional responses. The dataset contains three basic emotion categories: happy, sad, and angry. The EEG data was recorded by 62 channels with a sampling rate of 1000 Hz. To obtain eye movement data synchronized with the EEG data, SMI eye-tracking glasses were used to collect the corresponding eye movement data.
[0147] SEED-FRA: Eight native French subjects were recruited for this study and the experiment was conducted three times, with 21 stimulus material segments used each time.
[0148] In the experiment, the embodiments of this application adopt the leave-one-out cross-validation method to evaluate the performance. That is, one subject is alternately selected as the target domain population, and the remaining subjects are combined into a source domain population to evaluate the classification accuracy. All samples of the source domain population are used as the training set. At the same time, k labeled samples are randomly selected from each emotion category in the target domain population, that is, a total of N × k samples are selected as the training set, 300 samples are extracted from the remaining samples as the validation set, and the rest of the samples are used as the test set. In the meta-training stage, the support set and query set of each emotion category are both extracted from the source domain training set, and a small number of labeled samples in the target domain are only used to calculate the MMD loss. In the meta-validation and meta-test stages, the N * k labeled samples of the target domain training set are used as the support set, and the query set is extracted from the validation set or test set respectively.
[0149] For the setting of N-category k-samples, M is set to 3, and k is set to 1, 5, 10, 20. The embodiments of this application set the number of layers of the dynamic interaction attention layer to 3. The Adam optimizer is used to minimize the loss function. The learning rate is set to 1e-4. The λ in the loss function is dynamically adjusted as λ = , where p is the ratio of the current training epoch to the total number of training epochs, and θ is fixed at 10 during the training process. An early stopping scheme is adopted to save training time. The training is carried out for 50 epochs, and the number of tasks in each epoch is set to 20.
[0150] The trained multi-modal emotion recognition model (denoted as HIA-Net) proposed by the embodiments of this application will be systematically compared with the existing emotion recognition methods through experiments on the SEED and SEED-FRA datasets. Table 1 shows the performance of HIA-Net in the multi-modal emotion recognition task. The embodiments of this application use accuracy as the evaluation index, and the best results under different tasks are bolded to clearly show the performance differences of each model. Based on the reported experimental results, the following observations are summarized:
[0151] (1) HIA-Net shows the best emotion recognition effect in most cases. Under different datasets and different k-sample settings, HIA-Net is always superior to other comparison models. In the SEED dataset, the average emotion recognition accuracies of HIA-Net under different k-sample conditions are 86.60%, 95.85%, 98.34% and 99.00% respectively; while on the SEED-FRA dataset, the average emotion recognition accuracies of HIA-Net are 73.66%, 82.73%, 91.17% and 93.57% respectively. This series of results shows that HIA-Net demonstrates excellent ability in effectively combining EEG and EM modal information, significantly reducing the individual differences among subjects, and adaptively eliminating the differences between multi-layer interactive modal information, thus achieving higher accuracy in emotion recognition tasks.
[0152] By comparing the multi-modal emotion recognition results of HIA-Net and SDA-FSL, it is found that HIA-Net is significantly superior to SDA-FSL in most cases. This indicates that the hierarchical adaptive interactive attention module in HIA-Net can effectively combine the information of EEG and EM modalities, thus fitting the optimal fusion representation of modal interactive features.
[0153] Table 1:
[0154]
[0155] Further referring to Figure 6 , as an implementation of the methods shown in the above figures, this application provides an embodiment of a multi-modal emotion recognition device based on a hierarchical interactive alignment network. This device embodiment corresponds to the method embodiment shown in Figure 1 , and this device can be specifically applied to various electronic devices.
[0156] This application embodiment provides a multi-modal emotion recognition device based on a hierarchical interactive alignment network, including:
[0157] A model construction module 1, configured to construct a multi-modal emotion recognition model and a hierarchical representation distribution alignment layer. The multi-modal emotion recognition model includes a feature extraction module, a hierarchical adaptive interactive attention module, and a few-shot learning module. The feature extraction module includes an electroencephalogram network and an eye movement sub-network. The hierarchical adaptive interactive attention module includes several dynamically connected interactive attention layers in sequence;
[0158] The training module 2 is configured to obtain source domain data and target domain data. The source domain data includes pairs of electroencephalogram (EEG) data and eye movement data of the source domain population and their corresponding true emotion category labels. The target domain data includes pairs of EEG data and eye movement data of the target domain population. The pairs of EEG data and eye movement data in the source domain data and the pairs of EEG data and eye movement data in the target domain data are respectively input into the multi-modal emotion recognition model. First, the EEG sub-network and the eye movement sub-network in the feature extraction module are used to respectively extract the EEG depth features and eye movement depth features corresponding to the source domain data and the EEG depth features and eye movement depth features corresponding to the target domain data. The EEG depth features and eye movement depth features corresponding to the source domain data and the EEG depth features and eye movement depth features corresponding to the target domain data are respectively input into the hierarchical adaptive interaction attention module. After passing through each layer of dynamic interaction attention layer, the cross-modal features of each layer corresponding to the source domain data, the cross-modal features of each layer corresponding to the target domain data, the final cross-modal features corresponding to the source domain data, and the final cross-modal features corresponding to the target domain data are output. The cross-modal features of each layer corresponding to the source domain data and the cross-modal features of each layer corresponding to the target domain data are input into the hierarchical representation distribution alignment layer to calculate the multi-layer maximum mean discrepancy loss function. Support sets and query sets are extracted from the source domain data. The final cross-modal features corresponding to the support set and the final cross-modal features corresponding to the query set are determined according to the final cross-modal features corresponding to the source domain data and input into the few-shot learning module to obtain the probability values of each query sample in the query set belonging to each emotion category. The cross-entropy loss function is calculated according to the probability values of each query sample in the query set belonging to each emotion category and their corresponding true emotion category labels. A total loss function is constructed based on the multi-layer maximum mean discrepancy loss function and the cross-entropy loss function. The multi-modal emotion recognition model is trained using the total loss function to obtain the trained multi-modal emotion recognition model.
[0159] The prediction module 3 is configured to obtain a pair of EEG data and eye movement data of one of the persons to be recognized in the target domain population and input it into the trained multi-modal emotion recognition model. After passing through the feature extraction module and the hierarchical adaptive interaction attention module in sequence, the final cross-modal features corresponding to the person to be recognized are obtained. The final cross-modal features corresponding to the person to be recognized and the final cross-modal features corresponding to the target domain data are input into the few-shot learning module to obtain the probability values of the person to be recognized belonging to each emotion category. The emotion category corresponding to the largest probability value is selected as the predicted emotion category of the person to be recognized.
[0160] Figure 7 It is a schematic hardware structure diagram of the electronic device provided by the embodiment of the present invention. As Figure 7As shown in the figure, the electronic device of this embodiment includes: a processor 701 and a memory 702; wherein the memory 702 is used to store computer-executable instructions; the processor 701 is used to execute the computer-executable instructions stored in the memory to implement each step performed by the electronic device in the above embodiment. For details, please refer to the relevant descriptions in the foregoing method embodiments.
[0161] Optionally, the memory 702 can be either independent or integrated with the processor 701.
[0162] When the memory 702 is independently provided, the electronic device further includes a bus 703 for connecting the memory 702 and the processor 701.
[0163] This embodiment of the present invention also provides a computer storage medium, in which computer-executable instructions are stored. When the processor 701 executes the computer-executable instructions, the above method is implemented.
[0164] This embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by the processor 701, the above method is implemented.
[0165] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or modules can be in electrical, mechanical or other forms.
[0166] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.
[0167] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each module exists physically alone, or two or more modules are integrated in one unit. The units formed by the above modules can be implemented in the form of hardware or in the form of a hardware plus software functional unit.
[0168] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above-mentioned software functional modules stored in a storage medium include several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor 701 to execute some steps of the methods of various embodiments of the present application.
[0169] It should be understood that the above-mentioned processor 701 may be a central processing unit (CPU for short), or may also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor may be a microprocessor or the processor 701 may also be any conventional processor 701, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the execution of the hardware processor 701, or can be implemented by the combination of the hardware and software modules in the processor 701.
[0170] The memory 702 may include a high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.
[0171] The bus 703 may be an Industry Standard Architecture (ISA for short), a Peripheral Component Interconnect (PCI for short) bus, or an Extended Industry Standard Architecture (EISA for short) bus, etc. The bus 703 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus 703 in the drawings of the present application is not limited to only one bus 703 or one type of bus 703.
[0172] The above-mentioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0173] An exemplary storage medium is coupled to the processor 701, enabling the processor 701 to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 701. The processor 701 and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor 701 and the storage medium can also exist as discrete components in an electronic device or a master control device.
[0174] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0175] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-modal sentiment recognition method based on a hierarchical interactive alignment network, characterized in that, Including the following steps: Construct a multi-modal emotion recognition model and a hierarchical representation distribution alignment layer. The multi-modal emotion recognition model includes a feature extraction module, a hierarchical adaptive interactive attention module, and a few-shot learning module. The feature extraction module includes an electroencephalogram (EEG) sub-network and an eye movement sub-network. The hierarchical adaptive interactive attention module includes several dynamically interactive attention layers connected in sequence; Obtain source domain data and target domain data. The source domain data includes EEG data and eye movement data pairs of source domain population and their corresponding true emotion category labels. The target domain data includes EEG data and eye movement data pairs of target domain population. Input the EEG data and eye movement data pairs in the source domain data and the EEG data and eye movement data pairs in the target domain data into the multi-modal emotion recognition model respectively. First, the EEG deep features corresponding to the source domain data, the eye movement deep features corresponding to the source domain data, the EEG deep features corresponding to the target domain data, and the eye movement deep features corresponding to the target domain data are respectively extracted through the EEG sub-network and the eye movement sub-network in the feature extraction module. The EEG deep features corresponding to the source domain data, the eye movement deep features corresponding to the source domain data, the EEG deep features corresponding to the target domain data, and the eye movement deep features corresponding to the target domain data are respectively input into the hierarchical adaptive interactive attention module. Through each dynamically interactive attention layer, the cross-modal features of each layer corresponding to the source domain data, the cross-modal features of each layer corresponding to the target domain data, the final cross-modal features corresponding to the source domain data, and the final cross-modal features corresponding to the target domain data are output; The cross-modal features of each layer corresponding to the source domain data and the cross-modal features of each layer corresponding to the target domain data are input into the hierarchical representation distribution alignment layer to calculate the multi-layer maximum mean discrepancy loss function. Extract the support set and query set from the source domain data. Determine the final cross-modal features corresponding to the support set and the final cross-modal features corresponding to the query set according to the final cross-modal features corresponding to the source domain data and input them into the few-shot learning module to obtain the probability values of each query sample in the query set belonging to each emotion category. Calculate the cross-entropy loss function according to the probability values of each query sample in the query set belonging to each emotion category and their corresponding true emotion category labels. Construct a total loss function based on the multi-layer maximum mean discrepancy loss function and the cross-entropy loss function. Use the total loss function to train the multi-modal emotion recognition model to obtain a trained multi-modal emotion recognition model; Obtain the EEG data and eye movement data pair of one of the persons to be recognized in the target domain population and input it into the trained multi-modal emotion recognition model. Pass through the feature extraction module and the hierarchical adaptive interactive attention module in sequence to obtain the final cross-modal features corresponding to the person to be recognized. The final cross-modal features corresponding to the person to be recognized and the final cross-modal features corresponding to the target domain data are input into the few-shot learning module to obtain the probability values of the person to be recognized belonging to each emotion category. Select the emotion category corresponding to the largest probability value as the predicted emotion category of the person to be recognized.
2. The multimodal sentiment recognition method based on the hierarchical interaction alignment network according to claim 1, wherein The brain electronic network includes a residual block and a hybrid attention module connected in sequence, and the eye movement sub-network adopts a DenseNet structure; The electroencephalogram data in the electroencephalogram data and eye movement data pair is input into the brain electronic network to obtain corresponding electroencephalogram depth features, as shown in the following formula: ; Among them, represents electroencephalogram data, represents electroencephalogram depth features, represents electroencephalogram sub-networks; The eye movement data in the electroencephalogram data and eye movement data pair is input into the eye movement sub-network to obtain corresponding eye movement depth features, as shown in the following formula: ; Among them, represents eye movement data, represents eye movement depth features, represents the eye movement sub-network.
3. The multimodal sentiment recognition method based on the hierarchical interaction alignment network according to claim 1, characterized in that, Each dynamic interaction attention layer in the hierarchical adaptive interaction attention module includes an attention module, a gating module, a splicing layer, and a multi-layer perceptron; The cross-modal features of the (i - 1)-th layer output by the dynamic interaction attention layer of the (i - 1)-th layer are input into the dynamic interaction attention layer of the i-th layer, and after a linear transformation in the attention module, a query matrix is obtained, as shown in the following formula: ; Among them, , represents the cross-modal feature of the (i - 1)-th layer, represents the query matrix of the i-th layer, represents the query weight matrix of the i-th layer. When i = 1, the cross-modal feature output by the dynamic interaction attention layer of the 0-th layer is the EEG depth feature The eye movement depth features are input into the dynamic interaction attention layer of the i-th layer, and after a linear transformation in the attention module, a key matrix and a value matrix are obtained, as shown in the following formula: ; ; Among them, and respectively represent the key weight matrix of the i-th layer and the value weight matrix of the i-th layer; and respectively represent the key matrix of the i-th layer and the value matrix of the i-th layer, represents the eye movement depth feature; The following formula is used to calculate the attention features of the i-th layer: ; Among them, (⋅) represents a function, d represents the feature dimension, T represents the transposed matrix, denotes the attention feature of the i-th layer; The attention features of the i-th layer are input into the gating module, and the channel weights of the i-th layer are calculated through the gating attention mechanism, as shown in the following formula: ; where, σ(⋅) represents the sigmoid activation function, represents the gating weight, represents the channel weight of the i-th layer; The attention features of the i-th layer are gated and selected through the channel weights of the i-th layer to obtain the selected features of the i-th layer, as shown in the following formula: ; Among them, represents the selection feature of the i-th layer, and ⊙ represents channel-wise multiplication; The selected features of the i-th layer are spliced with the cross-modal features of the (i - 1)-th layer and then input into the multi-layer perceptron to obtain the fusion features of the i-th layer, as shown in the following formula: ; where [⋅;⋅] represents the feature concatenation operation, represents the function corresponding to the multi-layer perceptron, represents the fused feature of the i-th layer; The fusion features of the i-th layer and the cross-modal features of the (i - 1)-th layer are connected through a residual connection to obtain the cross-modal features of the i-th layer, as shown in the following formula: ; Among them, represents the cross-modal feature of the i-th layer.
4. The multimodal sentiment recognition method based on a hierarchical interaction alignment network according to claim 1, wherein In the hierarchical representation distribution alignment layer, the maximum mean discrepancy loss function is used to measure the inter-domain difference of the cross-modal features of each layer, as shown in the following formula: ; Among them, is the cross-modal feature of the i-th layer corresponding to one of the samples in the source domain data, is a set of, represents the cross-modal feature of the i-th layer corresponding to one of the samples in the target domain data, is a set of. The cross-modal features of each layer corresponding to the source domain data and the cross-modal features of each layer corresponding to the target domain data, represents the MMD loss of the i-th layer, represents the mapping function that transforms the feature points into a reproducing kernel Hilbert space in, represents the reproducing kernel Hilbert space the square of the norm in; Introduce a hierarchical weighting mechanism based on logarithmic decay: ; Secondly, represents the alignment weight of the i-th layer, and respectively represent the weight reference parameter and the logarithmic growth parameter; The following formula is used to calculate the multi-layer maximum mean discrepancy loss function: ; where I represents the total number of layers of the dynamic interactive attention layer, represents the multi-layer maximum mean discrepancy loss function.
5. The multimodal emotion recognition method based on a hierarchical interactive alignment network according to claim 1, wherein The cross-modal features of the I-th layer output by the dynamic interaction attention layer of the last layer are used as the final cross-modal features; During the inference process of the trained multi-modal emotion recognition model, the person to be recognized is used as a query sample in the query set, and a support set is extracted from the target domain data; In the few-shot learning module, prototype vectors for each emotion category are calculated based on the final cross-modal features corresponding to the support set, as shown in the following formula: ; Among them, represents one of the support samples in the support set that belongs to the c-th emotion category, represents one of the support samples in the support set that belongs to the c-th emotion category corresponding final cross-modal feature, represents the set of support samples in all support sets that belong to the c-th emotion category, represents the prototype vector of the c-th emotion category, c = 1, 2, …, N; Based on the final cross-modal features corresponding to the query set and the prototype vectors for each emotion category, the probability values of each query sample in the query set belonging to each emotion category are calculated, as shown in the following formula: ; Among them, represents one of the query samples in the query set, represents one of the query samples in the query set the corresponding final cross-modal feature, represents the prototype vector of the \(u\)-th emotion category, represents the prototype vector of the \(v\)-th emotion category, where \(u = 1, 2, \ldots, N\) and \(v = 1, 2, \ldots, N\), represents one of the query samples in the query set the probability value belonging to the \(u\)-th emotion category, represents the predicted emotion category, represents the calculation function of the Euclidean distance, represents the exponential function with base \(e\).
6. The multimodal sentiment recognition method based on a hierarchical interaction alignment network according to claim 5, characterized in that The following formula is used to calculate the cross-entropy loss function: ; Among them, represents the cross-entropy loss function, represents one of the query samples in the query set of the one-hot encoding of the true sentiment class label. If one of the query samples in the query set has a true sentiment class label of u, then , otherwise ; The following formula is used to calculate the total loss function: ; Among them, represents the multi-layer maximum mean discrepancy loss function, represents the total loss function, represents the balance weight.
7. A multi-modal sentiment recognition device based on a hierarchical interactive alignment network, characterized in that, Including: A model construction module configured to construct a multi-modal emotion recognition model and a hierarchical representation distribution alignment layer. The multi-modal emotion recognition model includes a feature extraction module, a hierarchical adaptive interaction attention module, and a few-shot learning module. The feature extraction module includes a brain electronic network and an eye movement sub-network. The hierarchical adaptive interaction attention module includes several layers of dynamically connected interaction attention layers in sequence; A training module, configured to obtain source domain data and target domain data, where the source domain data includes pairs of electroencephalogram (EEG) data and eye movement data of a source domain population and their corresponding true emotion category labels, and the target domain data includes pairs of EEG data and eye movement data of a target domain population; input the pairs of EEG data and eye movement data in the source domain data and the pairs of EEG data and eye movement data in the target domain data into the multi-modal emotion recognition model respectively, and first extract the EEG depth features and eye movement depth features corresponding to the source domain data and the EEG depth features and eye movement depth features corresponding to the target domain data through the EEG sub-network and the eye movement sub-network in the feature extraction module respectively; input the EEG depth features and eye movement depth features corresponding to the source domain data and the EEG depth features and eye movement depth features corresponding to the target domain data into the hierarchical adaptive interaction attention module respectively, and output the cross-modal features of each layer corresponding to the source domain data, the cross-modal features of each layer corresponding to the target domain data, the final cross-modal features corresponding to the source domain data, and the final cross-modal features corresponding to the target domain data through each layer of dynamic interaction attention layer; Input the cross-modal features of each layer corresponding to the source domain data and the cross-modal features of each layer corresponding to the target domain data into the hierarchical representation distribution alignment layer, and calculate the multi-layer maximum mean discrepancy loss function; extract a support set and a query set from the source domain data, determine the final cross-modal features corresponding to the support set and the final cross-modal features corresponding to the query set according to the final cross-modal features corresponding to the source domain data, and input them into the few-shot learning module to obtain the probability values of each query sample in the query set belonging to each emotion category, and calculate the cross-entropy loss function according to the probability values of each query sample in the query set belonging to each emotion category and their corresponding true emotion category labels; construct a total loss function based on the multi-layer maximum mean discrepancy loss function and the cross-entropy loss function, and use the total loss function to train the multi-modal emotion recognition model to obtain a trained multi-modal emotion recognition model; A prediction module, configured to obtain a pair of EEG data and eye movement data of one person to be recognized in the target domain population and input it into the trained multi-modal emotion recognition model, and successively pass through the feature extraction module and the hierarchical adaptive interaction attention module to obtain the final cross-modal features corresponding to the person to be recognized, input the final cross-modal features corresponding to the person to be recognized and the final cross-modal features corresponding to the target domain data into the few-shot learning module to obtain the probability values of the person to be recognized belonging to each emotion category, and select the emotion category corresponding to the largest probability value as the predicted emotion category of the person to be recognized.
8. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Video multi-modal emotion recognition method and device based on cross-modal dynamic convolution and computer equipment
CN114511906A
Multi-modal emotion recognition method, device and equipment and readable storage medium
CN115270849A