Emotional state detection method based on multi-modal information fusion

By using pre-trained models and multimodal bottleneck fusion network, data heterogeneity and feature fusion redundancy in multimodal emotion recognition are solved, and efficient and accurate information fusion and accurate emotion recognition are achieved.

CN120337114APending Publication Date: 2025-07-18SOUTHWEST UNIVERSITY FOR NATIONALITIES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510223319.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

There are problems in multimodal emotion recognition with strong heterogeneity of data sets, insufficient representation ability of feature extraction method, and redundant feature fusion calculation, which affects the performance and accuracy of the model.

Method used

The pre-trained model is used to extract text, audio and video features, feature alignment and dynamic weight allocation are realized through the information aggregation module, and a multimodal bottleneck fusion network is introduced, and a multimodal bottleneck fusion network is used to improve cross-modal information interaction, reducing computing complexity and noise interference.

Benefits of technology

It improves the accuracy and efficiency of multimodal emotion recognition, effectively integrates multimodal information, reduces computational complexity and noise interference, and improves the representation learning efficiency and accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337114A_ABST
    Figure CN120337114A_ABST
Patent Text Reader

Abstract

The invention relates to an emotional state detection method based on multi-modal information fusion, and the method comprises the steps: firstly employing a pre-training model as a feature extractor, and extracting the multi-modal feature expression capability; then feature alignment and dynamic weight distribution in a multi-modal emotion recognition task are achieved through an information aggregation module, the contribution degree of each modal data to a classification result is accurately evaluated through dynamic weight distribution, and then complementarity information between modals is more fully utilized to suppress noise; and finally, modal information integration and compression are carried out through a multi-modal bottleneck network, a bottleneck unit is introduced, a model is forced to carry out information interaction through the unit when cross-modal interaction is carried out, cross-modal information interaction is carried out on the basis, and more efficient and accurate information fusion can be realized. According to the method, the problems of insufficient representation capability of a traditional feature extraction method, high heterogeneity of multi-modal data, calculation redundancy during multi-modal feature fusion and the like in multi-modal emotion recognition are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and in particular, to an emotion state detection method based on multimodal information fusion. Background Art

[0002] Emotion recognition technology aims to achieve a more accurate human emotion recognition effect by deeply analyzing and fusing emotion information of multiple modalities. As an important branch of emotion recognition, multimodal emotion recognition captures the unique emotion expression characteristics of each modality and utilizes the information complementarity between modalities to improve the accuracy, stability, and applicability of emotion recognition. Currently, multimodal emotion recognition research mainly focuses on two aspects: multimodal feature representation and multimodal feature fusion. Multimodal emotion recognition aims to comprehensively utilize information from different modalities to more comprehensively understand and recognize human emotion states. Multimodal learning jointly models various modality data and establishes a shared representation space by changing the expression of these modality information to better correlate information between different modalities.

[0003] In domestic research, Liu Tianbao et al. proposed a CNN-LSTM model embedded with an attention mechanism to learn the audiovisual representation of emotional content contained in facial expressions and auditory signals. Lan et al. proposed a deep generalized canonical correlation analysis method with an attention mechanism, which has made new progress in the field of MER. Zhang et al. proposed a fuzzy weighted regression support vector machine method, which emphasizes the complementarity between multimodals and considers the possible negative impact between audio and visual signals. Liu Jiamin et al. proposed a new type of emotion response (ER) system based on multimodal information such as facial expressions and electroencephalogram signals.

[0004] In foreign research, Rupauliha et al. proposed a multimodal emotion recognition method based on a neural network architecture, which fuses multimodal information such as body language, gestures, facial expressions, and speech data. Lei and Cao introduced a preference learning framework that simultaneously considers prediction and actual labels to solve the problem of label imbalance in emotion recognition datasets. In foreign research, Nemati et al. proposed a hybrid fusion method for audio and visual modality emotion recognition based on a latent space linear graph. Hsu et al. proposed a bimodal Transformer encoder trained based on signal-level and segment-level emotion labels. This method improves the emotion recognition model by incorporating time series information. Chaudhari et al. proposed a unique fusion method based on Transformer and attention. This method incorporates multimodal self-supervised learning features into multimodal emotion classification, greatly improving the model performance.

[0005] Multi-modal feature representation is the basis and prerequisite for multi-modal feature fusion. If multi-modal data cannot be fully represented in terms of features, it will greatly affect the effect of multi-modal feature fusion. In addition, there are numerous multi-modal feature representation methods, and the feature representation methods for different modalities vary greatly, resulting in extremely high heterogeneity of data in different modalities, making it very difficult to construct a unified and effective feature representation. Zadeh et al. designed a multi-modal fusion model to learn the discriminative and generative representations of each modality. Tsai et al. used a stacked Transformer network to perform soft alignment on multi-modal data in the time series. Wang et al. used three bidirectional gated recurrent units to capture the context information in each modality of data.

[0006] Multi-modal feature fusion is the subsequent step of multi-modal feature representation. Data of multiple modalities not only have specific information for emotional expression, but there is also complementary information for emotional expression between the information of different modalities. Therefore, how to give full play to the specificity and complementarity of multi-modal information is an important topic in multi-modal feature fusion. The common methods of multi-modal feature fusion mainly include feature-level fusion, decision-level fusion, and model-level fusion. Feature-level fusion is the earliest multi-modal fusion method. For example, the tensor fusion network designed by Zadeh et al. is a typical example of feature-level fusion, which has a high computational complexity. At the same time, the model structure is very complex, requiring more parameters and a more complex training process. Decision-level fusion is a method of fusing the decision results of different modalities. For example, the low-rank matrix fusion method designed by Liu et al. is a typical example of decision-level fusion, which has problems such as high algorithm complexity and limited performance upper bound. With the development of deep learning technology, model-level fusion has received much attention in multi-modal emotion recognition tasks. For example, Rahman et al. directly fused word vectors with audio and visual information, changing the position of words in the semantic space to achieve multi-modal information fusion. Hazarika et al. divided the information of each modality into modality-invariant information and modality-specific information, and combined the Transformer architecture for fusion. Yuan et al. used the idea of masking, randomly masked the input modality segments, and made the model more focused on the key segments through the attention mechanism.

[0007] Deficiencies of existing technical solutions:

[0008] 1. The datasets in the field of multi-modal emotion recognition are highly heterogeneous

[0009] Multimodal representation learning aims to map information from different modalities into a unified representation space to facilitate subsequent sentiment classification tasks. However, existing multimodal representation learning methods are usually rather rough and fail to effectively distinguish between valid and invalid information within each modality. This leads to the model potentially confusing important features and noise when fusing information from different modalities, affecting the performance and accuracy of the model. In addition, the heterogeneity among multiple modality information also greatly impacts the multimodal learning performance of subsequent models. Therefore, it is necessary to design new multimodal representation learning methods that can better distinguish valid information within each modality and effectively map highly heterogeneous multimodal information into the same representation space, thereby improving the efficiency and accuracy of the model's representation learning.

[0010] 2. Insufficient representation ability of feature extraction methods

[0011] In multimodal sentiment recognition, data from different modalities have different features and expression forms. Traditional feature extraction methods are often limited to surface features and cannot fully capture the deep semantic information within the modality. In addition, the scale of multimodal sentiment datasets is usually small and the class distribution is uneven. Directly applying traditional feature extraction methods may not be able to extract features with rich semantic information. Therefore, it is necessary to find new feature extraction methods that can extract effective feature information for each modality when the dataset is small.

[0012] 3. Computational redundancy problem in feature fusion

[0013] In multimodal sentiment recognition, effectively fusing information from different modalities is a crucial task. The conventional approach is to simply concatenate multimodal data and then feed it into the model. This method may lead to an increase in computational complexity and the introduction of redundant information. In addition, there may be differences in the scales and feature representation methods between different modality data, which also increases the complexity of the fusion process. Therefore, it is necessary to design more effective multimodal fusion methods that can make full use of the complementary information between modalities and improve the efficiency and accuracy of the fusion process. Summary of the Invention

[0014] Aiming at the deficiencies of the existing technology, an emotion state detection method based on multimodal information fusion, the method includes: constructing a state detection network including an information aggregation module and a multimodal bottleneck fusion network, using three pre-trained models of text, speech, and video frames as feature extractors to improve the multimodal feature expression ability, the information aggregation module realizes feature alignment and dynamic weight allocation in the multimodal sentiment recognition task, suppressing noise, the multimodal bottleneck fusion network introduces a bottleneck unit, enabling the state detection network to effectively organize and compress information within each modality, and finally taking the mean of the sub-classification results of the three modalities as the final classification result, the detection method specifically includes:

[0015] Step 1: The original text, audio waveform, and video frames are respectively passed through three pre-trained models, namely the Bert model, Wav2vec2.0 model, and Fab-Net model, to extract features corresponding to each modality. The extracted features contain temporal information and context semantic information, specifically including the context encoding representation of the original text, the context representation of the original audio waveform, and the video feature representation of the original video frames;

[0016] Step 11: Split the input original text into a series of word tokens, add a [CLS] token at the beginning of the split sentence for aggregating information for classification tasks, and add a [SEP] token at the end of the sentence to indicate the end of the sentence. Each word token is converted into an embedding representation through an embedding layer, and the embedding representation is input into the Bert pre-trained model for processing to generate the corresponding context encoding representation;

[0017] Step 12: Input the original audio waveform into the Wav2vec2.0 pre-trained model, process the original audio waveform through a convolutional neural network to extract hidden latent speech representations, and then input the latent speech representations into a masked Transformer model. The masked Transformer model performs context processing on the latent speech representations to generate a context representation containing global semantics and temporal dependencies;

[0018] Step 13: Use the Fab-Net pre-trained model to extract features from the video frames. Input the start frame and the end frame together into the encoder to extract features, generating a start frame embedding representation and an end frame embedding representation. Merge the embedding representations of the start frame and the end frame through a concatenation operation to form a joint feature representation containing temporal and spatial relationships. Input the joint feature representation into the decoder, and the decoder processes it to obtain a high-dimensional representation containing temporal features and spatial information. Finally, use the output of the last hidden layer of the Fab-Net pre-trained model as the video feature, and perform a linear transformation through a fully connected layer to obtain the video feature representation;

[0019] Step 2: Implement feature alignment and dynamic weight allocation in the multi-modal emotion recognition task for the multi-modal features output by the pre-trained models through the information aggregation module. The information aggregation module includes a feature alignment unit and a modality-level weight dynamic allocation unit, including:

[0020] Step 21: Using the text data as the standard, map the time steps of the audio data and facial expression data to the time steps of the text data through the long short-term memory network of the feature alignment unit, and then apply the Softmax function to the output of the long short-term memory network to obtain the relative position probability distribution P of the audio data and facial expression data relative to the text data in terms of time position;

[0021] Transpose the relative position probability distribution P to obtain the transposed probability distribution P T , and multiply the transposed probability distribution P T by the audio data and facial expression data of the original input in batch matrix multiplication to obtain the aligned multi-modal data X m , realizing the alignment of audio data and facial expression data with text data in terms of time steps;

[0022] Apply co-attention operation to the multi-modal data X m , apply attention weights to the data of each modality to obtain the aggregated modality features f M ;

[0023] Step 22: Input the modality features f M into the modality-level weight dynamic allocation unit for dynamic weight allocation of each modality feature. The weight dynamic allocation unit includes an encoder network E M , a classifier and a confidence classifier T M , including:

[0024] Input the modality features f M into the encoder network E M for dynamic estimation to obtain their respective corresponding modality weight matrices W M ;

[0025] Construct a corresponding classifier for each modality Classifier Process the modality weight matrix W M to obtain the prediction distribution p M ;

[0026] Input the prediction distribution p M into the confidence classifier T M to obtain the confidence of the current modality, and multiply the confidence of the current modality by the features of the current modality to achieve weight allocation for the current modality, and finally obtain the multi-modal features F of the current modality M ;

[0027] Step 3: The multi-modal bottleneck fusion network consists of a video modality encoder, an audio modality encoder, a text modality encoder, a bottleneck unit, and an average fusion module. Among them, the video modality encoder includes multiple stacked video attention modules, the audio modality encoder includes multiple stacked audio attention modules, and the text modality encoder includes multiple stacked text attention modules. Input the video multi-modal features, audio multi-modal features, and text multi-modal features in the multi-modal features F M into the video modality encoder, audio modality encoder, and text modality encoder for processing;

[0028] Step 31: For the multi-modal features F of each modality M , the shallow feature information of the current modality is processed by the attention module before the bottleneck unit;

[0029] Step 32: Combine the shallow feature information with the bottleneck unit to obtain the fused feature information;

[0030] Step 33: Feed the fused feature information combined with the bottleneck unit into the encoder layer behind the bottleneck unit of the current modality. Finally, extract the output of the last layer of each modality, pass through the fully connected layer respectively and apply the Softmax activation function to obtain the sub-classification results of the three modalities. Input the sub-classification results into the average fusion module for averaging to obtain the final classification result.

[0031] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0032] 1. The present invention uses pre-trained models such as wav2vec2.0, BERT, and FAb-Net that have been deeply trained with a large amount of data to extract features from the text, audio, and video data of the present invention. By fine-tuning these pre-trained models, features with rich semantic information can be extracted from each modality, effectively making up for the limitations of the small data volume of the multi-modal emotion recognition dataset and the insufficient expressiveness of modal features of traditional feature extraction methods.

[0033] 2. The present invention designs an information aggregation module including a feature alignment and aggregation unit and a weight dynamic allocation unit. Align multi-modal heterogeneous data through the long short-term memory network; distinguish the information in the modality that contributes more to the classification result through the collaborative attention mechanism, and use the weight dynamic allocation unit to allocate weights for the data of multiple modalities at the modality level to improve the integration efficiency of information between modalities.

[0034] 3. Introduce a multi-modal bottleneck fusion network improved based on Transformer. By introducing a bottleneck unit, force the model to perform information interaction through this unit when performing cross-modal interaction, so that the model can effectively organize and compress information within each modality, and then perform cross-modal information interaction on this basis, aiming to achieve more efficient and accurate information fusion, thereby reducing the computational complexity and noise interference. Brief Description of the Drawings

[0035] Figure 1 is the overall structural schematic diagram of the emotion state detection network proposed by the present invention;

[0036] Figure 2 is the structural schematic diagram of the BERT preprocessing model;

[0037] Figure 3It is a schematic structural diagram of the Word2vec preprocessing model;

[0038] Figure 4 It is a schematic structural diagram of the Fab-Net preprocessing model;

[0039] Figure 5 It is a schematic structural diagram of the information aggregation module of the present invention;

[0040] Figure 6 It is a schematic diagram of the processing flow of the feature alignment unit in the information aggregation module of the present invention,

[0041] Figure 7 It is a schematic diagram of the process of the modality-level weight dynamic allocation unit in the information aggregation module of the present invention;

[0042] Figure 8 It is a schematic structural diagram of the multi-modal bottleneck fusion network of the present invention. Detailed implementation manners

[0043] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the specific implementation manners and with reference to the accompanying drawings. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.

[0044] The Token in the present invention refers to a word token.

[0045] The multi-modal sentiment recognition technology has many advantages, but the traditional feature extraction methods are insufficient in capturing the deep semantic information within modalities, the representation learning methods are difficult to effectively process heterogeneous data, and there are problems such as increased computational complexity and introduction of redundant information in the feature fusion process. The present invention mainly solves the problems of insufficient representation ability of traditional feature extraction methods, strong heterogeneity of multi-modal data, and computational redundancy in multi-modal feature fusion in multi-modal sentiment recognition. Although the multi-modal sentiment recognition technology has many advantages, the current feature extraction and representation learning methods still face challenges in capturing context semantic information, processing data heterogeneity, and avoiding fusion redundancy. In view of these challenges, the present invention explores and develops a multi-modal bottleneck fusion network. By introducing a bottleneck unit, the model can effectively organize and compress information within each modality, and on this basis, perform cross-modal information interaction to achieve more efficient and accurate information fusion. At the same time, the present invention also solves the problem of multi-modal data heterogeneity based on the feature extraction method and information aggregation module of transfer learning. In addition, a multi-modal bottleneck fusion network based on an improved Transformer is explored to achieve effective fusion of multi-modal features.

[0046] Figure 1It is a schematic structural diagram of the emotional state detection network proposed by the present invention. The emotional state detection method of the present invention will be described in detail below with reference to the accompanying drawings. As Figure 1 shown, the emotional state detection network mainly consists of an information aggregation module and a multi-modal bottleneck fusion network.

[0047] The present invention uses a pre-trained model to extract facial expression features of text, audio, and video. The information aggregation module (Information Aggregation Module, IAM), and at the same time, based on the improved Transformer to reduce redundancy. The multi-modal bottleneck aggregation network proposed by the present invention suppresses noise through the information aggregation module, maps heterogeneous features to a unified space, and dynamically allocates weights. Subsequently, the information is input into the improved Transformer model for learning.

[0048] Step 1: Respectively extract the features of the corresponding modalities of the original text, audio, and video frames through three pre-trained models: the Bert model, the Wav2vec2.0 model, and the Fab-Net model. The extracted features contain temporal information and context semantic information, specifically including the context encoding vector of the original text, the context representation of the original audio, and the video feature representation of the original video frame.

[0049] In the multi-modal emotion recognition task, the audio, text, and video frame sequences are all highly temporal data. Therefore, the feature extraction method must be able to fully express the temporal information and context semantic information inside these three modalities of information. The three pre-trained models, the Bert model, the Wav2vec2.0 model, and the Fab-Net model, have achieved excellent performance in their respective fields. The features of the corresponding modalities extracted by these three pre-trained models contain sufficient temporal information and context semantic information. Therefore, for text, audio, and video frames, the present invention respectively uses the Bert, Wav2vec2.0, and Fab-Net pre-trained models, and takes the output of the last hidden layer as the feature.

[0050] Step 11: Split the input text into a series of Token word tokens, specifically word token 1, word token 2... word token N; where N is the number of tokens after text splitting. Add the [CLS] token at the beginning of the split sentence to aggregate information for the classification task, and add the [SEP] token at the end of the sentence to indicate the end of the sentence. Each word token is converted into an embedding representation through the embedding layer. The embedding representation consists of word embedding, position embedding, and segment embedding. These embedding representations not only contain the semantic information of the token, but also contain its position in the sentence and the information of the sentence to which it belongs. Subsequently, the embedding representation is input into the Bert pre-trained model for processing to generate the corresponding context encoding representation. The processing flow is as Figure 2 shown, where Token represents a word token, and E1 to EN represents the embedded representation of word tokens, from T1 to T N represents the context-encoded representation of each word token after processing.

[0051] For example: The sentence "The rules you already knew that" is decomposed into word token 1 as "The", word token 2 as "rules", word token 3 as "you", word token 4 as "already", word token 5 as "knew", and word token 6 as "that".

[0052] Step 12: Input the original audio waveform into the Wav2vec2.0 pre-trained model. Process the original audio waveform through a convolutional neural network to extract hidden latent speech representations. Subsequently, input the latent speech representations into a masked Transformer model, which performs context processing on the latent speech representations to generate context representations containing global semantics and temporal dependencies. The processing flow chart is as Figure 3 shown.

[0053] By combining the feature extraction of a convolutional neural network (CNN) and the semantic modeling of a Transformer in Step 12, the entire process can convert the original speech signal into a high-level representation with rich semantic information.

[0054] Step 13: Use the Fab-Net pre-trained model to extract features from facial video frames.

[0055] As Figure 4 shown, the Fab-Net pre-trained model is a self-supervised learning framework that learns the temporal information in the frame sequence by predicting the optical flow field between the starting frame and the ending frame without the need for label information. Input the starting frame and the ending frame together into the encoder to extract features and generate the embedded representation of the starting frame and the embedded representation of the ending frame. These embedded representations not only contain the spatial information of each frame but also reflect the temporal changes between them through optical flow and feature interaction. Next, merge the embedded representations of the starting frame and the ending frame through a concatenation operation to form a joint feature representation containing temporal and spatial relationships. Input the merged feature representation into the decoder, which further processes it to generate a high-dimensional representation containing rich temporal features and spatial information. Finally, use the output of the last hidden layer of the network as the video feature and perform a linear transformation through a fully connected layer to obtain the video feature representation.

[0056] This feature extraction method ensures that the extracted facial expression features can capture the dynamic changes over time, thus improving the performance in different tasks. Among them, the starting frame is the initial input frame in the video sequence, and the network extracts its feature information by encoding the starting frame. The ending frame is the target frame in the video sequence, representing the state of the face at a certain time point. The relationship between the two helps the FAb-Net pre-trained model understand and generate the natural dynamic changes of the face, so as to learn temporal features without label information.

[0057] Step 2: Align the features corresponding to the multi-modalities preprocessed in Step 1 and perform dynamic weight allocation in the multi-modal emotion recognition task through the information aggregation module designed by the present invention. The information aggregation module includes a feature alignment unit and a weight dynamic allocation unit. The structural schematic diagram of the information aggregation module is as Figure 5 shown, Figure 5 where M ∈ (T, A, V), where M represents the modality, T represents text, A represents audio, and V represents video.

[0058] Step 21: Since the present invention extracts the features of the data of three modalities, namely text, audio, and video frame sequences respectively, the three modalities of data obtained are out of sync in terms of time steps. Input the modality data h m containing text data, audio data, and facial expression data into the long short-term memory neural network LSTM, map the time steps of the audio data and facial expression data to the time steps of the text data through the LSTM network, and then apply the Softmax function to the output of the LSTM layer to obtain the probability distribution P of the audio data and facial expression data relative to the text data in the time position. After that, transpose the obtained relative position probability distribution P to obtain the transposed probability distribution P T , and perform batch matrix multiplication (BMM) on the transposed probability distribution P T and the original input audio data and facial expression data to obtain the aligned multi-modal data X m , realizing the alignment of the audio data and facial expression data to the text data in terms of time steps.

[0059] For the aligned multi-modal data X m , m ∈ (T, A, V), perform co-attention operation on it, apply attention weights to the data of each modality to obtain the aggregated modality feature f M , and the specific operation is as follows.

[0060] First, for the data of each modality, calculate the correlation matrix C M between every two modalities and the attention enhancement matrix H M, the mathematical expression is as follows.

[0061] C M = X i ·(W bi ·X j )

[0062]

[0063] where i and j respectively represent any two of the three modalities, X i and X j respectively represent the corresponding modality data, and W is the weight matrix.

[0064] Then, calculate the attention distribution of each modality. Let the attention weight matrix corresponding to each modality be w hi , and input it into the Softmax activation function together with the attention enhancement matrix H M to obtain the attention weight a M .

[0065]

[0066] Finally, perform matrix multiplication on the attention weight a M and the aligned multi-modal data X m to obtain the aggregated modality feature f M .

[0067]

[0068] Step 21 corresponds to the processing flow of the feature alignment and aggregation unit, as Figure 6 shown. In the figure, h M is the multi-modal data, M ∈ (T, A, V), where M represents the modality, T represents the text, A represents the audio, and V represents the video. The function of Step 21 is to align the data of the three modalities in the time step, and allocate attention weights to distinguish the effective information and redundant information within the modality and aggregate them. The feature alignment and aggregation unit pseudo-aligns the data of different modalities and uses the co-attention mechanism to distinguish the information that contributes more to the classification result within the modality.

[0069] Step 22: In order to dynamically estimate the information in the input text and audio features, the present invention trains an encoder network E M : f M → W M , where f M is the modality feature aggregated in Step 21, and W MRefers to the modality weight matrices, M ∈ (T, A), and uses the sigmoid activation function to activate them. To obtain the classification confidence of different modalities, for each modality, a dedicated classifier is constructed Classifier Process each modality through the encoder E M The encoded modality weight matrix W M , and then convert the input modality weight matrix W M into the prediction distribution p M .

[0070] W M = E M (Sifmoid(f M ), M ∈ (T, A, V)

[0071]

[0072] To solve the problem that the maximum class probability p M obtained by classifying through the dedicated classifier for each modality may lead to the problem of false dependence on confidence, the present invention designs a confidence classifier T M that approximates the true class probability for each modality. The confidence classifier T M contains a pre-classifier that pre-classifies the current modality data, and uses its prediction result as the confidence at this time to estimate the contribution degree of the current modality data to the classification result. The mathematical expression is as follows

[0073]

[0074] Finally, input the contribution degree T M of the current modality into the confidence calculator g M , calculate the confidence of the current modality, and multiply the confidence of the current modality by the features of the current modality to achieve the purpose of weight allocation for the current modality, and finally obtain the multi-modal feature F M of the current modality. The mathematical expression is as follows

[0075] F M = W M · g M (T M ) (6)

[0076] Step 22 corresponds to the processing flow of the weight dynamic allocation unit, as Figure 7 shown. The function of step 22 is to achieve dynamic weight allocation for each modality feature through the encoder network and the confidence calculator

[0077] To address the high computational complexity of Transformer and the problem of redundant cross-modal interaction information, the present invention introduces a bottleneck unit. By introducing information from different modalities into the bottleneck unit, the computational complexity is reduced and the noise generated by cross-modal interaction is decreased.

[0078] Figure 8 It is a schematic diagram of the processing flow of the multi-modal bottleneck fusion network of the present invention. Its main function is to perform average fusion on the features from different modalities (video, audio, text), and take the mean of the classification results of the three modalities as the final classification result. The following combines Figure 8 to illustrate the principle of the multi-modal bottleneck network.

[0079] Step 3: The multi-modal bottleneck fusion network consists of a video modality encoder, an audio modality encoder, a text modality encoder, a bottleneck unit, and an average fusion module. Among them, the video modality encoder includes multiple stacked video attention modules, the audio modality encoder includes multiple stacked audio attention modules, and the text modality encoder includes multiple stacked text attention modules.

[0080] Input the video multi-modal feature, audio multi-modal feature, and text multi-modal feature in the multi-modal feature F M obtained in Step 2 into the video modality encoder, audio modality encoder, and text modality encoder respectively for processing to obtain corresponding sub-classification results.

[0081] Among them, the attention module of each modality is an encoder based on the attention mechanism, composed of a multi-head attention block and a multi-layer perceptron. Take the multi-modal feature F M output in Step 2 as the input for processing.

[0082] AttentionBlock = (F M + MLP(Attention(LN(F M )))) (7)

[0083] Connect the data processed by the attention module and the multi-layer perceptron with the multi-modal feature F M through residual connection to obtain the final output. Residual connection is widely used in deep networks, aiming to alleviate the problem of gradient disappearance or gradient explosion caused by too many layers.

[0084] The encoder layer is the key part for multi-modal information interaction. Build a multi-layer encoder stacked by attention modules for each modality to encode the data of each modality.

[0085] First, for the multimodal features of each modality, before the specified cross-modal interaction layer, i.e., the bottleneck unit, the attention module processes the feature information δ of the current modality M .

[0086] In the deep processing, after the feature information δ is processed in the shallow layer M , first it is combined with the bottleneck unit to obtain the fused feature information B M .

[0087] Then, the fused feature information B combined with the bottleneck unit M is fed into the subsequent encoder layer of the current modality. At this time, in the processing of the encoder, the shallow feature information of the current modality will be forced to flow into the bottleneck unit. Restrict the attention flow of all cross-modal information in the model to the bottleneck unit, that is, all cross-modal information interactions are carried out through the bottleneck unit, and keep the number of bottlenecks in the network always less than the data length of each modality. Finally, extract the output η of the last layer of each modality M , respectively pass through the fully connected layer and apply the Softmax activation function to obtain the sub-classification results of the three modalities, and the sub-classification results are input into the average fusion module for averaging to obtain the final classification result.

[0088] This processing method can enable the encoders of each modality to obtain information from other modalities and achieve cross-modal interaction. Compared with the processing method of splicing various modality data and then feeding them into the Transformer network, the method of cross-modal interaction through the bottleneck unit greatly reduces the excessive flow of attention and reduces the computational redundancy.

[0089] In this calculation method, the attention module of each modality can only perform cross-modal information interaction through the bottleneck unit. Therefore, these tightly combined bottleneck units enable the model to sort and compress the information from different modalities and only share the most critical information. During this process, ensure that the number of bottleneck marks in the network is always less than the data length of each modality.

[0090] To verify the beneficial effects of the method of the present invention, experiments are carried out on the public dataset. Select IEMOCAP and CMU-MOSEI as the datasets. The IEMOCAP dataset classifies the emotion categories into four categories: angry, happy, sad, and neutral, and contains 151 videos, recording the conversations between two speakers. The dataset is divided into 70% training data, 10% validation data, and 20% test data. Use accuracy and F1 score as indicators to evaluate the overall effect in four-class emotion recognition.

[0091] The CMU-MOSEI dataset is currently the largest publicly available dataset for multi-modal sentiment recognition, containing 23,259 utterance-video clips. After excluding unaligned and mismatched data samples, there are 22,856 videos remaining in the dataset. Five metrics are used on this dataset: mean absolute error (MAE), Pearson correlation coefficient (Corr), binary classification accuracy, seven-class classification accuracy, and F1 score.

[0092] To evaluate the effectiveness of the multi-modal bottleneck aggregation network proposed in the present invention, a comparative experiment was conducted between the model of the present invention and multiple mainstream models. Nine models that have performed well in this task in recent years, such as TFN (Tensor Fusion Network), LMF (Low-rank Matrix Fusion), MFN (Memory Fusion Network), and MulT (Multimodal Transformer), were selected as benchmark models.

[0093] A comparative experiment was conducted between the present invention and the above nine benchmark models on the CMU-MOSEI dataset, and the results are shown in Table 1. It can be seen that the method proposed in the present invention achieved the best performance in most metrics.

[0094] Table 1 Experimental results of various models on the CMU-MOSEI dataset

[0095]

[0096]

[0097] The comparative experimental results between the method of the present invention and the benchmark models on the IEMOCAP dataset are shown in Table 2. It can be seen that the multi-modal bottleneck aggregation network proposed in the present invention showed the best performance in the accuracy and F1 score of each emotion category, with an average increase of 2.7% compared to the baseline method. Among them, in the recognition of neutral emotions where other models performed poorly, the model of the present invention achieved an obvious improvement of 3.7% - 5.7%. This further verifies the obvious effectiveness of the model of the present invention in bridging the heterogeneity differences of multi-modal data.

[0098] Table 2 Comparison with other baseline methods on the IEOCAP dataset

[0099]

[0100] An ablation experiment was conducted on the Information Aggregation Module IAM and the Multimodal Bottleneck Fusion Network (MBFN) proposed in the present invention and the Transformer model. The experimental results are shown in Table 3:

[0101] Table 3 Ablation experiment results of each module

[0102]

[0103] As can be seen from Table 3, after adding the Information Aggregation Module IAM, the model performance was improved by 10% on average. After replacing the Transformer model with the Multimodal Bottleneck Fusion Network of the present invention, the performance was improved by 1.5% on average.

[0104] It should be noted that the above specific embodiments are exemplary. Those skilled in the art can come up with various solutions inspired by the disclosed content of the present invention, and these solutions also fall within the scope of the disclosure of the present invention and the protection scope of the present invention. Those skilled in the art should understand that the specification and drawings of the present invention are illustrative and do not constitute a limitation on the claims. The protection scope of the present invention is defined by the claims and their equivalents.

Claims

1. A method for detecting emotional states based on multimodal information, characterized in that, Construct a state detection network including an information aggregation module and a multi-modal bottleneck fusion network, and use three pre-trained models of text, speech, and video frames as feature extractors to improve the multi-modal feature expression ability. The information aggregation module realizes feature alignment and dynamic weight allocation in the multi-modal emotion recognition task and suppresses noise. The multi-modal bottleneck fusion network introduces a bottleneck unit, enabling the state detection network to effectively organize and compress information within each modality. Finally, the mean of the sub-classification results of the three modalities is taken as the final classification result. The specific detection method includes: Step 1: Respectively extract the features of the corresponding modalities of the original text, audio waveform, and video frame through three pre-trained models, namely the Bert model, the Wav2vec2.0 model, and the Fab-Net model. The extracted features contain temporal information and context semantic information, specifically including the context encoding representation of the original text, the context representation of the original audio waveform, and the video feature representation of the original video frame. Step 11: Split the input original text into a series of word tokens, add a [CLS] token at the beginning of the split sentence for aggregating information for the classification task, and add a [SEP] token at the end of the sentence to indicate the end of the sentence. Each word token is converted into an embedding representation through the embedding layer, and the embedding representation is input into the Bert pre-trained model for processing to generate the corresponding context encoding representation. Step 12: Input the original audio waveform into the Wav2vec2.0 pre-trained model, process the original audio waveform through a convolutional neural network to extract the hidden potential speech representation, and then input the potential speech representation into the masked Transformer model. The masked Transformer model performs context processing on the potential speech representation to generate a context representation containing global semantics and temporal dependence. Step 13: Use the Fab-Net pre-trained model to extract features from the video frame. Input the start frame and the end frame together into the encoder to extract features, generate the start frame embedding representation and the end frame embedding representation. Merge the embedding representations of the start frame and the end frame through a concatenation operation to form a joint feature representation containing temporal and spatial relationships. Input the joint feature representation into the decoder, and the decoder processes it to obtain a high-dimensional representation containing temporal features and spatial information. Finally, use the output of the last hidden layer of the Fab-Net pre-trained model as the video feature, and perform a linear transformation through a fully connected layer to obtain the video feature representation. Step 2: Implement feature alignment and dynamic weight allocation in the multi-modal emotion recognition task for the multi-modal features output by the pre-trained model through the information aggregation module. The information aggregation module includes a feature alignment unit and a modality-level weight dynamic allocation unit, including: Step 21: Using the text data as a standard, map the time steps of the audio data and facial expression data to the time steps of the text data through the long short-term memory network of the feature alignment unit, and then apply the Softmax function to the output of the long short-term memory network to obtain the relative position probability distribution P of the audio data and facial expression data relative to the text data in terms of time position; Transpose the relative position probability distribution \(P\) to obtain the transposed probability distribution \(P^T\). T , and multiply the transposed probability distribution \(P^T\) T by the original input audio data and facial expression data in batch matrix multiplication to obtain the aligned multimodal data \(X\). m , achieving the alignment of audio data and facial expression data to text data in terms of time steps; Perform co-attention operation on the multi-modal data X m Apply attention weights to the data of each modality to obtain the aggregated modality feature f M ; Step 22: Input the modal feature f M into the modal-level weight dynamic allocation unit for dynamic weight allocation of each modal feature. The weight dynamic allocation unit includes an encoder network E M , a classifier and a confidence classifier T M , and includes: Input the modal feature f M into the encoder network E M for dynamic estimation to obtain their respective corresponding modal weight matrices W M ; Construct a corresponding classifier for each modality Classifier Process the modality weight matrix W M to obtain the prediction distribution p M ; Input the prediction distribution p M into the confidence classifier T M to obtain the confidence of the current modality, and multiply the confidence of the current modality by the features of the current modality to achieve weight assignment for the current modality, and finally obtain the multimodal feature F of the current modality M ; Step 3: The multi-modal bottleneck fusion network consists of a video modality encoder, an audio modality encoder, a text modality encoder, a bottleneck unit, and an average fusion module. Among them, the video modality encoder includes multiple stacked video attention modules, the audio modality encoder includes multiple stacked audio attention modules, and the text modality encoder includes multiple stacked text attention modules. The multi-modal features F M in the video multi-modal features, audio multi-modal features, and text multi-modal features are respectively input into the video modality encoder, audio modality encoder, and text modality encoder for processing; Step 31: For the multi-modal features F of each modality M , the shallow feature information of the current modality is obtained by processing with the attention module before the bottleneck unit; Step 32: Combine the shallow feature information with the bottleneck unit to obtain the fused feature information; Step 33: Feed the fused feature information combined with the bottleneck unit into the encoder layer behind the bottleneck unit of the current modality. Finally, extract the outputs of the last layer of each modality, pass them through the fully connected layer and apply the Softmax activation function respectively to obtain the sub-classification results of the three modalities, and input the sub-classification results into the average fusion module for averaging to obtain the final classification result.