Multimodal Sentiment Analysis Method and System Based on Cross-Modal Attention and Hierarchical Fusion
Through the method of cross-modal attention and hierarchical fusion, text, visual and acoustic features are extracted, cross-attention and gated loop hierarchical fusion are performed, which solves the redundancy problem in the modal fusion stage and improves the accuracy of multimodal sentiment analysis.
Patent Information
- Application Number
- CN202210390047.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-04-14
AI Technical Summary
The existing multimodal sentiment analysis methods ignore the differences between modes in the modal fusion stage, resulting in redundant information retention, affecting model performance and accuracy.
A method based on cross-modal attention and hierarchical fusion is adopted to extract text, visual and acoustic features, cross-attention and gated loop hierarchical fusion network processing is performed, redundant information is eliminated, and synergistic representation information is obtained.
The accuracy and effectiveness of multimodal sentiment analysis are improved, and redundant information is eliminated through the gating mechanism, which enhances the synergy between the information interaction and overall emotional orientation between the modals.
Smart Images

Figure CN115063709B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field, and in particular relates to a multi-modal sentiment analysis method and system based on cross-modal attention and hierarchical fusion. Background Art
[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] Every day, a huge amount of meaningful information is generated around us. Most of this information is generated on the network, and social media is a concentrated area of information on the network, which covers many topics, opinions, emotions and moods closely related to our lives. Multi-modal sentiment analysis (MSA) has always been an active branch field in natural language processing and is widely used in fields such as intelligent healthcare and chatbot recommendations. Compared with traditional sentiment analysis, MSA uses multiple signal sources (extracted original text, acoustics, and vision) to predict the sentiment expressed by a specific object within a specific time period. Two challenges of MSA: 1) How to model the interaction between different modalities, especially supplementary and complementary information; 2) The fusion of data in the case of missing values, misalignments, etc. in the visual and auditory modalities.
[0004] In recent years, researchers have designed complex fusion models; Zadeh et al. designed a tensor fusion network to fuse the feature vectors of three modalities using the Cartesian product; Tasi et al. designed a multi-modal transformer to process all modalities together to obtain the predicted sentiment score; although these methods have achieved good results, there is also a problem that cannot be ignored: the differences between different modalities are ignored, resulting in the loss of key prediction information in the modality representation acquisition stage; Hazarika et al. designed a modality-specific and modality-invariant feature space to combine two types of representations with several losses and evaluate the model effect by means of distance, etc.; Yu et al. used a multi-task form and introduced a modality label automatic generation module in the training stage to assist the main task channel, saving the time of manual annotation and thus improving efficiency; although these studies have achieved exciting results, they lack the inter-modal information interaction in the modality fusion stage, resulting in redundant information being retained until the final prediction stage, affecting the model performance and accuracy. Summary of the Invention
[0005] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a multi-modal sentiment analysis method and system based on cross-modal attention and hierarchical fusion, enabling the modalities to obtain representation information that synergistically affects the overall sentiment orientation during the temporal interaction stage, extracting cross-modal interaction information from the three representations, and eliminating redundant information through a gating mechanism to achieve effective multi-modal representation fusion, thereby improving the fusion result and enhancing the accuracy of sentiment analysis.
[0006] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:
[0007] The first aspect of the present invention provides a multi-modal sentiment analysis method based on cross-modal attention and hierarchical fusion;
[0008] The multi-modal sentiment analysis method based on cross-modal attention and hierarchical fusion includes:
[0009] Extract the text features, visual features, and acoustic features from the video to be analyzed;
[0010] Perform cross-attention on the text features and acoustic features, and text features and visual features to obtain the acoustic modality representation and the visual modality representation;
[0011] The gated recurrent hierarchical fusion network extracts information through pairwise interaction of the acoustic modality representation, visual modality representation, and text features to obtain a one-dimensional vector;
[0012] Use the one-dimensional vector as the sentiment score to predict the sentiment label and obtain the analysis result.
[0013] Furthermore, use the pre-trained 12-layer BERT to extract text features from the video to be analyzed;
[0014] Select the first word vector of the last layer of BERT as the finally extracted text feature.
[0015] Furthermore, for the acoustic features and visual features, use a pre-trained toolkit to process the video to be analyzed to obtain the initial visual features and acoustic features. The specific steps are as follows:
[0016] Obtain the acoustic features and visual features through one-dimensional temporal convolution;
[0017] Embed the temporal information into the features through position encoding.
[0018] Furthermore, the cross-attention is to perform cross-modal cross-fusion of the text features with the acoustic features and visual features respectively to extract the features of interest.
[0019] Furthermore, the specific steps of cross-attention are as follows:
[0020] Parallel attention calculation is performed, and logistic regression is carried out on the Query vector of acoustic and visual features and the Key and Value vectors of text features;
[0021] The head is obtained, and weighted averaging is performed on the output of the parallel attention;
[0022] All the heads are concatenated, and multi-head self-attention connection is performed to obtain the acoustic modality representation and the visual modality representation.
[0023] Furthermore, the acoustic modality representation, the visual modality representation, and the text features are pairwise concatenated and input into a bidirectional gated recurrent network. Different modality information fully interacts, and redundant and irrelevant information in the representation is effectively removed through the gating mechanism to obtain three representations.
[0024] Furthermore, the three concatenated representations are processed using a two-layer RELU activation function to obtain a final one-dimensional vector for sentiment analysis prediction.
[0025] The second aspect of the present invention provides a multi-modal sentiment analysis system based on cross-modal attention and hierarchical fusion.
[0026] The multi-modal sentiment analysis system based on cross-modal attention and hierarchical fusion includes: a feature extraction module, a cross-attention module, and a gated recurrent hierarchical fusion network module;
[0027] The feature extraction module is configured to: extract text features, visual features, and acoustic features in the video to be analyzed;
[0028] The cross-attention module is configured to: perform cross-attention on the text features and the acoustic features, and the text features and the visual features to obtain the acoustic modality representation and the visual modality representation;
[0029] The gated recurrent hierarchical fusion network module is configured to: the gated recurrent hierarchical fusion network extracts information through pairwise interaction of the acoustic modality representation, the visual modality representation, and the text features to obtain a one-dimensional vector.
[0030] The analysis and prediction module uses the one-dimensional vector as the sentiment score to perform sentiment label prediction to obtain the analysis result.
[0031] The third aspect of the present invention provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the steps in the multi-modal sentiment analysis method based on cross-modal attention and hierarchical fusion as described in the first aspect of the present invention are implemented.
[0032] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, the steps in the multi-modal sentiment analysis method based on cross-modal attention and hierarchical fusion as described in the first aspect of the present invention are implemented.
[0033] The above one or more technical solutions have the following beneficial effects:
[0034] Based on the idea of distribution matching, the present invention enables the modality to obtain representation information that has a synergistic effect on the overall sentiment orientation during the time interaction stage, extracts the inter-modal interaction information for 3 pairs of bimodal pairs, and eliminates redundant information through a gating mechanism to achieve effective multi-modal representation fusion.
[0035] The advantages of the additional aspects of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention, and the schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0037] Figure 1 is a flowchart of the method for the first embodiment;
[0038] Figure 2 is a hierarchical gating recurrent network for the first embodiment;
[0039] Figure 3 is a system structure diagram for the second embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0041] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.
[0042] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0043] The general idea proposed by the present invention:
[0044] Based on the sentiment analysis of three modalities of text, vision, and acoustics in the video, after first extracting three features, the text is respectively cross-attentioned with vision and acoustics to obtain the differences between modalities, and then the three features are spliced and fused pairwise to extract the inter-modal interaction information, eliminate redundant information, and the obtained perfect and accurate fusion information is input into the RELU activation function to obtain the sentiment score.
[0045] Embodiment 1
[0046] This embodiment discloses a multi-modal sentiment analysis method based on cross-modal attention and hierarchical fusion;
[0047] As Figure 1 shown, the multi-modal sentiment analysis method based on cross-modal attention and hierarchical fusion includes:
[0048] S1: Extract text features, visual features, and acoustic features from the video to be analyzed;
[0049] S2: Cross-attention the text features with the acoustic features and the text features with the visual features to obtain the acoustic modal representation and the visual modal representation;
[0050] S3: The gated recurrent hierarchical fusion network extracts information through pairwise interactions of the acoustic modal representation, the visual modal representation, and the text features to obtain a one-dimensional vector;
[0051] S4: Use the one-dimensional vector as the sentiment score to perform sentiment label prediction and obtain the analysis result.
[0052] In step S1, first obtain the text sequence, video sequence, and audio sequence from the video to be analyzed, and then input them into the text channel, video channel, and audio channel respectively to extract features:
[0053] For the text channel, use the pre-trained BERT to extract its high-dimensional semantics, and select the first word vector f t of the last layer as the finally extracted feature. The formula is as follows:
[0054]
[0055] Among them, U t represents the initial sequence of the text, θ t bert represents the hyperparameters of the BERT pre-trained model, represents the feature space of the text, d is the space dimension, and t is the text.
[0056] For the acoustic and visual channels, use the pre-trained toolkit to process the original data, learn sufficient perceptual and temporal information, and obtain the initial vector features. The specific steps are as follows:
[0057] 1) One-dimensional temporal convolution: Feed the initial sequence into the one-dimensional temporal convolution. The formula is as follows:
[0058]
[0059] Among them, Conv1D(·) is the one-dimensional temporal convolution function, k m is the size of the convolution kernel used for modality m, U m is the input sequence of modality m, and d is the common dimension, Tm Denotes the utterance length of modality m, where m ∈ {a, v}, a is the acoustic modality, and v is the visual modality.
[0060] 2) Positional embedding: To endow the sequence with temporal information, the positional embedding (PE) is extended to The formula is as follows:
[0061]
[0062] The aim is to calculate the embedding for each position index, where PE(·) represents the positional embedding function, T m Denotes the utterance length of modality m, d is the common dimension, and m ∈ {a, v}.
[0063] In step S2, cross-modal cross-attention is performed on the extracted features to obtain the latent representation information of the acoustic and visual modalities, which has a synergistic effect on the overall sentiment orientation. The specific steps are as follows:
[0064] 1) Parallel attention calculation: Logistic regression is performed on the Query vectors of the acoustic and visual features and the Key and Value vectors of the text features. The formula is as follows:
[0065]
[0066]
[0067] Where Q a , Q v Represent the Query vectors of the acoustic and visual modalities respectively, K t , V t Represent the Key and Value vectors of the text modality respectively, softmax(·) represents the softmax function, and d h Represents the dimension of the modality, and T represents the transpose.
[0068] 2) Obtain the head: Weighted averaging is performed on the output of the parallel attention; the output of each attention is called a head. The calculation formula for the i-th head is:
[0069]
[0070] Here Is the weight matrix of Q m When calculating the head of the i-th m modality; Is the weight matrix of K m When calculating the head of the i-th m modality; Is the weight matrix of V mThe weight matrix; used to linearly project the matrix into a specific space, where m ∈ {a, v}.
[0071] 3) Concatenate all the heads and perform multi-head self-attention connection to obtain the acoustic modality representation and the visual modality representation. The formula is as follows:
[0072]
[0073] is the weight matrix multiplied after concatenating the heads of the m modality. n represents the number of self-attention heads used, n = 10, Concat(·) is the concatenation operation, and m ∈ {a, v}.
[0074] Through the above three steps, the audio modality and the video modality are represented. The general formula is as follows:
[0075]
[0076]
[0077] Among them, represents the main hyperparameters required for the cross-attention module.
[0078] In step S3, that is, a complete and accurate one-dimensional vector is obtained through the gated recurrent hierarchical fusion network; in previous studies, after obtaining effective representations, most directly concatenated the modality representations for final prediction, which would add redundant information and affect the final prediction result. In order to effectively eliminate the redundant information in the representation, as Figure 2 shown, the present invention designs a gated recurrent fusion network to process pairwise combinations of the three representations and feeds them into the gated recurrent hierarchical fusion network to obtain the interaction information between the three feature pairs. The specific steps are as follows:
[0079] 1) Combine the three modality representations obtained pairwise. The formula is as follows:
[0080]
[0081]
[0082]
[0083] Among them, f t are the acoustic representation after cross-modal cross-attention with the text, the visual representation after cross-modal cross-attention with the text, and the text representation respectively.
[0084] 2) It is sent into a bidirectional gated recurrent network to obtain three interaction representations, and the formula is as follows:
[0085]
[0086]
[0087]
[0088] Among them, Bi-GRU(·) represents a bidirectional gated recurrent unit network, and θ gru represents the hyperparameters of the gated recurrent unit network.
[0089] 3) After splicing the three interaction representations, project them into a low-dimensional feature space as follows:
[0090] f s = concat(f t-a , f t-v , f a-v )(16)
[0091]
[0092] Among them, is the parameter matrix, ReLU is the ReLU activation function, represents element-wise multiplication, is the bias term.
[0093] Finally, use the fused representation to predict the multi-modal sentiment:
[0094]
[0095] Among them, is the parameter matrix, ReLU is the ReLU activation function, represents element-wise multiplication, is the bias term.
[0096] In step S4, use the one-dimensional vector y' as the sentiment score to perform sentiment label prediction and obtain the analysis result.
[0097] Preferably, the label score rule is set as follows: when the sentiment score is in the range of (0 - 3], it is a positive sentiment; when the score is in the range of [-3 - 0), it is a negative sentiment; when the score is 0, it is a neutral sentiment.
[0098] Example Two
[0099] This example discloses a multi-modal sentiment analysis system based on cross-modal attention and hierarchical fusion;
[0100] As Figure 3As shown, a multimodal sentiment analysis system based on cross-modal attention and hierarchical fusion includes: a feature extraction module, a cross-attention module, a gated recurrent hierarchical fusion network module, and an analysis and prediction module;
[0101] The feature extraction module is configured to: extract text features, visual features, and acoustic features from the video to be analyzed;
[0102] The cross-attention module is configured to: perform cross-attention on text features and acoustic features, and text features and visual features to obtain acoustic modality representations and visual modality representations;
[0103] The gated recurrent hierarchical fusion network module is configured to: the gated recurrent hierarchical fusion network extracts information through pairwise interaction of acoustic modality representations, visual modality representations, and text features to obtain a one-dimensional vector.
[0104] The analysis and prediction module is configured to: use the one-dimensional vector as the sentiment score to perform sentiment label prediction to obtain the analysis result.
[0105] Embodiment III
[0106] The purpose of this embodiment is to provide a computer-readable storage medium.
[0107] A computer-readable storage medium stores a computer program, and when the program is executed by a processor, it implements the steps in the multimodal sentiment analysis method based on cross-modal attention and hierarchical fusion as described in Embodiment 1 of the present disclosure.
[0108] Embodiment IV
[0109] The purpose of this embodiment is to provide an electronic device.
[0110] An electronic device includes a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the multimodal sentiment analysis method based on cross-modal attention and hierarchical fusion as described in Embodiment 1 of the present disclosure.
[0111] The steps involved in the devices in the above Embodiments II, III, and IV correspond to those in Method Embodiment 1. For specific implementation manners, reference may be made to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0112] Those skilled in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computer device. Optionally, they can be implemented by program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0113] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts on the basis of the technical solutions of the present invention are still within the protection scope of the present invention.
Claims
1. A multimodal sentiment analysis method based on cross-modal attention and hierarchical fusion, characterized in that: Extract text features, visual features, and acoustic features from the video to be analyzed; Perform cross-attention on the text features and acoustic features, and text features and visual features to obtain acoustic modality representations and visual modality representations; the specific steps of the cross-attention are as follows: perform parallel attention calculation, perform logistic regression on the Query vectors of acoustic and visual features and the Key and Value vectors of text features; obtain heads, perform weighted average on the output of parallel attention; splice all heads, perform multi-head self-attention connection to obtain acoustic modality representations and visual modality representations; The gated recurrent hierarchical fusion network extracts information by pairwise interaction of acoustic modality representations, visual modality representations, and text features to obtain a one-dimensional vector. The specific process is as follows: pairwise splice the acoustic modality representations, visual modality representations, and text features, and input them into a bidirectional gated recurrent network. Different modality information fully interacts, and redundant and irrelevant information in the representations is effectively removed through the gating mechanism to obtain three interaction representations; use a two-layer RELU activation function to process the three spliced interaction representations to obtain the final one-dimensional vector; Use the one-dimensional vector as the sentiment score to perform sentiment label prediction to obtain the analysis result.
2. The multimodal sentiment analysis method based on cross-modal attention and hierarchical fusion according to claim 1, wherein Use the pre-trained 12-layer BERT to extract text features from the video to be analyzed; Select the first word vector of the last layer of BERT as the finally extracted text feature.
3. The multimodal sentiment analysis method based on cross-modal attention and hierarchical fusion according to claim 1, wherein For the acoustic features and visual features, use a pre-trained toolkit to process the video to be analyzed to obtain initial visual features and acoustic features. The specific steps are as follows: Obtain acoustic features and visual features through one-dimensional temporal convolution; Embed the position information into the features.
4. The multi-modal sentiment analysis method based on cross-modal attention and hierarchical fusion according to claim 1, wherein The cross-attention is to perform cross-modal cross-fusion of text features with acoustic features and visual features respectively to extract interesting features.
5. A multimodal sentiment analysis system based on cross-modal attention and hierarchical fusion, characterized in that: Including: A feature extraction module, a cross-attention module, and a gated recurrent hierarchical fusion network module; The feature extraction module is configured to: extract text features, visual features, and acoustic features from the video to be analyzed; The cross-attention module is configured to: perform cross-attention on the text features and acoustic features, and text features and visual features to obtain acoustic modality representations and visual modality representations; the specific steps of the cross-attention are as follows: perform parallel attention calculation, perform logistic regression on the Query vectors of acoustic and visual features and the Key and Value vectors of text features; obtain heads, perform weighted average on the output of parallel attention; splice all heads, perform multi-head self-attention connection to obtain acoustic modality representations and visual modality representations; The gated recurrent hierarchical fusion network module is configured to: The gated recurrent hierarchical fusion network extracts information through pairwise interaction of the acoustic modality representation, the visual modality representation, and the text features to obtain a one-dimensional vector. The specific process is as follows: The acoustic modality representation, the visual modality representation, and the text features are pairwise concatenated and input into a bidirectional gated recurrent network. Different modality information interacts sufficiently, and the redundant information and irrelevant information in the representation are effectively removed through the gating mechanism to obtain three interaction representations; The three interaction representations after concatenation are processed by a two-layer RELU activation function to obtain the final one-dimensional vector; The analysis and prediction module uses the one-dimensional vector as the sentiment score to perform sentiment label prediction and obtain the analysis result.
6. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in the multimodal sentiment analysis method based on cross-modal attention and hierarchical fusion as described in any one of claims 1-4.
7. An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the multimodal sentiment analysis method based on cross-modal attention and hierarchical fusion as described in any one of claims 1-4.
Citation Information
Patent Citations
Multi-modal fusion emotion recognition system and method based on multi-task learning and attention mechanism and experimental evaluation method
CN113420807A