Multimodal Sentiment Analysis Method and System Based on Fusion Decomposition and Trunk Aggregation
By adopting pyramid multi-head attention mechanism and modal translation technology in multimodal sentiment analysis, the feature extraction problem in the environment of uncertain mode loss is solved, and the accuracy and efficiency of the analysis are improved.
Patent Information
- Application Number
- CN202510228062.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The existing multimodal sentiment analysis methods are difficult to effectively take into account both local and global characteristics in an environment of uncertain mode loss, and the resource consumption of the multi-head attention mechanism is large, which affects the operating speed.
The multimodal emotion analysis method based on fusion decomposition and backbone gathering is adopted to calculate the weight of similar modal features through the pyramid multi-head attention mechanism, perform weighted average fusion, and improve the quality of each modality through modal translation and self-attention interaction.
It improves the accuracy and robustness of sentiment analysis, saves resource consumption of multiple attention, improves operation speed, and effectively completes and strengthens the quality of each modal.
Smart Images

Figure CN119720102B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a multi-modal sentiment analysis method and system based on fusion decomposition and backbone aggregation. Background Art
[0002] Multi-modal sentiment analysis is a research that mines users' opinions and emotional states from multi-modal data such as text, images, and audio. Existing research has achieved good research results in multi-modal emotion recognition, mainly including multi-modal sentiment analysis based on recurrent neural networks, multi-modal sentiment analysis based on Transformer, multi-modal sentiment analysis based on graph convolutional neural networks, multi-modal sentiment analysis based on contrastive learning, multi-modal sentiment analysis based on cross-modal transfer learning, etc. Multi-modal sentiment analysis models usually assume that the three modalities (text, image, and audio) are always available. In fact, due to some uncontrollable environmental factors and human factors, multi-modal data is lost and incompletely collected, resulting in the problem of uncertain modality missing often occurring in real applications.
[0003] In existing multi-modal sentiment analysis methods in an environment with uncertain modality missing, most of the core components are multi-head attention, but when performing multi-head operations, multi-head attention only uses a fixed number of heads and cannot well balance local and global features. And in most cases, the keys (K) and values (V) in multi-head attention are the same, and the repeated calculations and operations waste a lot of time and resources. In addition, existing methods usually only focus on how to complete the missing modality, and often ignore that the most fundamental factor affecting multi-modal fusion effect and final sentiment analysis is the quality of each modality. Summary of the Invention
[0004] In order to solve the above-mentioned problems, the present invention provides a multi-modal sentiment analysis method and system based on fusion decomposition and backbone aggregation.
[0005] In the first aspect, the multi-modal sentiment analysis method based on fusion decomposition and backbone aggregation provided by the present invention adopts the following technical solutions:
[0006] The multi-modal sentiment analysis method based on fusion decomposition and backbone aggregation includes the following steps:
[0007] S1. Obtain multi-modal data with uncertain missing, including three modalities: image, audio, and text;
[0008] S2. Select modal features with high similarity to the current modal features from historical modal features by calculating cosine similarity; calculate the weights of the similar modal features using the pyramid multi-head attention mechanism; finally, perform weighted average fusion on the current modal features and historical modal features according to the weights to form the final modal sample, and replace the current modal sample with it;
[0009] S3. Perform internal interaction on the three modalities through splicing fusion and the pyramid multi-head attention mechanism to form fused features;
[0010] Separate the fused features back into their respective modality features, and perform fusion interaction with the original three modalities one by one through splicing fusion and the pyramid multi-head attention mechanism to form three fused and separated modality features;
[0011] S4. Translate the fused and separated audio modality into image features and text features respectively through the pyramid multi-head attention mechanism and the feed-forward neural network, and perform splicing fusion and self-attention interaction on the translated image features and text features with the original image features and text features respectively to form image-enhanced modality features and text-enhanced modality features;
[0012] Translate the image-enhanced modality and the text-enhanced modality into audio modality respectively through the pyramid multi-head attention mechanism, and perform self-attention interaction on the three audio modalities to form a multi-modal sentiment analysis model;
[0013] S5. Train the multi-modal sentiment analysis model, and use the trained model to output sentiment prediction results.
[0014] Furthermore, calculating the weights of similar modality features using the pyramid multi-head attention mechanism includes, within the sliding window range, screening out the two most similar historical samples with a cosine similarity between the current sample and the historical samples in the range of 0.90 to 0.95 as similar samples, and using the pyramid multi-head attention mechanism to evaluate the importance of each similar sample and the current sample with the similar samples and the current sample, and using this as the weight. The formula is as follows:
[0015] ;
[0016] where is the weight vector, which contains the weights corresponding to each historical sample, , is the weight of each corresponding historical sample, and the two historical samples are used as similar samples .
[0017] Furthermore, the calculation formula of the fused features is as follows:
[0018] ;
[0019] ;
[0020] where represents the splicing fusion feature, They are respectively represented as the text modality, the audio modality, and the image modality. It is represented as the fused feature.
[0021] Furthermore, the above-mentioned splicing fusion and the pyramid multi-head attention mechanism are fused and interacted with the original three modalities one by one, and the calculation formula is as follows:
[0022] ;
[0023] ;
[0024] Among them, It is represented as the re-spliced and fused feature. is the modality feature separated by melting, including images, audio, and text.
[0025] Furthermore, the above-mentioned pyramid multi-head attention mechanism and the feed-forward neural network are used to translate the fused and separated audio modality into the image modality and the text modality respectively, including using the pyramid multi-head attention mechanism for modality translation, taking the fused and separated text features and the fused and separated image features as the queries in the pyramid multi-head attention mechanism respectively, taking the fused and separated audio features as the keys in the pyramid multi-head attention, and then further feature extraction is performed through the feed-forward neural network to obtain the translated text features and the translated image features respectively. The formula is as follows:
[0026] ;
[0027] ;
[0028] Among them, It is represented as the translated modality features, including text and images; It is represented as the modality features after fusion and separation, including text and images; It is represented as the audio features separated by melting, and are weight matrices, and are bias terms, is the activation function.
[0029] Furthermore, the translated image features and text features are respectively spliced, fused, and self-attention interacted with the original image features and text features to form the image-enhanced modality features and the text-enhanced modality features. The formula is as follows:
[0030] ;
[0031] ;
[0032] Among them, To enhance modal features, including images and text, For the translated modal features, including images and text, For the original modal features, including images and text.
[0033] Furthermore, the training objectives of the multi-modal sentiment analysis model include classification loss, L2 regularization loss, and attention regularization loss.
[0034] In a second aspect, the multi-modal sentiment analysis system provided by the present invention based on fusion decomposition and backbone aggregation, which executes the method, includes:
[0035] A data acquisition module, configured to: acquire multi-modal data with uncertain missing values, including three modalities: images, audio, and text;
[0036] A feature module, configured to: perform internal interaction on the three modalities through splicing fusion and pyramid multi-head attention mechanism to form fusion features;
[0037] Separate the fusion features back into their respective modal features, and perform fusion interaction with the original three modalities one by one through splicing fusion and pyramid multi-head attention mechanism to form three fusion-separated modal features;
[0038] A feature enhancement module, configured to: translate the fusion-separated audio modality into image features and text features respectively through the pyramid multi-head attention mechanism and a feed-forward neural network, and perform splicing fusion and self-attention interaction on the translated image features and text features with the original image features and text features respectively to form image-enhanced modal features and text-enhanced modal features;
[0039] Translate the image-enhanced modality and text-enhanced modality into audio modality respectively through the pyramid multi-head attention mechanism, and perform self-attention interaction on the three audio modalities to form the current modal features;
[0040] A model module, configured to: select modal features with high similarity to the current modal features from historical modal features by calculating cosine similarity; calculate the weights of the similar modal features using the pyramid multi-head attention mechanism; and finally perform weighted average fusion on the current modal features and historical modal features according to the weights to form a multi-modal sentiment analysis model;
[0041] A training module, configured to: train the multi-modal sentiment analysis model, and use the trained model to output sentiment prediction results.
[0042] In a third aspect, the present invention provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are suitable for being loaded and executed by a processor of a terminal device to execute the method.
[0043] In a fourth aspect, a terminal device provided by the present invention includes a processor and a computer-readable storage medium. The processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor to perform the above-mentioned method.
[0044] In summary, the present invention has the following beneficial technical effects:
[0045] (1) The multi-modal sentiment analysis method based on fusion decomposition and backbone aggregation of the present invention proposes a pyramid multi-head attention mechanism. The pyramid multi-head attention mechanism gradually increases the number of attention heads in a progressive manner and adaptively combines the results of each layer, so as to better extract multi-level feature information in different modalities and improve the accuracy and robustness of sentiment analysis; in addition, in order to save the resource consumption of multi-head attention and improve the running speed, the value (V) in multi-head attention is removed, and the key (K) is directly used to replace V, so as to remove the operations related to V to improve the performance; the pyramid multi-head attention mechanism solves the problem that the current multi-head attention cannot well balance local and global features.
[0046] (2) The multi-modal sentiment analysis method based on fusion decomposition and backbone aggregation of the present invention proposes a fusion separation method. First, the three modalities are spliced and fused and pyramid self-attention interaction is performed to form a fusion feature. At this time, the fusion feature contains the initial and generated information of each modality, and then reverse transformation is performed. The fusion feature is separated through three fully connected layers, and the three modalities are re-separated, and the fusion and self-attention interaction are performed with the original three modalities one by one to complete the missing modality complementation and improve the quality of each modality.
[0047] (3) The multi-modal sentiment analysis method based on fusion decomposition and backbone aggregation of the present invention proposes a backbone aggregation method. First, the audio modality is translated into the image modality through the modality translation module based on pyramid multi-head attention proposed in this paper, and it is spliced and fused with the original image modality and self-attention interaction (CSI) to complete the missing supplement and quality improvement of the image modality to form an image enhanced modality; then, the audio modality performs the same operation on the text modality to form a text enhanced modality; after that, the image enhanced modality and the text enhanced modality are both translated into the audio modality; in this way, each modality is complemented and enhanced through the modality translation method, and finally the three modalities are all translated into the same modality for fusion to eliminate the differences between modalities; the backbone aggregation method improves the quality of each modality and eliminates the interference of sentiment analysis caused by the differences between modalities.
[0048] (4) The multi-modal sentiment analysis method based on fusion decomposition and backbone aggregation of the present invention proposes a pyramid prior enhancement method; this method aims to improve the quality of the initial three modalities by using similar samples in historical data. Specifically, this method first selects similar samples with high similarity by calculating the cosine similarity between the current sample and historical samples; then calculates the weights of each sample through pyramid multi-head attention; finally, performs weighted average fusion of the current sample and similar samples according to this weight; because the selected similar samples are highly similar to the current sample, this method can introduce more modal information for the current modal sample while retaining the information of the current sample, thereby improving the quality of the initial three modalities; in addition, this method only needs to retain part of the user's historical records and does not need to introduce additional samples, which better meets the requirements of the actual application scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a schematic diagram of the connection of the multi-modal sentiment analysis model of the multi-modal sentiment analysis method based on fusion decomposition and backbone aggregation in Embodiment 1 of the present invention.
[0050] Figure 2 is a schematic diagram of prior enhancement of the multi-modal sentiment analysis method based on fusion decomposition and backbone aggregation in Embodiment 1 of the present invention.
[0051] Figure 3 is a schematic diagram of the experimental results of single-modal loss of the multi-modal sentiment analysis method based on fusion decomposition and backbone aggregation in Embodiment 1 of the present invention; among them, Figure 3 (a) in represents a schematic diagram of the M-F1 values of each model when the loss rate is 0.3, Figure 3 (b) in represents a schematic diagram of the ACC values of each model when the loss rate is 0.3.
[0052] Figure 4 is a schematic diagram of the experimental results of multi-modal loss of the multi-modal sentiment analysis method based on fusion decomposition and backbone aggregation in Embodiment 1 of the present invention, among which, Figure 4 (a) in represents a schematic diagram of the M-F1 values of each model when the loss rate is 0.3, Figure 4 (b) in represents a schematic diagram of the ACC values of each model when the loss rate is 0.3. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] The present invention will be further described in detail below with reference to the accompanying drawings.
[0054] Embodiment 1
[0055] Referring to Figure 1 , the multi-modal sentiment analysis method based on fusion decomposition and backbone aggregation in this embodiment includes the following steps:
[0056] S1. Obtain multi-modal data with uncertain missingness, including three modalities: images, audio, and text.
[0057] S2. Select modality features with high similarity to the current modality features from historical modality features by calculating cosine similarity; calculate the weights of similar modality features using the pyramid multi-head attention mechanism; finally, perform weighted average fusion on the current modality features and historical modality features according to the weights to form the final modality samples, and use them to replace the current modality samples.
[0058] S3. Perform internal interaction on the three modalities through concatenation fusion and the pyramid multi-head attention mechanism to form fusion features;
[0059] Separate the fusion features back into their respective modality features, and perform fusion interaction with the original three modalities one by one through concatenation fusion and the pyramid multi-head attention mechanism to form three fusion-separated modality features.
[0060] S4. Translate the fusion-separated audio modality into image features and text features respectively through the pyramid multi-head attention mechanism and the feed-forward neural network, and perform concatenation fusion and self-attention interaction on the translated image features and text features with the original image features and text features respectively to form image-enhanced modality features and text-enhanced modality features;
[0061] Translate the image-enhanced modality and text-enhanced modality into audio modality respectively through the pyramid multi-head attention mechanism, and perform self-attention interaction on the three audio modalities to form a multi-modal sentiment analysis model.
[0062] S5. Train the multi-modal sentiment analysis model, and use the trained model to output sentiment prediction results.
[0063] The above pyramid multi-head attention mechanism is specifically as follows:
[0064] Directly remove the value (V) in the multi-head attention mechanism, and use the key (K) instead of the value (V), that is, assume the input matrix is , and use two different parameter matrices , Perform linear transformation on the input matrix , and define Querys as , Keys as , Values as , to obtain the Q (query), K (key), and V (value) matrices. Then perform a scaled dot product operation for attention calculation, and its formula is as follows:
[0065] ;
[0066] Among them, and is the weight matrix.
[0067] By performing attention calculations on multiple heads in parallel, then concatenating the outputs of multiple heads, and generating the final output through a linear transformation, the formula is as follows:
[0068] ;
[0069] ;
[0070] where, represents the -th head, and represent the weight matrices of the -th Query and Key, represents the weight matrix, represents the number of attention heads.
[0071] Through a multi-layer multi-head attention mechanism, the number of heads is gradually increased layer by layer. The specific steps are as follows:
[0072] The initial layer has 2 heads. With fewer heads, the feature dimension of each head is larger, and it can capture relatively rough global features that are crucial for the overall understanding of the sentiment analysis task:
[0073] ;
[0074] The middle layer has 4 heads. With a moderate number of heads, the feature dimension of each head is smaller, and it can capture some local features that help refine the sentiment analysis and provide richer context information:
[0075] ;
[0076] The deepest layer has 6 heads. With the largest number of heads, the feature dimension of each head is the smallest, and it can capture the finest-grained features that are crucial for the accuracy of sentiment analysis:
[0077] ;
[0078] Then, adaptively combine the outputs of different layers, and weight the outputs of each layer by learning weights (learned and adjusted through the gradient descent optimization algorithm in backpropagation). The formula is as follows:
[0079] ;
[0080] where, is the weight of the -th layer, which is normalized by softmax so that the sum of the weights is 1.
[0081] Specifically, the multi-modal sentiment analysis method based on fusion decomposition and backbone aggregation includes the following steps:
[0082] S1. Obtain multi-modal data with uncertain missing values, including three modalities: images, audio, and text.
[0083] Assume that the multi-modal data for sentiment analysis contains three modalities: , where , and represent the image, audio, and text modalities respectively. Without loss of generality, is used to represent any missing modality, where . For example, when the image modality does not exist, the multi-modal feature is represented as . When the image and text modalities are missing, the multi-modal data is represented as . The following steps use to represent the multi-modal data with uncertain modality missing.
[0084] S2. Select the modality features with high similarity to the current modality features from the historical modality features by calculating the cosine similarity; calculate the weights of the similar modality features using the pyramid multi-head attention mechanism; finally, perform weighted average fusion of the current modality features and the historical modality features according to the weights to form the final modality sample, and use it to replace the current modality sample.
[0085] As Figure 2 shown, first, considering the limited memory occupancy in the actual application scenario and the difficulty of storing all the historical records of users, a method combining a sliding window and prior enhancement is proposed. The sliding window, that is, select the historical data within a fixed-size window range on the time axis to pick the samples similar to the current data. Specifically, take the current training step as the reference point and set a specific size value as the start position and end position of the historical window.
[0086] Then, within the sliding window range, select the two samples with the highest similarity by calculating the cosine similarity between the current sample and the historical samples. The cosine similarity calculation formula is as follows:
[0087] ;
[0088] where and represent the current sample vector and the historical sample vector, represents the dot product of the vectors, and respectively represent and norms.
[0089] Considering that there may be duplicate samples, first, historical samples with a similarity between 0.90 and 0.95 are screened out, and then the two historical samples with the highest similarity are selected from them. As similar samples .
[0090] After that, the selected similar samples and the current sample are input into the pyramid multi-head attention mechanism together, so as to evaluate the importance of each historical sample and the current sample through attention and use it as a weight.
[0091] ;
[0092] Among them, is the weight vector, which contains the weights corresponding to each historical sample, , is the weight of each corresponding historical sample.
[0093] To ensure that the current sample always dominates in the final fusion process and at the same time can reasonably utilize the information of historical samples, the fixed weight of the current sample is set to 1, and it is combined with the weights of the similar samples .
[0094] ;
[0095] Among them, , and then for each weight value in is normalized.
[0096] ;
[0097] At this time, the final weight vector is obtained.
[0098] Finally, the current sample and the selected similar samples are weighted-averaged and fused to form the final sample , and it is used to replace the current modal sample, thus completing the prior enhancement of the current modal sample:
[0099] ;
[0100] Among them, is the weight of the th sample, is the current sample or the selected similar sample.
[0101] S3. Through splicing fusion and the pyramid multi-head attention mechanism, internal interactions are carried out on the three modalities to form fusion features;
[0102] Separate the fused features back into their respective modal features, and through concatenation fusion and the pyramid multi-head attention mechanism, fuse and interact with the original three modalities one by one to form three fused and separated modal features.
[0103] First, concatenate and fuse the text modality , audio modality and image modality , and perform pyramid multi-head self-attention interaction to form the fused features ,
[0104] ;
[0105] ;
[0106] Then, separate the fused features back into their respective modal features. A fully connected layer is used in each step. The process of fusing and separating the image modality is as follows:
[0107] ;
[0108] Among them, is a learnable weight matrix, is the separated image modality feature, with a shape of [batch_size, max_visual_len + max_audio_len + max_text_len, att_dim]. To ensure that the separated image modality feature is the same as the original and that the extracted feature is only the image modality feature to prevent confusion of information from different modalities, a slicing operation is performed to only take the first max_visual_len part,
[0109] .
[0110] Next, concatenate and fuse the original image modality with the fused and separated image modality and perform self-attention interaction to complement the missing modality and improve the quality of the image modality,
[0111] ;
[0112] .
[0113] The process of fusing and separating the audio modality is as follows:
[0114] ;
[0115] Among them, is a learnable weight matrix, It is the separated audio modality feature, with the shape of [batch_size, max_visual_len + max_audio_len + max_text_len, att_dim]. Next, a slicing operation is performed:
[0116] Then, the original audio modality is concatenated and fused with the separated audio modality, and self-attention interaction is performed to complement the missing modality and improve the quality of the audio modality.
[0117] ;
[0118] .
[0119] The process of text modality fusion and separation is as follows:
[0120] ;
[0121] Among them, is a learnable weight matrix, is the separated text modality feature, with the shape of [batch_size, max_visual_len + max_audio_len + max_text_len, att_dim]. Next, a slicing operation is performed:
[0122] .
[0123] Then, the original text modality is concatenated and fused with the separated text modality, and self-attention interaction is performed to complement the missing modality and improve the quality of the text modality.
[0124] ;
[0125] .
[0126] S4. The separated audio modality is translated into image features and text features respectively through the pyramid multi-head attention mechanism and the feed-forward neural network, and the translated image features and text features are concatenated and fused with the original image features and text features respectively, and self-attention interaction is performed to form image-enhanced modality features and text-enhanced modality features;
[0127] The image-enhanced modality and the text-enhanced modality are translated into audio modalities respectively through the pyramid multi-head attention mechanism, and self-attention interaction is performed on the three audio modalities to form a multi-modal sentiment analysis model.
[0128] The pyramid multi-head attention mechanism is used to achieve modality translation. As Q in the pyramid multi-head attention, is used as K in the pyramid multi-head attention, and then further feature extraction is performed through a feed-forward neural network to achieve the translation from the audio modality to the text modality. The formula is as follows:
[0129] ;
[0130] ;
[0131] where, and are weight matrices, and are bias terms, and ReLU is the activation function.
[0132] After obtaining the translated text modality it is concatenated and fused with the initial and self-attention interaction is performed to obtain the enhanced text modality , thereby enhancing and complementing the text modality through modality translation. The formula is as follows:
[0133] ;
[0134] .
[0135] S5. Train the multi-modal sentiment analysis model and use the trained model to output the sentiment prediction result.
[0136] The total loss of the multi-modal sentiment analysis model consists of multiple modules, including the classification loss ( ), the L2 regularization loss ( ), and the attention regularization loss ( ). Specifically as follows:
[0137] Classification loss: The classification loss is defined based on the Softmax cross-entropy loss function (Softmax Cross-Entropy Loss) and is used to measure the difference between the category predicted by the model and the true label. The specific process is as follows: For a given input sample, the model first generates the logits (un-normalized probability prediction values) corresponding to its classification, denoted as the vector , where C is the number of classifications, and each element is the prediction value of classification i. The true label of the sample is represented by one-hot encoding, denoted as . One-hot encoding converts the category label of each sample into a binary vector of length C, where only one element is 1, indicating the category to which the sample belongs, and the remaining elements are 0.
[0138] The Softmax cross-entropy loss function is used to calculate the classification loss for each sample. The formula is as follows:
[0139] .
[0140] To obtain more robust prediction results, we applied the above Softmax cross-entropy loss function to the logits of the model and further summed the losses of all samples to get the total classification loss . The calculation formula is as follows:
[0141] ;
[0142] where N is the number of samples in a batch, and represent the logits and labels of the nth sample respectively.
[0143] L2 regularization loss: To avoid overfitting of the model, we introduced L2 regularization loss during training. L2 regularization suppresses the complexity of the model by penalizing the squared values of the model parameters, helps to keep the model parameters smooth, and improves the generalization ability of the model. The L2 regularization loss is defined as:
[0144] ;
[0145] where W represents the set of all trainable parameters in the model, and λ is the regularization coefficient, which is used to control the weight of the regularization term in the total loss.
[0146] Attention regularization loss: To prevent the pyramid multi-head attention mechanism from over-focusing on certain specific features during training, we introduced attention regularization loss. The attention regularization loss aims to make the attention weight distribution more uniform, thereby improving the stability and generalization ability of the model. The specific calculation method is as follows:
[0147] ;
[0148] where K is the number of layers of the pyramid, is the number of attention heads in the kth layer, is the attention weight distribution of the ith attention head in the Kth layer, and U is the uniform distribution. is the regularization coefficient, which is used to control the weight of the attention regularization term in the total loss. KL represents the Kullback-Leibler Divergence.
[0149] Total loss function: The total loss function It consists of classification loss, L2 regularization loss, and attention regularization loss, ensuring that the model performs excellently in classification tasks and also has significant advantages in terms of robustness and generalization ability. The calculation formula is as follows:
[0150] .
[0151] Two recognized multi-modal sentiment analysis benchmark datasets, CMU-MOSI and MELD, were adopted to verify the effectiveness of the above multi-modal sentiment analysis model.
[0152] Data preprocessing:
[0153] The feature extraction methods for each modality are as follows:
[0154] In the image modality, the OpenFace2.0 toolbox was used to extract features from the face images in the images, which include timestamp, confidence, successful recognition flag, eye movement, head pose, and facial movement, etc. Each image sample finally formed a high-dimensional feature vector with 709 features.
[0155] For the text modality, feature extraction was performed through a pre-trained BERT model. The text feature dimension output by this model is 768, which can effectively capture the semantic information in the text.
[0156] In terms of audio modality feature extraction, the Librosa library was used for processing. The audio signal was converted to mono and resampled to 16000 Hertz. The processing of audio frames was based on a 512-sample division, and features such as zero-crossing rate, Mel-frequency cepstral coefficients (MFCC), and constant Q transform (CQT) were mainly extracted. The final generated audio feature vector dimension is 33.
[0157] The features of the three modalities were combined in a concatenation manner for subsequent multi-modal sentiment analysis. Through the above methods, the data characteristics under different modalities can be comprehensively analyzed to support the accuracy of sentiment classification.
[0158] Experimental settings:
[0159] The operating system is Windows 10, equipped with an Intel(R) Core(TM) i9-10900K and an Nvidia 3090 graphics processor, and the memory capacity is 96GB. The model framework adopted in the experiment is based on the TensorFlow 1.14.0 version, and the programming language is selected as Python 3.7. The key parameter configurations of the model are shown in Table 1.
[0160] When evaluating the performance of the model, this paper uses accuracy (Acc) and macro F1-score (M-F1) as evaluation metrics. The calculation formulas for Acc and M-F1 are as follows:
[0161] ;
[0162] ;
[0163] where the number of correctly predicted samples is denoted by the symbol and the total number of samples is denoted by the symbol . represents the positive predictive value, represents the recall value.
[0164] Table 1 Key parameter settings in the experiment
[0165] Related parameter category Flag Value Batch size b 32 Experiment period e 15 Hidden layer size d 300 Learning rate <![CDATA[l r > 0.001 Text modality size <![CDATA[l t > 25 Audio modality size <![CDATA[l a > 150 Image modality size <![CDATA[l v > 100 Loss value λ 0.1
[0166] Baseline models:
[0167] To prove the effectiveness of the multi-modal sentiment analysis model of this application, 11 state-of-the-art baseline models were compared. The descriptions of these 11 baseline models are as follows:
[0168] AE: A generalized autoencoder framework designed to optimize the consistency between the neural network output and input, dealing with linear and non-linear autoencoding problems.
[0169] CRA: A missing modality reconstruction model based on the cascaded residual structure of autoencoders, which approximates the input data through the residual connection mechanism.
[0170] MCTN: Promotes the interaction between modalities through modality translation, effectively supporting robust joint relationship learning.
[0171] TransM: A multi-modal feature fusion model based on end-to-end translation, realizing multi-modal interaction through cyclic transformation between modes.
[0172] MMIN: A feature reconstruction model for dealing with missing modalities, which uses cascaded residual autoencoders and forward and backward imagination modules to convert available and missing modalities into each other.
[0173] ICDN: Combines consistency and difference networks, mapping the information of other modalities to the target modality through cross-modal transformers to achieve modality interaction.
[0174] MRAN: This model uses multi-modal and missing index embeddings to guide the reconstruction of missing modalities, while aligning image, audio, and text features to address the challenges of missing modalities.
[0175] TATE_C: adopts label-assisted transformer encoder to handle various cases of uncertain missing modalities, and uses pre-trained models to guide the learning process of joint representation.
[0176] MTMSA: Modal Translation Network, converts image and audio modalities into text modalities, thereby handling the missing modalities and capturing the deep interactions between different modalities. In addition, the model fully utilizes the advantages of text modality.
[0177] TATE_J: Assign different weights to different modalities to make the best use of each modality.
[0178] SMCMSA: For missing modalities, the model selects samples similar to the missing modality through a pre-built complete modality sample database to complete the missing modality.
[0179] In addition, FPMSA: This application.
[0180] For the experiment of single modality loss, the modality loss rate is set to 0.3. The experimental results are shown in Table 2. From Table 2, it can be found that for the dataset CMU-MOSI, the multimodal sentiment analysis model of this application outperforms the other 11 baseline models in both evaluation indicators (ACC and M-F1).
[0181] In addition, for the MELD dataset, the multimodal sentiment analysis model of this application outperforms the other 11 baseline models in both evaluation indicators (ACC and M-F1). Therefore, according to the experimental results in Table 2, it can be concluded that the overall performance of the multimodal sentiment analysis model of this application on the two public datasets is better than that of other baseline models.
[0182] Table 2 Experimental results in the case of single-mode loss
[0183]
[0184] For the multimodal missing experiments, the modal missing rate was set to 0.3. Experiments with multiple modal uncertain missing were carried out on two public datasets, CMU-MOSI and MELD, and the experimental results are shown in Table 3. It can be found from Table 3 that for the dataset CMU-MOSI, the multimodal sentiment analysis model of the present application outperforms the other 11 baseline models in both evaluation indicators (ACC and M-F1). In addition, compared with other baseline models, when the missing rate is 0.3, the multimodal sentiment analysis model of the present application increases the M-F1 index value by 4.85% to 15.56%, and increases the ACC value by 3.38% to 11.29%.
[0185] For the MELD dataset, the multi-modal sentiment analysis model of this application outperforms 11 other baseline models in two evaluation metrics (ACC and M-F1). Based on the above experimental results, it can be concluded that the multi-modal sentiment analysis model of this application has better overall performance than other baseline models when solving multi-modal sentiment analysis with uncertain modality missing.
[0186] Table 3 Experimental results under multi-modal missing
[0187]
[0188] Theoretical analysis: As can be seen from Table 2 and Table 3, the MCTN and TransM models have better performance than the AE and CRA models. It shows that the cyclic translation mechanism adopted in the MCTN and TransM models can extract and integrate information from different modalities more effectively than the auto-encoder mechanism in the AE and CRA models. Compared with the MCTN, MTMSA, and TransM models, the multi-modal sentiment analysis model of this application shows more excellent results. This is because FPMSA enhances the modal information of the three modalities through fusion decomposition and cross-cyclic translation, improving the modal quality and thus enhancing the fusion performance.
[0189] Multi-classification verification:
[0190] To test the sentiment multi-classification performance of the multi-modal sentiment analysis model of this application, this paper selects happy, angry, sad, and neutral sentiment labels to conduct a four-classification experiment on the MELD dataset, as Figure 3 shown, where Figure 3 (a) represents the M-F1 values of each model when the missing rate is 0.3, Figure 3 (b) represents the ACC values of each model when the missing rate is 0.3. Among them, the red bar chart represents this embodiment.
[0191] Select happy, angry, sad, neutral, afraid, disgusted, and surprised sentiment labels to conduct a seven-classification experiment on the MELD dataset, as Figure 4 shown, where Figure 4 (a) represents the M-F1 values of each model when the missing rate is 0.3, Figure 4 (b) represents the ACC values of each model when the missing rate is 0.3. The detailed distribution of the multi-class labels of the MELD dataset is shown in Table 4. Among them, the red bar chart represents this embodiment.
[0192] Table 4 Detailed distribution of the MELD dataset
[0193]
[0194] The results of the four-class sentiment classification experiment are as Figure 3 shown. ByFigure 3 In (a) and Figure 3 In (b), it can be found that the performance of the multi-modal sentiment analysis model of this application is higher than that of other baseline models in all indicators.
[0195] The experimental results of seven types of sentiment classification are as Figure 4 shown. From Figure 4 In (a) and Figure 4 In (b), it can be found that the performance of the multi-modal sentiment analysis model of this application is higher than that of other baseline models in all indicators.
[0196] Based on the above experimental results, it can be concluded that the multi-modal sentiment analysis model of this application has better performance in multi-class sentiment classification.
[0197] Ablation study:
[0198] To verify the effectiveness of different modules in the multi-modal sentiment analysis model (FPMSA) of this application, a module ablation experiment was conducted. The module ablation experiment was carried out based on the MELD dataset, and "T", "A", and "V" were used to represent the text, audio, and image modalities respectively. The specific experimental settings and experimental results are as follows:
[0199] Different model variants were generated by removing some key modules from FPMSA, and the effectiveness of different modules in FPMSA was verified by testing the performance of the model variants. The generated model variants are as follows: (1) Remove the fusion decomposition module from FPMSA to generate the model variant FPMSA-FD. (2) Remove the pyramid prior enhancement from FPMSA to generate the model variant FPMSA-PA. (3) Replace all the pyramid multi-head attention in FPMSA with conventional multi-head attention to generate the model variant FPMSA-MH. (4) Remove the backbone aggregation module in FPMSA, do not perform translation enhancement on each modality, and directly perform three-modal fusion after the fusion decomposition is completed to obtain FPMSA-TR.
[0200] The experimental results of the module ablation experiment are shown in Table 5. It can be seen from Table 5 that both M-F1 and ACC of the FPMSA-FD model decrease compared with the FPMSA model. The above experimental results show that the fusion decomposition module in the FPMSA model is effective.
[0201] For FPMSA-PA, FPMSA-TR models, and FPMSA-MH, compared with FPMSA, it can be found that the values of M-F1 and ACC decrease at a 0.3 missing rate. These results verify that the pyramid prior enhancement module, the backbone aggregation module, and the pyramid multi-head attention can improve the performance of FPMSA.
[0202] Table 5 Experimental results of the module ablation experiment of FPMSA
[0203]
[0204] Example 2
[0205] This example provides a multi-modal sentiment analysis system based on fusion decomposition and backbone aggregation, which executes the above method, including:
[0206] A data acquisition module, configured to: acquire multi-modal data with uncertain missingness, including three modalities: images, audio, and text;
[0207] A feature module, configured to: perform internal interaction on the three modalities through splicing fusion and a pyramid multi-head attention mechanism to form fused features;
[0208] Separate the fused features back into their respective modality features, and perform fusion interaction with the original three modalities one by one through splicing fusion and a pyramid multi-head attention mechanism to form three fused and separated modality features;
[0209] A feature enhancement module, configured to: translate the fused and separated audio modality into image features and text features respectively through a pyramid multi-head attention mechanism and a feed-forward neural network, and perform splicing fusion and self-attention interaction on the translated image features and text features with the original image features and text features respectively to form image-enhanced modality features and text-enhanced modality features;
[0210] Translate the image-enhanced modality and text-enhanced modality into audio modality respectively through a pyramid multi-head attention mechanism, and perform self-attention interaction on the three audio modalities to form current modality features;
[0211] A model module, configured to: select modality features with high similarity to the current modality features from historical modality features by calculating cosine similarity; calculate the weights of the similar modality features using a pyramid multi-head attention mechanism; and finally perform weighted average fusion on the current modality features and historical modality features according to the weights to form a multi-modal sentiment analysis model;
[0212] A training module, configured to: train the multi-modal sentiment analysis model and use the trained model to output sentiment prediction results.
[0213] A computer-readable storage medium, in which multiple instructions are stored, and the instructions are suitable for being loaded and executed by a processor of a terminal device to execute the above method.
[0214] A terminal device, including a processor and a computer-readable storage medium, where the processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are suitable for being loaded and executed by the processor to execute the above method.
[0215] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention shall be covered within the protection scope of the present invention.
Claims
1. A multimodal sentiment analysis method based on fusion decomposition and trunk aggregation, characterized by: The following steps are involved: S1. Acquire multimodal data with uncertain missingness, including three modalities: image, audio, and text; S2, selecting a modal feature with high similarity to the current modal feature from the historical modal features by calculating the cosine similarity; Use the pyramid multi-head attention mechanism to calculate the weights of similar modality features; Finally, the current modal features and historical modal features are weighted averaged and fused according to the weights to form the final modal sample, which is then used to replace the current modal sample. S3, through splicing fusion and pyramid multi-head attention mechanism, the three modalities are internally interacted to form fusion features; Separate the fused features back to their respective modal features, and perform fusion interaction with the original three modalities one by one through splicing fusion and pyramid multi-head attention mechanism to form three fused and separated modal features; S4. The fused and separated audio modalities are translated into image features and text features respectively through the pyramid multi-head attention mechanism and the feedforward neural network, and the translated image features and text features are concatenated and fused with the original image features and text features respectively, and self-attention interaction is performed to form image-enhanced modal features and text-enhanced modal features; The image enhancement modality and text enhancement modality are translated into audio modalities respectively through the pyramid multi-head attention mechanism, and the three audio modalities are interacted with each other through self-attention to form a multimodal sentiment analysis model; S5. Train the multimodal sentiment analysis model and use the trained model to output sentiment prediction results.
2. The multimodal sentiment analysis method based on fusion decomposition and trunk aggregation according to claim 1 is characterized in that: The pyramid multi-head attention mechanism is used to calculate the weights of similar modal features, including within the sliding window, by calculating the cosine similarity of the current sample and the historical sample, selecting the two most similar historical samples with a similarity between 0.90 and 0.95 as similar samples, and using the pyramid multi-head attention mechanism to evaluate the importance of each similar sample and the current sample, and using this as the weight. The formula is as follows: ; in, is a weight vector, which contains the weight corresponding to each historical sample. , is the weight of each corresponding historical sample, two historical samples As a similar sample .
3. The multimodal sentiment analysis method based on fusion decomposition and trunk aggregation according to claim 1 is characterized in that: The calculation formula of the fusion feature is as follows: ; ; in, Represented as concatenated fusion features, Represented as text mode, audio mode and image mode respectively, Represented as fusion features.
4. The multimodal sentiment analysis method based on fusion decomposition and trunk aggregation according to claim 1 is characterized in that: The splicing fusion and pyramid multi-head attention mechanism are used to interact with the original three modes one by one. The calculation is shown as follows: ; ; in, It is represented by the concatenated fusion features again. To melt the separated modal features, including image, audio and text.
5. The multimodal sentiment analysis method based on fusion decomposition and trunk aggregation according to claim 1 is characterized in that: The method uses a pyramid multi-head attention mechanism and a feedforward neural network to translate the fused and separated audio modality into an image modality and a text modality, respectively, including using the pyramid multi-head attention mechanism for modality translation, using the fused and separated text features and the fused and separated image features as queries in the pyramid multi-head attention mechanism, using the fused and separated audio features as keys in the pyramid multi-head attention, and then performing further feature extraction through a feedforward neural network to obtain translated text features and translated image features, respectively, and the formula is as follows: ; ; in, Represented as translated modality features, including text and images; Represented as the fusion of separated modal features, including text and image; Represented as melt-separated audio features, and is the weight matrix, and is the bias term, is the activation function.
6. The multimodal sentiment analysis method based on fusion decomposition and trunk aggregation according to claim 1 is characterized in that: The translated image features and text features are concatenated and fused with the original image features and text features, respectively, and self-attention interaction is performed to form image enhancement modal features and text enhancement modal features. The formula is as follows: ; ; in, To enhance modality features, including images and text, The translated modality features include images and texts. are the original modality features, including images and text.
7. The multimodal sentiment analysis method based on fusion decomposition and trunk aggregation according to claim 1 is characterized in that: The training objectives of the multimodal sentiment analysis model include classification loss, L2 regularization loss, and attention regularization loss.
8. A multimodal sentiment analysis system based on fusion decomposition and trunk aggregation, performing the method according to any one of claims 1 to 7, characterized in that: include: The data acquisition module is configured to: acquire multimodal data with uncertain missing data, including three modalities: image, audio and text; The feature module is configured to: internally interact the three modalities through splicing fusion and pyramid multi-head attention mechanism to form fusion features; Separate the fused features back to their respective modal features, and perform fusion interaction with the original three modalities one by one through splicing fusion and pyramid multi-head attention mechanism to form three fused and separated modal features; The feature enhancement module is configured to: translate the fused and separated audio modalities into image features and text features respectively through a pyramid multi-head attention mechanism and a feedforward neural network, and concatenate and fuse the translated image features and text features with the original image features and text features respectively, and perform self-attention interaction to form image enhancement modal features and text enhancement modal features; The image enhancement modality and text enhancement modality are translated into audio modalities respectively through the pyramid multi-head attention mechanism, and the three audio modalities are interacted with each other through self-attention to form the current modality features; The model module is configured to: select a modal feature having a high similarity with the current modal feature from the historical modal features by calculating the cosine similarity; Use the pyramid multi-head attention mechanism to calculate the weights of similar modality features; Finally, the current modal features and historical modal features are weighted averaged and fused according to the weights to form a multimodal sentiment analysis model; The training module is configured to: train the multimodal sentiment analysis model and output the sentiment prediction result using the trained model.
9. A computer-readable storage medium storing a plurality of instructions, characterized in that: The instructions are suitable for being loaded by a processor of a terminal device and executing the method according to any one of claims 1 to 7.
10. A terminal device, comprising a processor and a computer-readable storage medium, wherein the processor is used to implement each instruction; and the computer-readable storage medium is used to store multiple instructions, characterized in that: The instructions are suitable for being loaded by a processor and executing the method according to any one of claims 1-7.
Citation Information
Patent Citations
Mongolian multi-modal sentiment analysis method based on pre-training model and Transform
CN118503774A
Multi-modal sentiment analysis method combining pre-training model and self-attention block
CN118898046A