Multi-modal sentiment analysis method and system under uncertain content deficiency
The feature fragments are screened through the multi-head attention mechanism and diffusion model feature aggregation method, and combined with the pyramid fusion and convolutional fusion decomposition mechanism, the problem of partial missing features of uncertain modal fragments in multimodal sentiment analysis is solved, and the accuracy and robustness of sentiment analysis are improved.
Patent Information
- Application Number
- CN202510607280.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-13
AI Technical Summary
When the existing multimodal sentiment analysis model deals with the situation where the partial missing features of uncertain modal fragments is found, it is difficult to effectively solve the problem and is easily disturbed by low-quality modal fragments.
The multi-head attention mechanism and diffusion model feature aggregation method are used to screen the multi-modal emotional data in feature fragments, and the three-layer feature fusion of different particle sizes is performed through the pyramid fusion method, and the three-modal decomposition mechanism is used to perform three-modal decomposition and fused with the original mode.
It effectively solves the problem of missing content, improves the quality of key fragments, reduces interference from low-quality modal fragments, and achieves a more comprehensive and detailed fusion of multimodal emotional information.
Smart Images

Figure CN120123993A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal sentiment analysis, and in particular to a multimodal sentiment analysis method and system under the absence of uncertain content. Background Art
[0002] Multimodal sentiment analysis can effectively overcome the limitations of single-modal data in emotional expression and improve the robustness of emotion recognition by fusing different modal data. Therefore, Multimodal Sentiment Analysis (MSA) has emerged and become a new research hotspot.
[0003] Existing MSA models have improved the accuracy of user sentiment analysis and promoted the rapid development of the multimodal sentiment analysis field. However, due to some uncontrollable environmental factors and human factors, multimodal data is lost, resulting in the frequent occurrence of the problem of uncertain modal absence in real applications. For example, due to user privacy issues, the user's image cannot be captured, etc. Therefore, how to solve the MSA problem in an environment of uncertain modal absence has become a research hotspot.
[0004] In multimodal data, each modality consists of segments composed of multiple features. Due to different data processing, the meanings of segments and features may be different, but modalities are all composed of segments composed of multiple features. However, existing research on uncertain modal absence only focuses on the case where the features of modal segments are completely missing, ignoring the situation where the features of modal segments are partially missing in real applications. For example, the upper part of a camera is covered with dust, fingerprints or oil stains, resulting in the inability to capture some areas of the image, thus causing partial loss of the features of the image modal segment. In a noisy environment, background noise may mask some voice signals, resulting in partial loss of audio features. When users input text, there may be grammar errors or typing mistakes, resulting in partial loss of text information, and so on.
[0005] In practical applications, while the MSA model can solve the problem of complete absence of features of uncertain modal segments, it also needs to be able to solve the problem of partial absence of features of uncertain modal segments. Both are problems of uncertain absence of modal content, and we collectively refer to them as the problem of uncertain content absence. Compared with the "uncertain modal absence problem" that only considers the complete absence of features of modal segments, the "uncertain content absence problem" more comprehensively and meticulously reflects the situations that the model may encounter when performing multimodal sentiment analysis in real applications. Therefore, how to solve the MSA problem in an environment of uncertain content absence has become a new challenge.
[0006] In summary, although existing methods can realize multimodal sentiment analysis in the absence of uncertain modalities to a certain extent, there are still the following shortcomings: Existing MSA research on uncertain modal missing only focuses on the case where the features of uncertain modal segments are completely missing, ignoring the fact that in practical applications, the MSA model needs to be able to solve the problem of partial feature missing of uncertain modal segments, that is, the problem of uncertain content missing, while being able to solve the problem of complete feature missing of uncertain modal segments.
[0007] Existing MSA research on uncertain modality loss uses all modal segments for sentiment analysis, which can easily cause the MSA model to be disturbed by low-quality modal segments.
[0008] Existing multimodal emotion fusion methods are mainly limited to the single-level fusion of three single modalities, ignoring the fact that the essence of multimodal fusion is the multi-level fusion of seven modal signals (three single modalities, three two-by-two bimodal fusion modalities, and one trimodal fusion modality). Summary of the invention
[0009] In order to solve the above-mentioned problems, the present invention provides a multimodal sentiment analysis method and system under uncertain content missing.
[0010] In a first aspect, the present invention provides a multimodal sentiment analysis method under uncertain content loss, which adopts the following technical solution: A multimodal sentiment analysis method under uncertain content loss, including: Acquire multimodal sentiment data, including images, audio, and text; The multi-head attention mechanism and diffusion model feature aggregation method are used to screen the feature fragments of the acquired multimodal sentiment data; The selected features are fused into three layers with different granularity through the pyramid fusion method; The fused features are decomposed into three modes through the convolution fusion decomposition mechanism and then fused with the original mode. Get the sentiment prediction results.
[0011] Furthermore, the multi-head attention mechanism and the diffusion model feature aggregation method are used to screen the feature fragments of the acquired multimodal sentiment data, including using the multi-head attention mechanism to process the input sequence, wherein the input sequence is mapped to a new space by linear transformation to generate a query matrix , and the query matrix is processed separately, expressed as Among them, num_heads is the number of attention heads, Will Divided into num_heads sub - matrices, each sub - matrix has a dimension of .
[0012] Furthermore, the method of using the multi - head attention mechanism and the diffusion model feature aggregation method to screen feature segments from the obtained multi - modal sentiment data also includes introducing the feature aggregation and transformation strategy in the diffusion model, using to represent the feature dimension of the diffusion model, and passing through the non - linear activation function to perform feature transformation on the query matrix to generate a new feature representation , and calculating the attention weights of each feature segment through the softmax function to complete segment aggregation, which is expressed as: where and are the weights and biases of the feature transformation.
[0013] Furthermore, the method of using the multi - head attention mechanism and the diffusion model feature aggregation method to screen feature segments from the obtained multi - modal sentiment data also includes splitting the attention weights by head, averaging them in the head dimension, retaining the unique information of each head and comprehensively considering the information of all heads to capture the temporal and spatial features of the data. Among them, the attention weights are divided into num_heads sub - matrices, each sub - matrix has a dimension of (N, T), and then an average operation is performed on the sub - matrices to smooth the attention distribution. Finally, the sequence is sorted according to the attention weights, and the top segments are selected, which is expressed as: where returns the top indices sorted in descending order, are the selected key segments.
[0014] Furthermore, the method of performing three - layer feature fusion with different granularities on the screened features through the pyramid fusion method includes, through concatenation and self - attention mechanism, enabling the three - modal features to be fused to form a high - level feature representation, that is, the three - modal fusion modality, and using the attention diffusion screening module to perform segment screening on it; among them, performing attention diffusion screening ensures that the dimensions of the segments fused with the bimodal are consistent, so as to better guide the fusion of pairwise modalities, which is expressed as: where V, A, T represent the three modalities.
[0015] Further, the three - layer different - granularity feature fusion of the selected features by the pyramid fusion method includes mapping the high - level features to the same dimension as the bimodal fusion through a fully - connected layer to generate a high - level guidance signal. Through matrix multiplication and the Softmax function, the attention weights between the bimodal fusion features and the high - level features are obtained, and the high - level features are weighted according to the transfer score, so that the important high - level feature parts for the bimodal fusion features can play a role. Then, through an addition operation, the transferred high - level feature information is combined with the bimodal fusion features, and self - attention interaction is performed on them to complete the fusion guidance and enhance the bimodal fusion.
[0016] Further, the three - layer different - granularity feature fusion of the selected features by the pyramid fusion method also includes performing attention diffusion screening on all modal signals to ensure avoiding interference from low - quality segments, and then performing the final fusion, which is expressed as: where represents three cases of pairwise modal fusion.
[0017] Further, the trilinear decomposition of the fused features by the convolutional fusion decomposition mechanism and then fusing with the original modalities includes using pointwise convolution to perform a linear transformation on the input feature vector. Among them, given the input feature vector , a new feature vector is obtained through a 1x1 convolution operation, which is expressed as: where is the convolutional kernel weight matrix, is the bias vector, is the separated image modality feature, with the shape of [batch_size, max_visual_len + max_audio_len + max_text_len, att_dim].
[0018] Further, the trilinear decomposition of the fused features by the convolutional fusion decomposition mechanism and then fusing with the original modalities also includes concatenating and fusing the original image modality with the convolution - separated image modality and performing self - attention interaction to complete the missing modality and improve the quality of the image modality; concatenating and fusing the original audio modality with the convolution - separated audio modality and performing self - attention interaction to complete the missing modality and improve the quality of the audio modality; concatenating and fusing the original text modality with the fusion - separated text modality and performing self - attention interaction to complete the missing modality and improve the quality of the text modality.
[0019] In a second aspect, a multi-modal sentiment analysis system under the absence of uncertain content includes: A data acquisition module, configured to acquire multi-modal sentiment data, including three modalities of images, audio, and text; A screening module, configured to perform feature segment screening on the acquired multi-modal sentiment data by using a multi-head attention mechanism and a diffusion model feature aggregation method; A fusion module, configured to perform three-layer feature fusion with different granularities on the screened features by using a pyramid fusion method; A prediction module, configured to perform three-modal decomposition on the fused features by using a convolutional fusion decomposition mechanism and then fuse them with the original modalities; To obtain a sentiment prediction result.
[0020] In a third aspect, the present invention provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are adapted to be loaded and executed by a processor of a terminal device to perform the multi-modal sentiment analysis method under the absence of uncertain content.
[0021] In a fourth aspect, the present invention provides a terminal device, including a processor and a computer-readable storage medium. The processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor to perform the multi-modal sentiment analysis method under the absence of uncertain content.
[0022] In summary, the present invention has the following beneficial technical effects: In order to solve the problem of the absence of uncertain content, the present invention innovatively proposes the MFAMSA model. MFAMSA improves the quality of key segments through convolutional fusion decomposition complementation, multi-level key segment interaction to form new segments, and the attention diffusion screening method for new and old segments, effectively solving the problem of the absence of uncertain content.
[0023] Through the attention diffusion screening method. This method first performs a multi-head attention operation on the input modalities to obtain a weight matrix of the feature of each modality segment. Then, in order to further improve the effect of the attention mechanism and ensure that high-quality modality segments can obtain greater weights, we introduce the feature aggregation and transformation strategy in the diffusion model to perform feature transformation, non-linear activation, and calculation and normalization of importance scores on the weight matrix. Finally, segment screening is performed according to the weight size.
[0024] Through the pyramid fusion method, seven modal signals are fully utilized to achieve multi-level fusion. The pyramid fusion module first preliminarily fuses three single-modal features to form high-level three-modal fusion features. Subsequently, the high-level fusion features are used to guide the fusion between pairwise modalities, thereby enhancing the representation ability of the bimodal fusion features. Finally, all seven modal signals are fused together. By gradually integrating multimodal information through multi-layer fusion with different granularities, the step-by-step enhancement and refinement of information are ensured, enabling the model to capture feature interactions at different levels and comprehensively understand modal information. Brief Description of the Drawings
[0025] Figure 1 is a schematic diagram of a multi-modal sentiment analysis method under uncertain content loss in Embodiment 1 of the present invention; Figure 2 is a schematic diagram of the model structure in Embodiment 1 of the present invention. Detailed Description of the Embodiments
[0026] The present invention will be further described in detail below with reference to the accompanying drawings.
[0027] Embodiment 1 Refer to Figure 1 , a multi-modal sentiment analysis method under uncertain content loss in this embodiment includes: Obtain multi-modal sentiment data, including three modalities: images, audio, and text; Use the multi-head attention mechanism and the diffusion model feature aggregation method to screen feature segments from the obtained multi-modal sentiment data; Perform three-layer feature fusion with different granularities on the screened features through the pyramid fusion method; Perform tri-modal decomposition on the fused features through the convolutional fusion decomposition mechanism and then fuse them with the original modalities; Obtain the sentiment prediction result.
[0028] Specifically: S1. Assume that the multi-modal data for sentiment analysis contains three modalities: , where , and represent the image, audio, and text modalities respectively. Without loss of generality, this article uses to represent the modality with uncertain content loss, where . For example, when the image modality is completely missing, the multi-modal features can be represented as . When the feature parts of the image and text modality segments are missing, the multi-modal data is represented as . The problem studied in this article can be defined as based on multi-modal data User sentiment analysis. For the sake of convenience, in the following sections, this paper uses to represent multimodal data with missing uncertain content.
[0029] S2. To solve the problem of multimodal sentiment analysis under missing uncertain content, this paper proposes an MSA model (MFAMSA) based on multi-level fusion completion and attention diffusion screening, as Figure 2 shown. The innovative inspiration of the MFAMSA model comes from the famous object detection model YOLOv11 (You Only Look Once). Drawing on the model architecture of YOLOv11, we designed the MFAMSA model. MFAMSA improves the quality of modal key segments through convolutional fusion decomposition completion, multi-level key segment interaction to form new segments, and new and old segment attention diffusion screening methods, effectively solving the problem of missing uncertain content.
[0030] Specifically, first, in order to retain key segments and avoid interference from low-quality segments, MFAMSA screens segments of the three modalities with missing uncertain content through the attention diffusion screening module (ADM). Then, in order to complete the initial fusion by filling in the completely missing modality, the three modalities after segment screening are concatenated and self-attention interaction (CS) is performed to form CON1; then MFAMSA re-separates the three modalities from CON1 through the convolutional fusion decomposition module (CSM) and performs CS operations on them with the initial three modalities, and then the new three modalities are pyramidally fused to form CON2. After that, the same decomposition operation is performed on CON2 to form CON3. To solve the problem of the quality decline of key segments caused by partial feature loss, MFAMSA completes the quality improvement of key segments through multi-level key segment interaction to form new segments and new and old segment attention diffusion screening methods. Specifically, MFAMSA performs ADM on CON3 to retain key segments to form CON4 with the same segment dimension as CON2 and performs CS to form CON5; then, ADM is performed on CON5 to form CON6 with the same segment dimension as CON1 and CS is performed to form CON7. Then, CS and ADM are performed on CON5 and CON7 containing all key segments to retain the dimension consistent with the initially fused segments, thereby completing the quality improvement of key segments and forming CONF. Finally, CONF is fed into the softmax function to obtain the final sentiment prediction result. The multi-head attention module, attention diffusion screening module, pyramid fusion module, and convolutional fusion decomposition module mentioned in the MFAMSA model are introduced in detail below.
[0031] S3. Multi-head attention mechanism, In MFAMSA, the multi-head attention mechanism (MHA) serves as a key module, responsible for mining the potential associations between different modalities (text, image, audio) and performing deep feature interactions. Through parallelizing multiple groups of attention calculations, this mechanism can capture global modality interaction features and enhance the model's expressive power. The core principle and calculation formula of the multi-head attention mechanism can be summarized as follows.
[0032] First, to effectively reduce the resource consumption of the multi-head attention mechanism and improve the model's running speed, we observed that in many MSA models, in most cases, the keys (K) and values (V) in the queries (Q), keys (K), and values (V) are the same. Based on this discovery, we proposed an optimization scheme, which is to directly remove the values (V) in the multi-head attention mechanism and use the keys (K) instead of the values (V). That is, assuming the input matrix is , and using two different parameter matrices , to perform linear transformations on the input matrix , and defining Querys as , Keys as , and Values as , we obtain the Q (query), K (key), and V (value) matrices. By doing so, we significantly reduce the memory footprint by reducing the storage of the V matrix and the corresponding weight matrix , and each head only requires two linear transformations instead of three in the traditional MHA, reducing the computational amount in each forward pass and thus accelerating the running speed. Then, the scaled dot product operation is performed for attention calculation, and the specific process is shown in formula (1): (1) where and are weight matrices.
[0033] The multi-head attention mechanism first enhances the model's ability by performing attention calculations on multiple heads in parallel, then concatenates the outputs of these heads, and produces the final output through a linear transformation, as shown in formulas (2) and (3).
[0034] (2) (3) where represents the th head, and represent the weight matrices of the th Query and Key. represents the weight matrix, Indicates the number of attention heads.
[0035] S4. Attention diffusion screening module When processing long sequence data, to improve computational efficiency and model performance, and to retain key modal segments while reducing interference from low-quality segments, we propose a segment screening method based on the multi-head attention mechanism and the diffusion model. This method can effectively shorten the length of the input sequence while screening key information.
[0036] Traditional screening methods usually adopt simple fixed-length screening or random sampling, which may lose important information. In contrast, we use the multi-head attention mechanism and the diffusion model feature aggregation method to dynamically select the most representative sequence segments, which can better retain important segments and features in the sequence and improve the performance of the model. Specifically as follows: First, we obtain the length of the input sequence , and compare it with the target length . If , it indicates that the input sequence is already short enough and does not need to be screened, and the original sequence is directly returned, thus avoiding unnecessary calculations and improving the algorithm efficiency.
[0037] Next, we use the multi-head attention mechanism to process the input sequence. The specific steps are as follows: First, to enable the subsequent attention mechanism to better capture important segments in the sequence, we use a linear transformation to map the input sequence to a new space to generate the query matrix .
[0038] Among them, is the input sequence, is the weight of the query matrix, is the batch size, is the sequence length, is the feature dimension.
[0039] Then, to enhance the expressive power of the model and prepare for subsequent parallel operations and better capture complex dependencies in the input data, the query matrix is processed separately by heads.
[0040] Among them, num_heads is the number of attention heads, The is divided into num_heads sub-matrices, and the dimension of each sub-matrix is .
[0041] Next, to ensure that high-quality modal segments can obtain greater weights, we introduce the feature aggregation and transformation strategy in the diffusion model, where d' represents the feature dimension of the diffusion model. The specific steps include feature transformation, non-linear activation, and the calculation and normalization of importance scores, as follows:
[0042] First, to enhance the expressive power of the segments, we perform feature transformation on the query matrix through a non-linear activation function ( ) to generate a new feature representation , enabling the model to better capture complex patterns.
[0043] Among them, and are the weights and biases of the feature transformation.
[0044] Then, to enable the model to dynamically focus on the most important modal segments in the sequence, we calculate the attention weights for each segment through the softmax function to complete segment aggregation.
[0045] Among them, is the weight vector of the attention weights.
[0046] Next, split S by head and then average over the head dimension, retaining the unique information of each head and comprehensively considering the information of all heads to better capture the temporal and spatial features of the data.
[0047] Specifically, divide S into num_heads submatrices, each with dimension (N, T), and then perform an average operation on these submatrices to smooth the attention distribution and complete the merging.
[0048] Among them, divides into num_heads submatrices, each with dimension , and then performs an average operation on these submatrices to smooth the attention distribution and complete the merging.
[0049] Finally, sort the sequence according to the attention weights and select the top segments.
[0050] Among them, returns the top in descending order.An index, is the filtered key segment.
[0051] S5. The essence of tri-modal fusion is the fusion of seven modal signals (three single-modal signals, three pairwise-combined modal signals, and one tri-modal combined signal). Therefore, in order to make full use of the seven signals for more sufficient fusion, we propose pyramid fusion (as Figure 2 shown). The pyramid fusion module first preliminarily fuses the three single-modal features to form a high-level tri-modal fusion feature. Subsequently, this high-level fusion feature is used to guide the fusion between pairwise modalities, thereby enhancing the representation ability of the bi-modal fusion feature. Finally, all seven modal signals are fused together to ensure that the feature information at different granularities is fully considered, thus achieving a more comprehensive and effective multi-modal information fusion.
[0052] Pyramid fusion gradually integrates multi-modal information through three layers of fusion with different granularities, from single-modal to pairwise-modal and then to the large fusion of all modalities, ensuring the gradual enhancement and refinement of information. This multi-level fusion method can capture feature interactions at different levels, enabling the model to more comprehensively understand modal information. The specific operations are as follows: (1) The first layer of fusion - the fusion of three modalities: Through concatenation and self-attention mechanism, the three modalities are fused to form a high-level feature representation - the tri-modal fusion modality, and the attention diffusion screening module is used to screen its segments, thereby providing a basis for subsequent bi-modal fusion.
[0053] First, perform the concatenation operation: Among them, V, A, and T represent the three modalities.
[0054] Then, perform self-attention and feed-forward network processing: Finally, perform attention diffusion screening to ensure that the dimensions of the segments for bi-modal fusion are consistent, so as to better guide the fusion between pairwise modalities: (2) The second layer of fusion - pairwise-modal fusion guided by high-level features: In this layer, we use high-level features to guide the fusion of pairwise-modal features. By using high-level features to guide the fusion of low-level features, global information can be transmitted to local features, enhancing the representation ability of local features. This guiding mechanism can help the model better understand the complementary information between different modalities, thereby enhancing the representation ability of local features and improving the accuracy of fusion. The specific steps are as follows: First, to guide the bimodal fusion, since the bimodal fusion can increase the global information and form complementary information, we map the high-level features to the same dimension as the bimodal fusion through a fully connected layer to generate a high-level guidance signal.
[0055] Next, through matrix multiplication and the Softmax function, the attention weights between the bimodal fusion features and the high-level features are obtained, so that the more important high-level features can play a greater role in guiding the fusion.
[0056] Among them, are the features after pairwise modality fusion. The pairwise modality fusion is achieved through concatenation operation and self-attention mechanism. Since this operation appears multiple times, it will not be elaborated here.
[0057] Then, according to the transfer score, the high-level features are weighted, so that the part of the high-level features that is more important for the bimodal fusion features can play a greater role.
[0058]
[0059] Finally, through the addition operation, the transferred high-level feature information is combined with the bimodal fusion features, and self-attention interaction is performed on them, so as to complete the fusion guidance and enhance the bimodal fusion.
[0060] .
[0061] (3) The third layer of fusion - the large fusion of all fusion features: After that, the seven-modal signals are screened by attention diffusion to avoid interference from low-quality segments, and then the final fusion is performed.
[0062] Among them represents three cases of pairwise modality fusion.
[0063] Finally, the ALL is screened by the attention diffusion screening module to retain the key segments, remove the redundant segments, and ensure that has the same dimension size as the three-modal fusion modality.
[0064] .
[0065] S6. Convolutional fusion decomposition module, To perform missing modality completion and improve the quality of each modality, we propose a convolutional fusion decomposition module. After the three modalities are concatenated and fused and internal interactions are achieved through the self-attention mechanism, the fused features formed contain the original information of each modality and the information generated by cross-modal interactions. We perform a linear transformation on the fused features through 1×1 convolution, re-decompose them into new three modalities, and perform fusion interactions with the original three modalities one by one, so as to complete the missing modality and improve the quality of each modality. The specific steps are as follows: First, we have obtained the fused multi-modal features through the multi-head self-attention mechanism , where N is the batch size, T is the sequence length, and D is the feature dimension. Then, we separate the information of the image, audio, and text modalities from this fused feature.
[0066] The process of separating the image modality is as follows: In the separation of the image modality, we use 1x1 convolution (also known as point convolution) to perform a linear transformation on the input feature vector. 1x1 convolution is a special convolution operation that does not slide on the time axis and only performs a linear transformation on the feature vector at each time step. This operation is achieved through function. Given the input feature vector , a new feature vector is obtained through the 1x1 convolution operation: where, is the convolution kernel weight matrix, is the bias vector. Here, 1x1 convolution is used, that is, the feature dimension at each time step remains unchanged. is the separated image modality feature, with the shape of [batch_size, max_visual_len + max_audio_len + max_text_len, att_dim]. To ensure that we only retain the part related to the image modality, we perform a slicing operation and only take the first max_visual_len part.
[0067] Then, the original image modality is concatenated and fused with the image modality separated by convolution and self-attention interaction is performed to complete the missing modality and improve the quality of the image modality.
[0068] The process of separating the audio modality is as follows: where, is the convolutional kernel weight matrix, is the bias vector. is the separated audio modality feature, with the shape of [batch_size, max_visual_len + max_audio_len + max_text_len, att_dim]. Next, a slicing operation is performed: Then, the original audio modality is concatenated and fused with the audio modality separated by convolution, as well as self-attention interaction, so as to complement the missing modality and improve the quality of the audio modality.
[0069] The process of text modality fusion and separation is as follows:
[0070] Among them, is the convolutional kernel weight matrix, is the bias vector. is the separated text modality feature, with the shape of [batch_size, max_visual_len + max_audio_len + max_text_len, att_dim]. Next, a slicing operation is performed: Then, the original text modality is concatenated and fused with the text modality separated by fusion, as well as self-attention interaction, so as to complement the missing modality and improve the quality of the text modality.
[0071]
[0072] S7. Training objective, The total loss of the MFAMSA model proposed in this embodiment consists of multiple modules, including the classification loss ( ), and the L2 regularization loss ( ). Next, we will introduce each part in detail.
[0073] (1) Classification loss: The classification loss is defined based on the Softmax cross-entropy loss function (Softmax Cross-Entropy Loss), which is used to measure the difference between the category predicted by the model and the true label. The specific process is as follows: For a given input sample, the model first generates the logits (un-normalized probability prediction values) corresponding to its classification, represented as a vector , where C is the number of classifications, and each element is the predicted value of classification i. The true label of the sample is represented by one-hot encoding, denoted as . One-hot encoding converts the class label of each sample into a binary vector of length C, where only one element is 1, indicating the class to which the sample belongs, and the remaining elements are 0.
[0074] The Softmax cross-entropy loss function is used to calculate the classification loss of each sample. The formula is as follows: To obtain more robust prediction results, we applied the above Softmax cross-entropy loss function to the logits of the model and further summed the losses of all samples to obtain the total classification loss . The calculation formula is as shown below: where N is the number of samples in a batch, and represent the logits and label of the nth sample, respectively.
[0075] (2) L2 regularization loss: To avoid overfitting of the model, we introduced L2 regularization loss during the training process. L2 regularization suppresses the complexity of the model by penalizing the squared values of the model parameters, helps to keep the model parameters smooth, and improves the generalization ability of the model. The L2 regularization loss is defined as:
[0076] where W represents the set of all trainable parameters in the model, and λ is the regularization coefficient, which is used to control the weight of the regularization term in the total loss.
[0077] (3) Total loss function: The total loss function is composed of the classification loss and the L2 regularization loss, ensuring that the model performs excellently in the classification task and also has significant advantages in terms of robustness and generalization ability. The calculation formula is as follows: .
[0078] S8. Experimental verification: To verify the performance of the proposed model MFAMSA in this paper, a large number of experiments were conducted on two public benchmark datasets (CMU-MOSI and IEMOCAP). Below, this paper first introduces the two public datasets and the process of data preprocessing. Then, this paper introduces the experimental settings and the baseline model. Finally, the experimental results are given and analyzed.
[0079] (1)Benchmark datasets In this paper, two well - recognized benchmark datasets for Multimodal Sentiment Analysis (MSA), namely CMU - MOSI and IEMOCAP, are adopted to verify the effectiveness of the proposed model. The following sections will elaborate on these two datasets and their feature extraction processes in detail.
[0080] CMU - MOSI: The dataset CMU - MOSI (Carnegie Mellon University Multimodal Opinion Sentiment and Intensity) is derived from 93 movie clips on YouTube and contains a total of 2,199 monologue images. In this paper, binary classification experiments are conducted on CMU - MOSI, including two labels: negative and positive.
[0081] IEMOCAP: The dataset IEMOCAP (Interactive Emotional Dyadic Motion Capture) is based on emotional conversations and exchanges among participants. It covers 5 sessions, with each session containing approximately 30 image segments, and each segment contains at least 24 utterances. The emotion labels in the IEMOCAP dataset include: neutral, frustrated, angry, happy, sad, excited, surprised, fearful, disappointed, etc. In this paper, binary classification experiments are conducted on IEMOCAP, including two labels: negative and positive.
[0082] (2)In this paper, the experiments are completed on a personal computer equipped with the following hardware: the operating system is Windows 10, equipped with an Intel(R) Core(TM) i9 - 10900K and an Nvidia 3090 graphics processor, and the memory capacity is 96GB. The model framework adopted in the experiment is based on TensorFlow 1.14.0 version, and the programming language is selected as Python 3.7. The key parameter configurations of the model are shown in Table 3 in detail. Among them, the batch size is adjusted to 32, and the hidden layer size is 300. The experimental cycle is set to 15, and the loss value is 0.1.
[0083] When evaluating the model performance, this paper uses accuracy (Acc) and macro F1 - score (Macro F1, M - F1) as evaluation indicators. The calculation formulas of Acc and M - F1 are defined as: In the formula, the number of correctly predicted samples is represented by the symbol The total number of samples is represented by the symbol P represents the positive prediction value, and R represents the recall value.
[0084] (3)To prove the effectiveness of the MFAMSA model proposed in this paper, this paper compares 11 most advanced baseline models, including AE (generalized autoencoder framework); CRA (missing modality reconstruction model based on cascaded residual structure of autoencoders); MCTN, TransM (multi-modal feature fusion models based on end-to-end translation), MMIN (feature reconstruction model for handling missing modalities), ICDN, MRAN, TATE, MTMSA, TATE_J, and SMCMSA.
[0085] For the experiments facing single-modal missing, the modality missing rates are set to 0, 0.1, 0.2, 0.3, 0.4, and 0.5 respectively. Among them, a missing rate of 0 means there is no modality missing; a missing rate of 0.1 means there are 10% uncertain data samples, each sample randomly completely misses one modality (all modality data becomes 0), and there are 10% uncertain data samples, each sample randomly selects one modality, and each segment of this modality randomly misses half of the modality features (randomly selects half of the feature data to become 0); a missing rate of 0.2 means there are 20% uncertain data samples, each sample randomly completely misses one modality, and there are 20% uncertain data samples, each sample randomly selects one modality, and each segment of this modality randomly misses half of the modality features; and so on for the remaining missing rates. To ensure that the experimental design fits the actual application, we first randomly select modalities for complete missing, and then randomly select modalities from all modalities (including the completely missing modalities) for partial missing of segment features.
[0086] The experimental results are shown in Table 1. It can be found from Table 1 that for the dataset CMU-MOSI, at all missing rates (0, 0.1, 0.2, 0.3, 0.4, and 0.5), the MFAMSA model proposed in this paper is superior to the other 11 baseline models in two evaluation metrics (ACC and M-F1). Compared with other baseline models, when the missing rate is 0.5, the MFAMSA model proposed in this paper increases the M-F1 metric value by 7.31% to 17.75% and increases the ACC value by 8.66% to 17.03%. Compared with the second-best performing model - SMCMSA, the MFAMSA model we proposed increases the M-F1 by 7.21% on average and increases the ACC by 7.43% on average.
[0087] In addition, for the IEMOCAP dataset, the MFAMSA model outperforms the other 11 baseline models in both evaluation metrics (ACC and M-F1) at all missing rates (0, 0.1, 0.2, 0.3, 0.4, and 0.5). Compared with other baseline models, the proposed MFAMSA model improves the M-F1 index value by 4.58% to 10.48% and the ACC value by 5.07% to 10.29% at a missing rate of 0.3. Compared with the second best performing model, SMCMSA, the proposed MFAMSA model improves M-F1 by an average of 4.26% and ACC by an average of 4.87%.
[0088] Therefore, according to the experimental results in Table 1, it can be concluded that the overall performance of the MFAMSA model proposed in this paper is better than other baseline models on the two public datasets.
[0089] Table 1 Single-mode missing experimental data For the multi-modal missing experiments, the modal missing rates are set to 0, 0.1, 0.2, 0.3, 0.4 and 0.5 respectively. This paper conducts experiments on multiple modal uncertain content missing by setting different modal missing rates on two public datasets, CMU-MOSI and IEMOCAP. The experimental results are shown in Table 1. A missing rate of 0 means no modal missing; a missing rate of 0.1 means that there are 10% uncertain data samples, each sample randomly completely lacks two modalities, and there are 10% uncertain data samples, each sample randomly selects two modalities, and each fragment of these two modalities randomly lacks half of the modal features; a missing rate of 0.2 means that there are 20% uncertain data samples, each sample randomly completely lacks two modalities, and there are 20% uncertain data samples, each sample randomly selects two modalities, and each fragment of these two modalities randomly lacks half of the modal features; the remaining missing rates are similar. To ensure that the experimental design is suitable for practical applications, we first randomly select modes for complete missingness, and then randomly select modes from all modes (including the completely missing modes) for partial missingness.
[0090] As can be seen from Table 1, for the CMU-MOSI dataset, at all missing rates (0, 0.1, 0.2, 0.3, 0.4, and 0.5), the proposed MFAMSA model in this paper outperforms the other 11 baseline models in terms of two evaluation metrics (ACC and M-F1). In addition, compared with other baseline models, when the missing rate is 0.3, the proposed MFAMSA model in this paper increases the M-F1 metric value by 7.65% to 16.37% and increases the ACC value by 9.52% to 16.93%. Compared with the second-best performing model - SMCMSA, the proposed MFAMSA model in this paper has an average increase of 7.70% in M-F1 and an average increase of 8.41% in ACC.
[0091] For the IEMOCAP dataset, at all missing rates (0, 0.1, 0.2, 0.3, 0.4, and 0.5), the proposed MFAMSA model in this paper outperforms the other 11 baseline models in terms of two evaluation metrics (ACC and M-F1). In addition, compared with other baseline models, when the missing rate is 0.5, the proposed MFAMSA model in this paper increases the M-F1 metric value by 5.40% to 11.48% and increases the ACC value by 7.46% to 13.82%. Compared with the second-best performing model - SMCMSA, the proposed MFAMSA model in this paper has an average increase of 4.64% in M-F1 and an average increase of 6.12% in ACC. Based on the above experimental results, it can be concluded that the proposed MFAMSA model in this paper has better overall performance than other baseline models in solving the problem of multi-modal sentiment analysis under uncertain content missing.
[0092] Table 2 Multi-modal missing experiment data Theoretical analysis. As can be seen from Table 1 and Table 2, the MCTN and TransM models have better performance than the AE and CRA models. This shows that the cyclic translation mechanism adopted in the MCTN and TransM models can extract and integrate information from different modalities more effectively than the auto-encoder mechanism in the AE and CRA models. Compared with the MCTN, MTMSA, and TransM models, the proposed MFAMSA model in this paper shows more excellent results. This is because the MFAMSA model solves the problem of the decline in the quality of modal segments caused by partial missing of modal segment features by using the convolutional fusion decomposition method to complete modal missing filling, forming new segments through multi-level key segment interactions, and improving the segment quality through the new and old segment attention diffusion screening method.
[0093] As can be seen from Table 1 and Table 2, on the CMU-MOSI and IEMOCAP datasets, when the missing rate is 0.5, the ACC and M-F1 of the SMCMSA model show a significant decline. The reason is that when the SMCMSA model lacks a particularly large number of modalities, it is difficult to perform similarity search and completion, and it is impossible to find suitable modalities for completion.
[0094] In addition, as can be seen from Table 1 and Table 2, when the missing rate is set to 0.5, the ACC and M-F1 values of the MTMSA model decrease significantly. This is because the core method of the MTMSA model translates images and audio into the text modality. However, when the missing rate increases, the situation of text modality missing also increases, resulting in a decline in the translation effect and affecting the overall performance of the model.
[0095] To verify the effectiveness of different modules in the MFAMSA model we proposed, a module ablation experiment was conducted in this section. The module ablation experiment was carried out based on the CMU-MOSI dataset. The specific experimental settings and experimental results are as follows.
[0096] Different model variants were generated by removing some key modules from MFAMSA, and the effectiveness of different modules in MFAMSA was verified by testing the performance of the model variants. The generated model variants are as follows: (1) Remove the attention diffusion screening module from MFAMSA and directly use the random sampling method for screening to generate the model variant MFAMSA-AF. (2) Remove the pyramid fusion from MFAMSA, use concatenation fusion instead and perform self-attention interaction after concatenation to generate the model variant MFAMSA-PF. (3) To verify that using convolutional decomposition completion twice in the model architecture is reasonable and effective, we adjusted the model to only retain one convolutional decomposition to generate the model variant MFAMSA-ONE. After convolutional decomposition completion, the three modalities are fused again and concatenated with the initial three-modal fusion result and self-attention interaction are performed and fed to Softmax to obtain the final sentiment analysis result. (4) To verify that the convolutional fusion decomposition module and the MFAMSA model architecture are effective, we directly fed the result after concatenating the three modalities to the Softmax function to obtain the final sentiment analysis result to generate the model variant MFAMSA-ALL.
[0097] The experimental results of the module ablation experiment are shown in Table 3. As can be seen from Table 3, when the missing rate is 0.5, compared with the MFAMSA model, the M-F1 of the MFAMSA-AF model decreases by 4.04% and the ACC decreases by 6.27%. The above experimental results show that the attention diffusion screening module in the MFAMSA model is effective.
[0098] For the MFAMSA-PF model and MFAMSA-ONE, compared with MFAMSA, it can be found that the values of M-F1 and ACC decrease at different missing rates. These results verify that the pyramid fusion module and the two convolutional decomposition completions can improve the performance of MFAMSA.
[0099] Compared with MFAMSA, when the missing rate is 0.3, MFAMSA-ALL decreases by 16.70% in M-F1 and 16.41% in ACC. When the missing rate is 0.4, the M-F1 value of MFAMSA-ALL drops to 17.16%. When the missing rate is set to 0.5, the ACC value of the MFAMSA-ALL model decreases by 19.06%. These results prove that the convolutional fusion decomposition module and the overall MFAMSA model architecture are effective.
[0100] Table 3 Module ablation experiment data
[0101] Example 2 To solve the problem of multi-modal sentiment analysis under uncertain content missing, this paper proposes an MSA model (MFAMSA) based on multi-level fusion completion and attention diffusion screening, as Figure 2 shown. The innovative inspiration of the MFAMSA model comes from the famous object detection model YOLOv11 (You Only Look Once). Drawing on the model architecture of YOLOv11, we designed the MFAMSA model. MFAMSA improves the quality of modal key segments through convolutional fusion decomposition completion, multi-level key segment interaction to form new segments, and the method of attention diffusion screening for new and old segments, effectively solving the problem of uncertain content missing.
[0102] Specifically, This paper proposes the MFAMSA model. First, to retain key segments and avoid interference from low-quality segments, MFAMSA screens the segments of the three modalities with missing uncertain content through the Attention Diffusion Screening Module (ADM). Then, to complete the initial fusion by filling in the completely missing modality, the three modalities after segment screening are concatenated and self-attention interaction (CS) is performed to form CON1. Then MFAMSA re-separates the three modalities from CON1 through the Convolutional Fusion Decomposition Module (CSM), performs CS operation on them with the initial three modalities, and then the new three modalities are pyramidally fused to form CON2. After that, the same decomposition operation is performed on CON2 to form CON3. To solve the problem of the quality decline of key segments caused by partial feature loss, MFAMSA completes the quality improvement of key segments through multi-level key segment interaction to form new segments and the attention diffusion screening method for new and old segments. Specifically, MFAMSA performs ADM on CON3 to retain key segments to form CON4 with the same segment dimension as CON2 and performs CS to form CON5. Then, ADM is performed on CON5 to form CON6 with the same segment dimension as CON1 and CS is performed to form CON7. Then, CS and ADM are performed on CON5 and CON7 containing all key segments to retain the dimension consistent with the initial fusion segments, thus completing the quality improvement of key segments and forming CONF. Finally, CONF is fed into the softmax function to obtain the final sentiment prediction result. The main contributions of this paper are summarized as follows: In practical applications, while the MSA model can solve the problem of complete feature loss of uncertain modality segments, it also needs to be able to solve the problem of partial feature loss of uncertain modality segments, which we call the "problem of missing uncertain content". To our knowledge, this is the first time to clearly propose the problem of "missing uncertain content" in the field of multi-modal sentiment analysis. To solve this problem, we innovatively propose the MFAMSA model. MFAMSA improves the quality of key segments through convolutional fusion decomposition and completion, multi-level key segment interaction to form new segments, and the attention diffusion screening method for new and old segments, effectively solving the problem of missing uncertain content.
[0103] To retain key modality segments and avoid interference from low-quality segments, we propose an attention diffusion screening method. This method first performs multi-head attention operation on the input modality to obtain the weight matrix of the features of each modality segment. Then, to further improve the effect of the attention mechanism and ensure that high-quality modality segments can obtain larger weights, we introduce the feature aggregation and transformation strategy in the diffusion model to perform feature transformation, non-linear activation, and calculation and normalization of importance scores on the weight matrix. Finally, segment screening is performed according to the weight size.
[0104] To improve the multi-modal fusion effect, we propose a pyramid fusion method that fully utilizes seven modal signals to achieve multi-level fusion. The pyramid fusion module first preliminarily fuses three single-modal features to form high-level three-modal fusion features. Subsequently, the high-level fusion features are used to guide the fusion between pairwise modalities, thereby enhancing the representation ability of the bimodal fusion features. Finally, all seven modal signals are fused together. By gradually integrating multi-modal information through multi-level fusion with different granularities, the gradual enhancement and refinement of information are ensured, enabling the model to capture feature interactions at different levels and comprehensively understand modal information.
[0105] To complement missing modalities, we propose a convolutional fusion decomposition method. After the three-modal fusion forms the fusion features, the fusion features at this time contain the original modal information of each modality and the information generated by cross-modal interactions. We perform a linear transformation on the fusion features through point convolution, re-decompose them into new three modalities, and perform fusion interactions with the original three modalities one by one to complete the complement of the missing modalities and improve the quality of each modality.
[0106] A computer-readable storage medium stores multiple instructions, and the instructions are adapted to be loaded and executed by a processor of a terminal device to perform the above method.
[0107] A terminal device includes a processor and a computer-readable storage medium. The processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor to perform the above method.
[0108] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited accordingly. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the protection scope of the present invention.
Claims
1. A multimodal sentiment analysis method under uncertain content loss, characterized in that: include: Acquire multimodal sentiment data, including images, audio, and text; The multi-head attention mechanism and diffusion model feature aggregation method are used to screen the feature fragments of the acquired multimodal sentiment data; The selected features are fused into three layers with different granularity through the pyramid fusion method; The fused features are decomposed into three modes through the convolution fusion decomposition mechanism and then fused with the original mode. Get the sentiment prediction results.
2. According to the multimodal sentiment analysis method under uncertain content loss according to claim 1, it is characterized in that: The method uses a multi-head attention mechanism and a diffusion model feature aggregation method to screen feature fragments of the acquired multimodal sentiment data, including using a multi-head attention mechanism to process the input sequence, wherein the input sequence is mapped to a new space by a linear transformation to generate a query matrix , and process the query matrix separately, expressed as: Among them, num_heads is the number of attention heads, Will Divide into num_heads sub-matrices, each with a dimension of .
3. According to the multimodal sentiment analysis method under uncertain content loss according to claim 2, it is characterized in that: The method uses a multi-head attention mechanism and a diffusion model feature aggregation method to screen feature fragments of the acquired multimodal sentiment data, and also includes introducing a feature aggregation and transformation strategy in the diffusion model. Represents the feature dimension of the diffusion model through a nonlinear activation function The query matrix Perform feature transformation to generate a new feature representation D, and calculate the attention weight of each feature fragment through the softmax function to complete the fragment aggregation, which is expressed as: in, and are the weights and biases of the feature transformation.
4. According to the multimodal sentiment analysis method under uncertain content loss according to claim 3, it is characterized in that: The method uses a multi-head attention mechanism and a diffusion model feature aggregation method to screen feature fragments of the acquired multimodal sentiment data, and also includes dividing the attention weight by head, averaging it on the dimension of the head, retaining the unique information of each head and comprehensively considering the information of all heads to capture the temporal and spatial characteristics of the data, wherein the attention weight is divided into num_heads sub-matrices, each sub-matrix has a dimension of (N, T), and then the sub-matrices are averaged to smooth the attention distribution. Finally, the sequence is sorted according to the attention weight, and the top fragments, represented by: in, Returns the first indexes, It is the key fragment after screening.
5. According to the multimodal sentiment analysis method under uncertain content loss according to claim 4, it is characterized in that: The pyramid fusion method is used to fuse the selected features at three levels with different granularities, including fusing the three-modal features to form a high-level feature representation, i.e., the three-modal fusion mode, through splicing and self-attention mechanism, and using the attention diffusion screening module to perform segment screening on it; wherein, the attention diffusion screening is performed to ensure that the segment dimension is consistent with the bimodal fusion, so as to better guide the fusion of the two modes, which is expressed as: Among them, V, A, and T represent three modes.
6. The multimodal sentiment analysis method under uncertain content loss according to claim 5, characterized in that: The pyramid fusion method is used to fuse the selected features in three layers with different granularity, including mapping the high-level features to the same dimension as the bimodal fusion through a fully connected layer to generate a high-level guidance signal, obtaining the attention weights between the bimodal fusion features and the high-level features through matrix multiplication and the Softmax function, and weighting the high-level features according to the transfer score, so that the high-level feature parts that are important to the bimodal fusion features can play a role, and then, through the addition operation, the transferred high-level feature information is combined with the bimodal fusion features, and self-attention interaction is performed on them, so as to complete the fusion guidance and enhance the bimodal fusion.
7. The multimodal sentiment analysis method under uncertain content loss according to claim 6, characterized in that: The pyramid fusion method is used to fuse the selected features in three layers with different granularity, and also includes attention diffusion screening of all modal signals to avoid interference from low-quality fragments, and then final fusion is performed, which is expressed as: in Represents three cases of pairwise modal fusion.
8. The multimodal sentiment analysis method under uncertain content loss according to claim 7, characterized in that: The convolution fusion decomposition mechanism is used to perform trimodal decomposition of the fused features and then fuse them with the original modalities, including using point convolution to perform linear transformation on the input feature vector, wherein, given the input feature vector , a new feature vector is obtained through a 1x1 convolution operation , expressed as: in, is the convolution kernel weight matrix, is the bias vector, It is the separated image modality feature, with the shape of [batch_size, max_visual_len + max_audio_len + max_text_len, att_dim].
9. The multimodal sentiment analysis method under uncertain content loss according to claim 8, characterized in that: The convolution fusion decomposition mechanism is used to perform trimodal decomposition of the fused features and then fuse them with the original modalities, and further includes splicing and fusing the original image modality with the image modality separated by convolution and performing self-attention interaction to complete the missing modality and improve the quality of the image modality; splicing and fusing the original audio modality with the audio modality separated by convolution and performing self-attention interaction to complete the missing modality and improve the quality of the audio modality; The original text modality and the fused and separated text modality are spliced and fused, and self-attention interacts to complete the missing modality and improve the quality of the text modality.
10. A multimodal sentiment analysis system under uncertain content loss, characterized in that: include: The data acquisition module is configured to acquire multimodal sentiment data, including three modalities: image, audio, and text; The screening module is configured to screen feature fragments of the acquired multimodal sentiment data using a multi-head attention mechanism and a diffusion model feature aggregation method; The fusion module is configured to fuse the selected features at three levels with different granularities through a pyramid fusion method; The prediction module is configured to perform trimodal decomposition of the fused features and then fuse them with the original modalities through a convolutional fusion decomposition mechanism; Get the sentiment prediction results.
Citation Information
Patent Citations
Multi-modal sentiment analysis method and system based on improved diffusion model, medium and product
CN118035940A
Multi-modal sentiment analysis method based on knowledge guidance and modal dynamic attention fusion
CN119443227A
Multi-modal sentiment analysis method and system based on fusion decomposition and trunk gathering
CN119720102A
Method for multimodal emotion classification based on modal space assimilation and contrastive learning
US20240119716A1
Cited By
Variable source voice separation method and system based on adaptive clustering diffusion
CN121122308A
A method and system for variable-source speech separation based on adaptive clustering diffusion
CN121122308B