A Multimodal Sentiment Analysis Method and System under the Absence of Uncertain Content

Through the multi-head attention mechanism and the MFAMSA model of feature aggregation screening, pyramid fusion and convolutional decomposition completion, the problem of partial missing features of uncertain modal fragments is solved, and the accuracy and robustness of multi-modal sentiment analysis are improved.

CN120123993BActive Publication Date: 2025-07-22YANTAI UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510607280.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-07-22
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

When faced with the partial absence of features of uncertain modal fragments, the existing multimodal sentiment analysis model is susceptible to interference from low-quality modal fragments, and ignores the nature of multi-level fusion, resulting in a decrease in the accuracy of emotion recognition.

Method used

The multi-head attention mechanism and diffusion model feature aggregation method are used to screen feature fragments, multi-level feature fusion is performed through the pyramid fusion method, and modal decomposition and completion are used to form the MFAMSA model.

Benefits of technology

It effectively solves the problem of missing content, improves the accuracy and robustness of sentiment analysis, can understand multimodal information more comprehensively, and enhances the model's expression ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123993B_ABST
    Figure CN120123993B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of multimodal sentiment analysis, and in particular to a multimodal sentiment analysis method and system under the absence of uncertain content. The method includes obtaining multimodal sentiment data, including three modalities of images, audio, and text; using a multi-head attention mechanism and a diffusion model feature aggregation method to screen feature segments of the obtained multimodal sentiment data; performing three-layer different-granularity feature fusion on the screened features through a pyramid fusion method; performing tri-modal decomposition on the fused features through a convolutional fusion decomposition mechanism and then fusing them with the original modalities; and obtaining a sentiment prediction result. In order to solve the problem of the absence of uncertain content, the present invention innovatively proposes the MFAMSA model. MFAMSA improves the quality of key segments through convolutional fusion decomposition complementation, multi-level key segment interaction to form new segments, and new and old segment attention diffusion screening methods, effectively solving the problem of the absence of uncertain content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal sentiment analysis, and in particular, to a multimodal sentiment analysis method and system under the absence of uncertain content. Background Art

[0002] Multimodal sentiment analysis can effectively overcome the limitations of single-modal data in emotional expression and improve the robustness of emotion recognition by fusing different modal data. Therefore, multimodal sentiment analysis (MSA) has emerged and become a new research hotspot.

[0003] Existing MSA models have improved the accuracy of user sentiment analysis and promoted the rapid development of the multimodal sentiment analysis field. However, due to some uncontrollable environmental factors and human factors, multimodal data is lost, resulting in the problem of uncertain modal absence frequently occurring in real applications. For example, due to user privacy issues, the user's image cannot be captured, etc. Therefore, how to solve the MSA problem under the environment of uncertain modal absence has become a research hotspot.

[0004] In multimodal data, each modality consists of segments composed of multiple features. Due to different data processing, the meanings of segments and features may be different, but modalities are all composed of segments composed of multiple features. However, existing research on uncertain modal absence only focuses on the case where the features of modal segments are completely missing, ignoring the case where the features of modal segments are partially missing in real applications. For example, the upper part of the camera is covered with dust, fingerprints or oil stains, resulting in the inability to capture some areas of the image, thus causing partial loss of the features of the image modal segment. In a noisy environment, background noise may mask some voice signals, resulting in partial loss of audio features. When users input text, there may be grammar errors or typing mistakes, resulting in partial loss of text information, and so on.

[0005] In practical applications, while the MSA model can solve the problem of complete absence of features of uncertain modal segments, it also needs to be able to solve the problem of partial absence of features of uncertain modal segments. Both are problems of uncertain absence of modal content, and we collectively refer to them as the problem of uncertain content absence. Compared with the "uncertain modal absence problem" that only considers the complete absence of features of modal segments, the "uncertain content absence problem" more comprehensively and meticulously reflects the situations that the model may encounter when performing multimodal sentiment analysis in real applications. Therefore, how to solve the MSA problem under the environment of uncertain content absence has become a new challenge.

[0006] In summary, although the existing methods can achieve multi-modal sentiment analysis in the case of missing uncertain modalities to a certain extent, there are still the following deficiencies:

[0007] The existing MSA research on missing uncertain modalities only targets the case where the features of uncertain modality segments are completely missing, ignoring that in practical applications, while the MSA model can solve the problem of completely missing features of uncertain modality segments, it also needs to be able to solve the problem of partially missing features of uncertain modality segments, that is, the problem of missing uncertain content.

[0008] The existing MSA research on missing uncertain modalities all uses all modality segments for sentiment analysis, which easily makes the MSA model be interfered by low-quality modality segments.

[0009] The existing multi-modal sentiment fusion methods are mainly limited to the single-level fusion of three single modalities, ignoring that the essence of multi-modal fusion is the multi-level fusion among seven modality signals (three single modalities, three two-modal fusion modalities obtained by pairwise fusion, and one three-modal fusion modality). Summary of the Invention

[0010] To solve the above-mentioned problems, the present invention provides a multi-modal sentiment analysis method and system for missing uncertain content.

[0011] In the first aspect, a multi-modal sentiment analysis method for missing uncertain content provided by the present invention adopts the following technical solutions:

[0012] A multi-modal sentiment analysis method for missing uncertain content includes:

[0013] Obtain multi-modal sentiment data, including three modalities: images, audio, and text;

[0014] Use the multi-head attention mechanism and the diffusion model feature aggregation method to screen feature segments from the obtained multi-modal sentiment data;

[0015] Perform three-layer feature fusion with different granularities on the screened features through the pyramid fusion method;

[0016] Perform three-modal decomposition on the fused features through the convolutional fusion decomposition mechanism and then fuse them with the original modalities;

[0017] Obtain the sentiment prediction result.

[0018] Further, the use of the multi-head attention mechanism and the diffusion model feature aggregation method to screen feature segments from the obtained multi-modal sentiment data includes using the multi-head attention mechanism to process the input sequence, where the input sequence is mapped to a new space through a linear transformation to generate a query matrix , and process the query matrix separately, denoted as

[0019]

[0020] where num_heads is the number of attention heads, Split into num_heads sub-matrices, and the dimension of each sub-matrix is .

[0021] Furthermore, the feature segment screening of the obtained multi-modal sentiment data by using the multi-head attention mechanism and the diffusion model feature aggregation method further includes introducing the feature aggregation and transformation strategy in the diffusion model, using to represent the feature dimension of the diffusion model, and passing through the non-linear activation function to perform feature transformation on the query matrix to generate a new feature representation , and calculating the attention weight of each feature segment through the softmax function to complete segment aggregation, denoted as:

[0022]

[0023] where, and are the weights and biases of the feature transformation.

[0024] Furthermore, the feature segment screening of the obtained multi-modal sentiment data by using the multi-head attention mechanism and the diffusion model feature aggregation method further includes splitting the attention weights by head, averaging them in the dimension of the head, retaining the unique information of each head and comprehensively considering the information of all heads to capture the temporal and spatial features of the data. Among them, the attention weights are divided into num_heads sub-matrices, and the dimension of each sub-matrix is (N, T). Then, an average operation is performed on the sub-matrices to smooth the attention distribution. Finally, the sequence is sorted according to the attention weights, and the top segments are screened, denoted as:

[0025]

[0026] where, returns the top indices sorted in descending order, are the key segments after screening.

[0027] Furthermore, the three - layer different - granularity feature fusion of the screened features by the pyramid fusion method includes fusing the three - modality features through splicing and self - attention mechanism to form a high - level feature representation, that is, the three - modality fusion modality, and using the attention diffusion screening module to perform segment screening on it; among them, performing attention diffusion screening ensures that the dimensions of the segments fused with the bimodal are consistent, so as to better guide the fusion of pairwise modalities, which is expressed as:

[0028]

[0029] Among them, V, A, and T represent three modalities.

[0030] Furthermore, the three - layer different - granularity feature fusion of the screened features by the pyramid fusion method includes mapping the high - level features to the same dimension as the bimodal fusion through a fully - connected layer to generate a high - level guidance signal, obtaining the attention weights between the bimodal fusion features and the high - level features through matrix multiplication and the Softmax function, and weighting the high - level features according to the transfer score, so that the important high - level feature parts for the bimodal fusion features play a role. Then, through an addition operation, the transferred high - level feature information is combined with the bimodal fusion features, and self - attention interaction is performed on them, so as to complete the fusion guidance and enhance the bimodal fusion.

[0031] Furthermore, the three - layer different - granularity feature fusion of the screened features by the pyramid fusion method also includes performing attention diffusion screening on all modality signals to ensure avoiding interference from low - quality segments, and then performing the final fusion, which is expressed as:

[0032]

[0033] Among them represents three cases of pairwise modality fusion.

[0034] Furthermore, the three - modality decomposition of the fused features by the convolution fusion decomposition mechanism and then fusing with the original modalities includes using point convolution to perform a linear transformation on the input feature vector. Among them, given the input feature vector , a new feature vector is obtained through a 1x1 convolution operation, which is expressed as:

[0035]

[0036] Among them, is the convolution kernel weight matrix, is the bias vector, It is the separated image modality feature, with the shape of [batch_size, max_visual_len + max_audio_len + max_text_len, att_dim].

[0037] Further, after the fused features are decomposed into three modalities by the convolutional fusion decomposition mechanism and fused with the original modalities, it also includes splicing and fusing the original image modality with the convolution-separated image modality and self-attention interaction to complete the missing modality and improve the quality of the image modality; splicing and fusing the original audio modality with the convolution-separated audio modality and self-attention interaction to complete the missing modality and improve the quality of the audio modality; splicing and fusing the original text modality with the fusion-separated text modality and self-attention interaction to complete the missing modality and improve the quality of the text modality.

[0038] In a second aspect, a multi-modal sentiment analysis system under uncertain content loss includes:

[0039] A data acquisition module, configured to acquire multi-modal sentiment data, including three modalities of images, audio, and text;

[0040] A screening module, configured to screen feature segments of the acquired multi-modal sentiment data by using the multi-head attention mechanism and the diffusion model feature aggregation method;

[0041] A fusion module, configured to perform three-layer different granularity feature fusion on the screened features by the pyramid fusion method;

[0042] A prediction module, configured to decompose the fused features into three modalities by the convolutional fusion decomposition mechanism and fuse them with the original modalities;

[0043] To obtain the sentiment prediction result.

[0044] In a third aspect, the present invention provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are suitable for being loaded and executed by a processor of a terminal device for the multi-modal sentiment analysis method under uncertain content loss.

[0045] In a fourth aspect, the present invention provides a terminal device, including a processor and a computer-readable storage medium. The processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are suitable for being loaded and executed by the processor for the multi-modal sentiment analysis method under uncertain content loss.

[0046] In summary, the present invention has the following beneficial technical effects:

[0047] To solve the problem of missing uncertain content, the present invention innovatively proposes the MFAMSA model. MFAMSA improves the quality of key segments through convolution fusion decomposition and complementation, multi-level key segment interaction to form new segments, and the attention diffusion screening method for new and old segments, effectively solving the problem of missing uncertain content.

[0048] Through the attention diffusion screening method. This method first performs multi-head attention operations on the input modality to obtain the weight matrix of the feature of each modality segment. Then, in order to further improve the effect of the attention mechanism and ensure that high-quality modality segments can obtain greater weights, we introduce the feature aggregation and transformation strategies in the diffusion model to perform feature transformation, non-linear activation, and calculation and normalization of the importance scores on the weight matrix. Finally, segment screening is performed according to the weight size.

[0049] Through the pyramid fusion method, seven modality signals are fully utilized to achieve multi-level fusion. The pyramid fusion module first performs preliminary fusion on the three single-modal features to form high-level three-modal fusion features. Subsequently, the high-level fusion features are used to guide the fusion between pairwise modalities, thereby enhancing the representation ability of the bimodal fusion features. Finally, all seven modality signals are fused. Through multi-level fusion with different granularities, multi-modal information is gradually integrated, ensuring the gradual enhancement and refinement of information, enabling the model to capture feature interactions at different levels and comprehensively understand modality information. Description of the Drawings

[0050] Figure 1 It is a schematic diagram of a multi-modal sentiment analysis method under the missing of uncertain content in Embodiment 1 of the present invention;

[0051] Figure 2 It is a schematic diagram of the model structure in Embodiment 1 of the present invention. Detailed Embodiments

[0052] The present invention will be further described in detail below with reference to the accompanying drawings.

[0053] Embodiment 1

[0054] Refer to Figure 1 , a multi-modal sentiment analysis method under the missing of uncertain content in this embodiment includes:

[0055] Obtain multi-modal sentiment data, including three modalities: images, audio, and text;

[0056] Use the multi-head attention mechanism and the diffusion model feature aggregation method to screen the feature segments of the obtained multi-modal sentiment data;

[0057] Perform three-layer feature fusion with different granularities on the screened features through the pyramid fusion method;

[0058] After performing trimodal decomposition on the fused features through a convolution fusion decomposition mechanism, they are fused with the original modalities;

[0059] The sentiment prediction result is obtained.

[0060] Specifically:

[0061] S1. Assume that the multimodal data for sentiment analysis contains three modalities: , where , and represent the image, audio, and text modalities respectively. Without loss of generality, in this paper, is used to represent the modality with uncertain content missing, where . For example, when the image modality is completely missing, the multimodal features can be represented as . When the feature parts of the image and text modality segments are missing, the multimodal data is represented as . The problem studied in this paper can be defined as user sentiment analysis based on multimodal data with uncertain content missing. For the sake of convenient representation, in the following parts, this paper uses to represent the multimodal data with uncertain content missing.

[0062] S2. To solve the problem of multimodal sentiment analysis under uncertain content missing, this paper proposes an MSA model (MFAMSA) based on multi-level fusion completion and attention diffusion screening, as shown in Figure 2 . The innovative inspiration of the MFAMSA model comes from the famous object detection model YOLOv11 (You Only Look Once). Drawing on the model architecture of YOLOv11, we designed the MFAMSA model. MFAMSA improves the quality of modal key segments through convolution fusion decomposition completion, multi-level key segment interaction to form new segments, and new and old segment attention diffusion screening methods, effectively solving the problem of uncertain content missing.

[0063] Specifically, first, in order to retain key segments and avoid interference from low-quality segments, MFAMSA screens segments of the trimodal with missing uncertain content through the Attention Diffusion Screening Module (ADM). Next, to complete the missing modality, the trimodal after segment screening is subjected to a splicing operation and self-attention interaction (CS) to complete the initial fusion, forming CON1. Then MFAMSA re-separates the trimodal from CON1 through the Convolutional Fusion Decomposition Module (CSM), performs the CS operation on it and the initial trimodal, and then the new trimodal undergoes pyramid fusion to form CON2. After that, the same decomposition operation is performed on CON2 to form CON3. To address the problem of the quality decline of key segments due to partial feature loss, MFAMSA completes the quality improvement of key segments through multi-level key segment interaction to form new segments and the attention diffusion screening method for new and old segments. Specifically, MFAMSA performs ADM on CON3 to retain key segments to form CON4 with the same segment dimension as CON2 and performs CS to form CON5. Then, ADM is performed on CON5 to form CON6 with the same segment dimension as CON1 and CS is performed to form CON7. Then, CON5 and CON7 containing all key segments are subjected to CS and ADM to retain the dimension consistent with the initially fused segments, thus completing the quality improvement of key segments and forming CONF. Finally, CONF is fed into the softmax function to obtain the final sentiment prediction result. The multi-head attention module, attention diffusion screening module, pyramid fusion module, and convolutional fusion decomposition module mentioned in the MFAMSA model are introduced in detail below.

[0064] S3. Multi-Head Attention Mechanism

[0065] In MFAMSA, the multi-head attention mechanism (Multi-Head Attention, MHA) serves as a key module, responsible for exploring the potential associations between different modalities (text, image, audio) and performing deep feature interactions. Through parallelizing multiple groups of attention calculations, this mechanism can capture global modality interaction features and enhance the model's expressive power. The core principle and calculation formula of the multi-head attention mechanism can be summarized as follows.

[0066] First, to effectively reduce the resource consumption of the multi-head attention mechanism and improve the model's running speed, we observe that in many MSA models, in most cases, the key (K) and value (V) in the query (Q), key (K), and value (V) are the same. Based on this discovery, we propose an optimization scheme, that is, directly removing the value (V) in the multi-head attention mechanism and using the key (K) instead of the value (V). That is, let the input matrix be , and use two different parameter matrices , for the input matrix Perform a linear transformation, and define Querys as , Keys as , and Values as to obtain the Q (query), K (key), and V (value) matrices. By doing so, the memory occupancy is significantly reduced by reducing the storage of the V matrix and the corresponding weight matrix . Moreover, each head only requires two linear transformations instead of three in the traditional MHA, reducing the computational amount in each forward pass and thus accelerating the running speed. Then, perform the scaled dot-product operation for attention calculation, and the specific process is shown in Equation (1):

[0067] (1)

[0068] where and are weight matrices.

[0069] The multi-head attention mechanism first enhances the model's ability by performing attention calculations on multiple heads in parallel, then concatenates the outputs of these heads, and generates the final output through a linear transformation. The process is shown in Equations (2) and (3).

[0070] (2)

[0071] (3)

[0072] where represents the -th head, and represent the weight matrices of the -th Query and Key. represents the weight matrix, represents the number of attention heads.

[0073] S4. Attention Diffusion Screening Module

[0074] When dealing with long sequence data, to improve the computational efficiency and model performance, and retain key modal segments while reducing interference from low-quality segments, we propose a segment screening method based on the multi-head attention mechanism and the diffusion model. This method can effectively shorten the length of the input sequence while screening key information.

[0075] Traditional screening methods usually adopt simple fixed-length screening or random sampling, and these methods may lose important information. In contrast, we use the multi-head attention mechanism and the diffusion model feature aggregation method to dynamically select the most representative sequence segments, which can better retain important segments and features in the sequence and improve the model's performance. The specific process is as follows:

[0076] First, we get the length of the input sequence , and with the target length For comparison. , indicating that the input sequence is short enough and does not need to be screened. The original sequence is returned directly, thus avoiding unnecessary calculations and improving algorithm efficiency.

[0077]

[0078] Next, we use the multi-head attention mechanism to process the input sequence. The specific steps are as follows:

[0079] First, in order to enable the subsequent attention mechanism to better capture the important fragments in the sequence, we use linear transformation to map the input sequence to a new space to generate the query matrix .

[0080]

[0081] in, is the input sequence, is the weight of the query matrix, is the batch size, is the sequence length, is the characteristic dimension.

[0082] Next, in order to enhance the expressiveness of the model and prepare for subsequent parallel operations and to better capture the complex dependencies in the input data, the query matrix is processed separately.

[0083]

[0084] Among them, num_heads is the number of attention heads, Will Divide into num_heads sub-matrices, each sub-matrix has a dimension of .

[0085] Next, in order to ensure that high-quality modal segments can obtain greater weight, we introduced the feature aggregation and transformation strategy in the diffusion model, and used d' to represent the feature dimension of the diffusion model. The specific steps include feature transformation, nonlinear activation, and calculation and normalization of importance scores, as shown below:

[0086] First, to enhance the expressiveness of the fragment, we use a nonlinear activation function ( ) for the query matrix Perform feature transformation to generate new feature representation , so that the model can better capture complex patterns.

[0087]

[0088] Among them, and are the weights and biases of the feature transformation.

[0089] Next, in order to enable the model to dynamically focus on the most important modal segments in the sequence, we calculate the attention weights of each segment through the softmax function to complete segment aggregation.

[0090]

[0091] Among them, is the weight vector of the attention weights.

[0092] Next, split S by head, and then average it over the head dimension, retaining the unique information of each head and comprehensively considering the information of all heads to better capture the temporal and spatial features of the data.

[0093] Specifically, divide S into num_heads submatrices, each submatrix with dimension (N, T), and then perform an average operation on these submatrices to smooth the attention distribution and complete the merging.

[0094]

[0095] Among them, divides into num_heads submatrices, each submatrix with dimension , and then performs an average operation on these submatrices to smooth the attention distribution and complete the merging.

[0096] Finally, sort the sequence according to the attention weights and select the top segments.

[0097]

[0098] Among them, returns the top indices sorted in descending order, and

[0099] S5. The essence of the three-modal fusion is the fusion of seven modal signals (three single-modal signals, three pairwise combined modal signals, and one three-modal combined signal). Therefore, in order to make full use of the seven signals to make the fusion more sufficient, we propose pyramid fusion (as Figure 2As shown). The pyramid fusion module first preliminarily fuses the three unimodal features to form high-level trimodal fusion features. Subsequently, this high-level fusion feature is used to guide the fusion between pairwise modalities, thereby enhancing the representation ability of the bimodal fusion features. Finally, all seven modal signals are fused comprehensively to ensure that the feature information at different granularities is fully considered, thus achieving more comprehensive and effective multimodal information fusion.

[0100] Pyramid fusion gradually integrates multimodal information through three layers of fusion with different granularities, from unimodal to pairwise modality and then to the comprehensive fusion of all modalities, ensuring the gradual enhancement and refinement of information layer by layer. This multi-level fusion method can capture feature interactions at different levels, enabling the model to understand modal information more comprehensively. The specific operations are as follows:

[0101] (1)The first layer of fusion - the fusion of three modalities:

[0102] Through concatenation and self-attention mechanism, the three modalities are fused to form a high-level feature representation - the trimodal fusion modality, and the attention diffusion screening module is used to screen its segments, thus providing a basis for subsequent bimodal fusion.

[0103] First, perform the concatenation operation:

[0104]

[0105] Among them, V, A, and T represent the three modalities.

[0106] Then, perform self-attention and feed-forward network processing:

[0107]

[0108] Finally, perform attention diffusion screening to ensure that the segment dimensions are consistent with bimodal fusion, so as to better guide the fusion of pairwise modalities:

[0109]

[0110] (2)The second layer of fusion - pairwise modality fusion guided by high-level features: In this layer, we use high-level features to guide the fusion of pairwise modality features. By using high-level features to guide the fusion of low-level features, global information can be transmitted to local features, enhancing the representation ability of local features. This guiding mechanism can help the model better understand the complementary information between different modalities, thereby enhancing the representation ability of local features and improving the accuracy of fusion. The specific steps are as follows:

[0111] First, to guide the bimodal fusion, since the bimodal fusion can increase the global information and form complementary information, we map the high-level features to the same dimension as the bimodal fusion through a fully connected layer to generate a high-level guidance signal.

[0112]

[0113] Next, through matrix multiplication and the Softmax function, we obtain the attention weights between the bimodal fusion features and the high-level features, so that the more important high-level features can play a greater role in guiding the fusion.

[0114]

[0115] Among them, are the features after pairwise bimodal fusion. The pairwise bimodal fusion is achieved through concatenation operations and self-attention mechanisms. Since this operation appears multiple times, it will not be elaborated here.

[0116] Then, according to the transfer score, the high-level features are weighted, so that the more important high-level feature part for the bimodal fusion features can play a greater role.

[0117]

[0118] Finally, through an addition operation, the transferred high-level feature information is combined with the bimodal fusion features, and self-attention interaction is performed on them to complete the fusion guidance and enhance the bimodal fusion.

[0119]

[0120] 。

[0121] (3) The third layer of fusion - the large fusion of all fusion features:

[0122] After that, the seven-modal signals are screened by attention diffusion to avoid interference from low-quality segments, and then the final fusion is performed.

[0123]

[0124] Among them represents three cases of pairwise bimodal fusion.

[0125] Finally, the ALL is screened by the attention diffusion screening module to retain the key segments, remove the redundant segments, and ensure that

[0126] is consistent with the dimension size of the trimodal fusion modality.

[0127] 。

[0128] S6. Convolutional Fusion and Decomposition Module

[0129] To perform missing modality completion and improve the quality of each modality, we propose a convolutional fusion and decomposition module. After the three modalities are concatenated, fused, and internally interacted through the self-attention mechanism, the resulting fused features contain the information of each original modality and the information generated by cross-modal interaction. We perform a linear transformation on the fused features through 1x1 convolution, re-decompose them into new three modalities, and perform fusion interaction with the original three modalities one by one to complete the missing modality and improve the quality of each modality. The specific steps are as follows:

[0130] First, we have obtained the fused multi-modal features through the multi-head self-attention mechanism , where N is the batch size, T is the sequence length, and D is the feature dimension. Then, we separate the information of the image, audio, and text modalities from this fused feature.

[0131] The process of separating the image modality is as follows:

[0132] In the separation of the image modality, we use 1x1 convolution (also known as point convolution) to perform a linear transformation on the input feature vector. 1x1 convolution is a special convolution operation that does not slide on the time axis and only performs a linear transformation on the feature vector at each time step. This operation is implemented through function. Given the input feature vector , a new feature vector is obtained through 1x1 convolution operation:

[0133]

[0134] where, is the convolutional kernel weight matrix, is the bias vector. Here, 1x1 convolution is used, that is, the feature dimension at each time step remains unchanged. is the separated image modality feature, with the shape of [batch_size, max_visual_len + max_audio_len + max_text_len, att_dim]. To ensure that we only retain the part related to the image modality, we perform a slicing operation and only take the first max_visual_len part.

[0135]

[0136] Then, the original image modality is concatenated, fused, and self-attention interacted with the image modality separated by convolution to complete the missing modality and improve the quality of the image modality.

[0137]

[0138]

[0139] The audio modality separation process is as follows:

[0140]

[0141] Among them, is the convolutional kernel weight matrix, is the bias vector. is the separated audio modality feature, with the shape of [batch_size, max_visual_len + max_audio_len + max_text_len, att_dim]. Next, perform slicing operations:

[0142]

[0143] Then, splice and fuse the original audio modality with the audio modality separated by convolution and perform self-attention interaction to complete the missing modality and improve the quality of the audio modality.

[0144]

[0145]

[0146] The text modality fusion and separation process is as follows:

[0147]

[0148] Among them, is the convolutional kernel weight matrix, is the bias vector. is the separated text modality feature, with the shape of [batch_size, max_visual_len + max_audio_len + max_text_len, att_dim]. Next, perform slicing operations:

[0149]

[0150] Then, splice and fuse the original text modality with the text modality separated by fusion and perform self-attention interaction to complete the missing modality and improve the quality of the text modality.

[0151]

[0152]

[0153] S7. Training objective

[0154] The total loss of the MFAMSA model proposed in this embodiment consists of multiple modules, including the classification loss ( ), and the L2 regularization loss ( ). Next, we will introduce each part in detail.

[0155] (1) Classification loss: The classification loss is defined based on the Softmax Cross-Entropy Loss function, which is used to measure the difference between the categories predicted by the model and the true labels. The specific process is as follows: For a given input sample, the model first generates the logits (un-normalized probability prediction values) corresponding to its classification, denoted as the vector , where C is the number of classifications, and each element is the prediction value for classification i. The true label of the sample is represented by one-hot encoding, denoted as . One-hot encoding converts the category label of each sample into a binary vector of length C, where only one element is 1, indicating the category to which the sample belongs, and the remaining elements are 0.

[0156] The Softmax cross-entropy loss function is used to calculate the classification loss for each sample. The formula is as follows:

[0157]

[0158] To obtain more robust prediction results, we applied the above Softmax cross-entropy loss function to the logits of the model and further summed the losses of all samples to obtain the total classification loss . The calculation formula is as follows:

[0159]

[0160] where N is the number of samples in a batch, and represent the logits and labels of the nth sample, respectively.

[0161] (2) L2 regularization loss: To avoid overfitting of the model, we introduced the L2 regularization loss during training. L2 regularization suppresses the complexity of the model by penalizing the squared values of the model parameters, helps to maintain the smoothness of the model parameters, and improves the generalization ability of the model. The L2 regularization loss is defined as:

[0162]

[0163] Among them, \(W\) represents the set of all trainable parameters in the model, and \(\lambda\) is the regularization coefficient, which is used to control the weight of the regularization term in the total loss.

[0164] (3) Total loss function: The total loss function consists of the classification loss and the L2 regularization loss, ensuring that the model performs excellently in the classification task and also has significant advantages in terms of robustness and generalization ability. The calculation formula is as follows:

[0165] .

[0166] S8. Experimental verification:

[0167] To verify the performance of the proposed model MFAMSA in this paper, a large number of experiments were carried out on two public benchmark datasets (CMU-MOSI and IEMOCAP). Below, this paper first introduces the two public datasets and the process of data preprocessing. Then, this paper introduces the experimental settings and the baseline model. Finally, the experimental results are given and analyzed.

[0168] (1) Benchmark datasets,

[0169] This paper adopted two recognized multi-modal sentiment analysis (MSA) benchmark datasets - CMU-MOSI and IEMOCAP - to verify the effectiveness of the proposed model. The following sections will elaborate on these two datasets and their feature extraction processes in detail.

[0170] CMU-MOSI: The dataset CMU-MOSI (Carnegie Mellon University Multimodal Opinion Sentiment and Intensity) is derived from 93 movie clips on YouTube and contains a total of 2,199 monologue images. This paper conducted binary classification experiments on CMU-MOSI, including two labels: negative and positive.

[0171] IEMOCAP: The dataset IEMOCAP (Interactive Emotional Dyadic Motion Capture) is based on emotional conversations and exchanges among participants. It covers 5 sessions, with each session containing approximately 30 image segments, and each segment contains at least 24 utterances. The emotion labels in the IEMOCAP dataset include: neutral, frustrated, angry, happy, sad, excited, surprised, frightened, disappointed, etc. This paper conducted binary classification experiments on IEMOCAP, including two labels: negative and positive.

[0172] (2)In this paper, the experiments were conducted on a personal computer equipped with the following hardware: Windows 10 operating system, Intel(R) Core(TM) i9-10900K and Nvidia 3090 graphics processor, and a memory capacity of 96GB. The model framework used in the experiment was based on TensorFlow version 1.14.0, and the programming language was chosen as Python 3.7. The key parameter configurations of the model are shown in Table 3. Among them, the batch size was adjusted to 32, and the hidden layer size was 300. The experimental period was set to 15, and the loss value was 0.1.

[0173] When evaluating the model performance, this paper used accuracy (Acc) and macro F1 score (Macro F1, M-F1) as evaluation indicators. The calculation formulas of Acc and M-F1 are defined as:

[0174]

[0175]

[0176] In the formula, the number of correctly predicted samples is represented by the symbol and the total number of samples is represented by the symbol P represents the positive prediction value, and R represents the recall value.

[0177] (3)To prove the effectiveness of the MFAMSA model proposed in this paper, this paper compared 11 state-of-the-art baseline models, including AE (generalized autoencoder framework); CRA (missing modality reconstruction model based on autoencoder cascaded residual structure); MCTN, TransM (multi-modal feature fusion models based on end-to-end translation), MMIN (feature reconstruction model for handling missing modalities), ICDN, MRAN, TATE, MTMSA, TATE_J, and SMCMSA.

[0178] Experiments for single-modal missingness set the missingness rates to 0, 0.1, 0.2, 0.3, 0.4, and 0.5 respectively. A missingness rate of 0 means there is no single-modal missingness; a missingness rate of 0.1 means that 10% of the data samples are uncertain, and for each of these samples, one modality is randomly and completely missing (all modality data becomes 0), and another 10% of the data samples are uncertain, and for each of these samples, one modality is randomly selected, and half of the modality features of each segment of this modality are randomly missing (randomly selecting half of the feature data to become 0); a missingness rate of 0.2 means that 20% of the data samples are uncertain, and for each of these samples, one modality is randomly and completely missing, and another 20% of the data samples are uncertain, and for each of these samples, one modality is randomly selected, and half of the modality features of each segment of this modality are randomly missing; the remaining missingness rates follow the same pattern. To ensure that the experimental design fits the actual application, we first randomly select modalities for complete missingness, and then randomly select modalities from all modalities (including the completely missing ones) for partial missingness of segment features.

[0179] The experimental results are shown in Table 1. It can be found from Table 1 that for the CMU-MOSI dataset, at all missingness rates (0, 0.1, 0.2, 0.3, 0.4, and 0.5), the proposed MFAMSA model in this paper outperforms the other 11 baseline models in terms of both evaluation metrics (ACC and M-F1). Compared with the other baseline models, when the missingness rate is 0.5, the proposed MFAMSA model in this paper increases the M-F1 metric value by 7.31% to 17.75% and increases the ACC value by 8.66% to 17.03%. Compared with the second-best performing model - SMCMSA, the proposed MFAMSA model in this paper increases the M-F1 by an average of 7.21% and increases the ACC by an average of 7.43%.

[0180] In addition, for the IEMOCAP dataset, at all missingness rates (0, 0.1, 0.2, 0.3, 0.4, and 0.5), the MFAMSA model outperforms the other 11 baseline models in terms of both evaluation metrics (ACC and M-F1). Compared with the other baseline models, when the missingness rate is 0.3, the proposed MFAMSA model in this paper increases the M-F1 metric value by 4.58% to 10.48% and increases the ACC value by 5.07% to 10.29%. The proposed MFAMSA model in this paper, compared with the second-best performing model - SMCMSA, increases the M-F1 by an average of 4.26% and increases the ACC by an average of 4.87%.

[0181] Therefore, based on the experimental results in Table 1, it can be concluded that the proposed MFAMSA model in this paper has better overall performance than the other baseline models on the two public datasets.

[0182] Table 1 Experimental data for single-modal missingness

[0183]

[0184] For the experiments on multi-modal missingness, the missingness rates are set to 0, 0.1, 0.2, 0.3, 0.4, and 0.5 respectively. In this paper, on two public datasets CMU-MOSI and IEMOCAP, by setting different missingness rates, multiple experiments on the missingness of uncertain multi-modal content are carried out, and the experimental results are shown in Table 1. Among them, a missingness rate of 0 means no modal missing; a missingness rate of 0.1 means that there are 10% uncertain data samples, each sample randomly and completely misses two modalities, and there are 10% uncertain data samples, each sample randomly selects two modalities, and each segment of these two modalities randomly misses half of the modal features; a missingness rate of 0.2 means that there are 20% uncertain data samples, each sample randomly and completely misses two modalities, and there are 20% uncertain data samples, each sample randomly selects two modalities, and each segment of these two modalities randomly misses half of the modal features; and so on for the remaining missingness rates. To ensure that the experimental design fits the actual application, we first randomly select modalities for complete missing, and then randomly select modalities from all modalities (including the completely missing modalities) for partial missing.

[0185] As can be seen from Table 1, for the dataset CMU-MOSI, at all missingness rates (0, 0.1, 0.2, 0.3, 0.4, and 0.5), the proposed MFAMSA model in this paper outperforms the other 11 baseline models in two evaluation metrics (ACC and M-F1). In addition, compared with other baseline models, when the missingness rate is 0.3, the proposed MFAMSA model in this paper increases the M-F1 metric value by 7.65% to 16.37% and increases the ACC value by 9.52% to 16.93%. Compared with the second-best performing model - SMCMSA, the proposed MFAMSA model in this paper increases by 7.70% on average in M-F1 and 8.41% on average in ACC.

[0186] For the IEMOCAP dataset, at all missing rates (0, 0.1, 0.2, 0.3, 0.4, and 0.5), the proposed MFAMSA model in this paper outperforms the other 11 baseline models in terms of two evaluation metrics (ACC and M-F1). In addition, compared with other baseline models, when the missing rate is 0.5, the proposed MFAMSA model increases the M-F1 metric value by 5.40% to 11.48% and the ACC value by 7.46% to 13.82%. Compared with the second-best performing model, SMCMSA, the proposed MFAMSA model improves by an average of 4.64% in M-F1 and 6.12% in ACC. Based on the above experimental results, it can be concluded that the proposed MFAMSA model in this paper has better overall performance than other baseline models when solving the problem of multi-modal sentiment analysis under uncertain content missing.

[0187] Table 2 Multi-modal Missing Experiment Data

[0188]

[0189] Theoretical analysis. As can be seen from Table 1 and Table 2, the MCTN and TransM models have better performance than the AE and CRA models. This shows that the cyclic translation mechanism adopted in the MCTN and TransM models can extract and integrate information from different modalities more effectively than the auto-encoder mechanism in the AE and CRA models. Compared with the MCTN, MTMSA, and TransM models, the proposed MFAMSA model in this paper shows more excellent results. This is because the MFAMSA model solves the problem of the decline in the quality of modal fragments caused by partial missing of modal fragment features by using the convolutional fusion decomposition method to complete modal missing complementation, forming new fragments through multi-level key fragment interaction, and improving the fragment quality through the old and new fragment attention diffusion screening method.

[0190] As can be found from Table 1 and Table 2, on the CMU-MOSI and IEMOCAP datasets, when the missing rate is 0.5, the ACC and M-F1 of the SMCMSA model show a significant decline. The reason is that when the SMCMSA model lacks a large number of modalities, it is difficult to perform similarity search and complementation, and it is impossible to find suitable modalities for complementation.

[0191] In addition, as can be found from Table 1 and Table 2, when the missing rate is set to 0.5, the ACC and M-F1 values of the MTMSA model drop significantly. This is because the core method of the MTMSA model translates images and audio into the text modality, but when the missing rate increases, the missing situation of the text modality also increases, resulting in a decline in the translation effect and affecting the overall performance of the model.

[0192] To verify the effectiveness of different modules in the proposed MFAMSA model, this section conducts module ablation experiments. The module ablation experiments are carried out based on the CMU-MOSI dataset. The specific experimental settings and results are as follows.

[0193] Different model variants are generated by removing some key modules from MFAMSA, and the effectiveness of different modules in MFAMSA is verified by testing the performance of the model variants. The generated model variants are as follows: (1) Remove the attention diffusion screening module from MFAMSA and directly use the random sampling method for screening to generate the model variant MFAMSA-AF. (2) Remove the pyramid fusion from MFAMSA, use concatenation fusion instead and perform self-attention interaction after concatenation to generate the model variant MFAMSA-PF. (3) To verify that using two convolutional decomposition completions in the model architecture is reasonable and effective, we adjust the model to only retain one convolutional decomposition to generate the model variant MFAMSA-ONE. After convolutional decomposition completion, the three modalities are fused again and concatenated with the initial three-modal fusion result for concatenation fusion and self-attention interaction and then fed into Softmax to obtain the final sentiment analysis result. (4) To verify the effectiveness of the convolutional fusion decomposition module and the MFAMSA model architecture, we directly feed the result after concatenation fusion of the three modalities into the Softmax function to obtain the final sentiment analysis result, generating the model variant MFAMSA-ALL.

[0194] The experimental results of the module ablation experiments are shown in Table 3. As can be seen from Table 3, when the missing rate is 0.5, compared with the MFAMSA model, the M-F1 of the MFAMSA-AF model decreases by 4.04% and the ACC decreases by 6.27%. The above experimental results show that the attention diffusion screening module in the MFAMSA model is effective.

[0195] For the MFAMSA-PF model and MFAMSA-ONE, compared with MFAMSA, it can be found that the values of M-F1 and ACC decrease at different missing rates, and these results verify that the pyramid fusion module and two convolutional decomposition completions can improve the performance of MFAMSA.

[0196] Compared with MFAMSA, when the missing rate is 0.3, MFAMSA-ALL decreases by 16.70% in M-F1 and 16.41% in ACC. When the missing rate is 0.4, the M-F1 value of MFAMSA-ALL drops to 17.16%. When the missing rate is set to 0.5, the ACC value of the MFAMSA-ALL model decreases by 19.06%. These results prove the effectiveness of the convolutional fusion decomposition module and the overall MFAMSA model architecture.

[0197] Table 3 Data of module ablation experiments

[0198]

[0199] Example 2

[0200] To solve the problem of multi-modal sentiment analysis under the absence of uncertain content, this paper proposes an MSA model (MFAMSA) based on multi-level fusion completion and attention diffusion screening, as Figure 2 shown. The innovative inspiration of the MFAMSA model comes from the famous object detection model YOLOv11 (You Only Look Once). Drawing on the model architecture of YOLOv11, we designed the MFAMSA model. MFAMSA improves the quality of modal key segments through convolutional fusion decomposition completion, multi-level key segment interaction to form new segments, and new and old segment attention diffusion screening methods, effectively solving the problem of missing uncertain content.

[0201] Specifically,

[0202] this paper proposes the MFAMSA model. First, in order to retain key segments and avoid interference from low-quality segments, MFAMSA screens segments of the three modalities with missing uncertain content through the attention diffusion screening module (ADM). Then, in order to complete the completely missing modality, the three modalities after segment screening are concatenated and self-attention interaction (CS) is performed to complete the initial fusion, forming CON1; then MFAMSA re-separates the three modalities from CON1 through the convolutional fusion decomposition module (CSM) and performs CS operation with the initial three modalities, and then the new three modalities are pyramidally fused to form CON2. After that, the same decomposition operation is performed on CON2 to form CON3. To solve the problem of the quality decline of key segments caused by partial feature loss, MFAMSA completes the quality improvement of key segments through multi-level key segment interaction to form new segments and new and old segment attention diffusion screening methods. Specifically, MFAMSA performs ADM on CON3 to retain key segments to form CON4 with the same segment dimension as CON2 and performs CS to form CON5; then, ADM is performed on CON5 to form CON6 with the same segment dimension as CON1 and performs CS to form CON7. Then, CS and ADM are performed on CON5 and CON7 containing all key segments to retain the dimension consistent with the initial fusion segments, thus completing the quality improvement of key segments and forming CONF. Finally, CONF is fed into the softmax function to obtain the final sentiment prediction result. The main contributions of this paper are summarized as follows:

[0203] In practical applications, while the MSA model can solve the problem of complete absence of features in uncertain modal fragments, it also needs to be able to solve the problem of partial absence of features in uncertain modal fragments, which we call the "uncertain content missing" problem. To our knowledge, this is the first time the "uncertain content missing" problem has been clearly proposed in the field of multimodal sentiment analysis. To solve this problem, we innovatively proposed the MFAMSA model. MFAMSA improves the quality of key fragments through convolution fusion decomposition and complementation, multi-level key fragment interaction to form new fragments, and the attention diffusion screening method for new and old fragments, effectively solving the uncertain content missing problem.

[0204] To retain key modal fragments and avoid interference from low-quality fragments, we proposed an attention diffusion screening method. This method first performs multi-head attention operations on the input modality to obtain the weight matrix of the features of each modal fragment. Then, to further improve the effect of the attention mechanism and ensure that high-quality modal fragments can obtain greater weights, we introduced the feature aggregation and transformation strategies in the diffusion model to perform feature transformation, non-linear activation, and calculation and normalization of importance scores on the weight matrix. Finally, fragment screening is performed according to the weight size.

[0205] To improve the multimodal fusion effect, we proposed a pyramid fusion method to fully utilize seven modal signals for multi-level fusion. The pyramid fusion module first preliminarily fuses three single-modal features to form high-level three-modal fusion features. Subsequently, the high-level fusion features are used to guide the fusion between pairwise modalities, thereby enhancing the representation ability of the bimodal fusion features. Finally, all seven modal signals are fused. By gradually integrating multimodal information through multi-level fusion at different granularities, it ensures the gradual enhancement and refinement of information, enabling the model to capture feature interactions at different levels and comprehensively understand modal information.

[0206] To complement missing modalities, we proposed a convolution fusion decomposition method. After the three-modal fusion forms the fusion features, the fusion features at this time contain the original modal information of each modality and the information generated by cross-modal interaction. We perform a linear transformation on the fusion features through point convolution, re-decompose them into new three-modalities, and perform fusion interaction with the original three-modalities one by one to complete the complementation of missing modalities and improve the quality of each modality.

[0207] A computer-readable storage medium stores multiple instructions, and the instructions are adapted to be loaded and executed by a processor of a terminal device to perform the described method.

[0208] A terminal device includes a processor and a computer-readable storage medium. The processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor to perform the described method.

[0209] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention shall be covered within the protection scope of the present invention.

Claims

1. A multi-modal sentiment analysis method under the absence of uncertain content, characterized in that Including: Obtain multi-modal sentiment data, including three modalities of image, audio, and text; Use the multi-head attention mechanism and the diffusion model feature aggregation method to screen feature segments from the obtained multi-modal sentiment data; Perform three-layer feature fusion with different granularities on the screened features through the pyramid fusion method; Perform tri-modal decomposition on the fused features through the convolutional fusion decomposition mechanism and then fuse them with the original modalities; Obtain the sentiment prediction result; The method for screening feature segments of the obtained multimodal sentiment data by using the multi-head attention mechanism and the diffusion model feature aggregation method includes processing the input sequence by using the multi-head attention mechanism. Among them, the input sequence is mapped to a new space by linear transformation to generate a query matrix , and the query matrix is processed separately, expressed as , Among them, is the number of attention heads, Divide into sub-matrices, and the dimension of each sub-matrix is ; The method for screening feature segments of the obtained multi-modal sentiment data by using the multi-head attention mechanism and the diffusion model feature aggregation method further includes introducing the feature aggregation and transformation strategies in the diffusion model, and passing through a non-linear activation function to perform feature transformation on the query matrix to generate a new feature representation , and calculating the attention weight of each feature segment through the softmax function to complete segment aggregation, which is expressed as: , Among them, and are the weights and biases of the feature transformation, is the feature dimension of the diffusion model.

2. The multimodal sentiment analysis method under the absence of uncertain content according to claim 1, characterized in that The method for screening feature segments of the obtained multi-modal sentiment data by using the multi-head attention mechanism and the diffusion model feature aggregation method further includes splitting the attention weights by head, averaging them in the dimension of the head, retaining the unique information of each head and comprehensively considering the information of all heads to capture the temporal and spatial features of the data. Among them, the attention weights are divided into num_heads sub-matrices, the dimension of each sub-matrix is (N, T), and then an average operation is performed on the sub-matrices to smooth the attention distribution. Finally, the sequence is sorted according to the attention weights, and the first segments are screened, which is expressed as: , Among them, returns the top indices sorted in descending order, which are the filtered key segments.

3. The multimodal sentiment analysis method under the absence of uncertain content according to claim 2, wherein The three-layer feature fusion with different granularities on the screened features through the pyramid fusion method includes, through splicing and the self-attention mechanism, fusing the tri-modal features to form a high-level feature representation, that is, the tri-modal fusion modality, and using the attention diffusion screening module to perform segment screening on it; among them, performing attention diffusion screening ensures that the dimensions of the segments fused with the bi-modal are consistent, so as to better guide the fusion of pairwise modalities, expressed as: 。 4. A multimodal sentiment analysis method under the absence of uncertain content, according to claim 3, characterized in that The three-layer feature fusion with different granularities on the screened features through the pyramid fusion method includes mapping the high-level features to the same dimension as the bi-modal fusion through a fully connected layer to generate a high-level guidance signal, obtaining the attention weights between the bi-modal fusion features and the high-level features through matrix multiplication and the Softmax function, and weighting the high-level features according to the transfer score, so that the important high-level feature part for the bi-modal fusion features plays a role. Through an addition operation, the transferred high-level feature information is combined with the bi-modal fusion features and self-attention interaction is performed on them, so as to complete the fusion guidance and enhance the bi-modal fusion.

5. A multimodal sentiment analysis method under the absence of uncertain content, according to claim 4, characterized in that The three-layer feature fusion with different granularities on the screened features through the pyramid fusion method also includes performing attention diffusion screening on all modal signals to ensure that the dimensions are all (32, 25, 300) and avoid interference from low-quality segments, and then performing the final fusion, expressed as: , Among them represents three cases of pairwise modality fusion; V, A, and T represent three modalities.

6. The multimodal sentiment analysis method under the absence of uncertain content according to claim 5, wherein, The tri-modal decomposition of the fused features by the convolution fusion decomposition mechanism is then fused with the original modality, including performing a linear transformation on the input feature vector using point convolution. Given the input feature vector , a new feature vector is obtained through a 1x1 convolution operation and is expressed as: , Among them, is the convolutional kernel weight matrix, is the bias vector, is the separated image modality feature, with the shape of [batch_size, max_visual_len + max_audio_len + max_text_len, att_dim].

7. A multimodal sentiment analysis method under the absence of uncertain content, according to claim 6, wherein The tri-modal decomposition of the fused features through the convolutional fusion decomposition mechanism and then fusing them with the original modalities also includes splicing and fusing the original image modality and the convolution-separated image modality and self-attention interaction to complement the missing modality and improve the quality of the image modality; splicing and fusing the original audio modality and the convolution-separated audio modality and self-attention interaction to complement the missing modality and improve the quality of the audio modality; Splicing and fusing the original text modality and the fusion-separated text modality and self-attention interaction to complement the missing modality and improve the quality of the text modality.

8. A multi-modal sentiment analysis system under uncertain content loss, which executes a multi-modal sentiment analysis method under uncertain content loss as described in claim 1, characterized in that, Including: A data acquisition module, configured to obtain multi-modal sentiment data, including three modalities of image, audio, and text; A screening module, configured to screen feature segments from the obtained multi-modal sentiment data by using the multi-head attention mechanism and the diffusion model feature aggregation method; A fusion module, configured to perform three-layer feature fusion with different granularities on the screened features through the pyramid fusion method; A prediction module, configured to perform tri-modal decomposition on the fused features through the convolutional fusion decomposition mechanism and then fuse them with the original modalities; Obtain the sentiment prediction result.

Citation Information

Patent Citations

  • Multi-modal sentiment analysis method and system based on improved diffusion model, medium and product

    CN118035940A

  • Multi-modal sentiment analysis method and system based on fusion decomposition and trunk gathering

    CN119720102A