Multi-modal fusion sentiment analysis method and device based on inter-modal information exchange

By using technologies such as sLSTM layer, acoustic vocabulary, visual vocabulary and cross-modal embedder in multimodal sentiment analysis, the problem of insufficient information interaction between modes in traditional methods is solved, and higher reliability and accuracy of sentiment analysis are achieved.

CN120145153APending Publication Date: 2025-06-13GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510277762.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The traditional multimodal fusion emotion analysis method has problems such as insufficient information interaction between modes, noise influence and information loss, resulting in low reliability.

Method used

The multimodal fusion sentiment analysis method based on intermodal information exchange is adopted, text features are extracted through the sLSTM layer, and the acoustic vocabulary and visual vocabulary are converted into audio and video features, and information interaction and feature fusion are performed through cross-modal embedded devices and multimodal exchangers.

Benefits of technology

It improves the reliability of multimodal sentiment analysis, enhances the information interaction and effectiveness between modes, reduces information loss, and improves the accuracy and robustness of sentiment analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145153A_ABST
    Figure CN120145153A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal fusion sentiment analysis method and device based on inter-modal information exchange, and relates to the technical field of sentiment analysis, and the method comprises the steps: generating an original text feature, an original audio feature and an original video feature of to-be-analyzed multi-modal sentiment data; determining intermediate text features of the original text features by adopting an sLSTM layer; determining an audio index tag sequence feature and a video index tag sequence feature of the original audio feature and the original video feature based on the acoustic vocabulary and the visual vocabulary, and performing cross-modal interaction with the intermediate text feature to output an audio text feature and a video text feature; and adding [CLS] to the intermediate text feature, the audio text feature and the video text feature, determining a standard audio exchange text feature, a standard text exchange audio feature, a standard video exchange text feature and a standard text exchange video feature through cross-membrane information exchange, and fusing, classifying and outputting an emotion analysis result. Based on the scheme, the reliability of multi-modal fusion sentiment analysis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sentiment analysis, and particularly to a multi-modal fusion sentiment analysis method and device based on information exchange between modalities. Background Art

[0002] With the rapid development of big data and artificial intelligence technologies, sentiment analysis has gradually become an important research direction in fields such as natural language processing, computer vision, and audio analysis. Traditional sentiment analysis methods mainly rely on single-modal text data and infer sentiment tendencies by analyzing features such as vocabulary and syntactic structures. However, single-modal data often has difficulty comprehensively capturing sentiment information in complex scenarios because sentiment expressions usually have multi-modal characteristics. Multi-modal sentiment analysis is exactly a method that fuses information from multiple modalities such as text, speech, and images for sentiment analysis. Among them, the text modality can convey explicit information and semantic information, the speech modality can reflect implicit sentiment features such as intonation and speech rate, and the image modality can capture non-verbal signals such as facial expressions and gestures. By fusing information from multiple modalities, sentiment expressions can be understood more precisely. Especially when facing different cultural backgrounds, contexts, or communication methods, single-modal analysis has relatively large limitations, and multi-modal fusion can significantly improve the accuracy and robustness of sentiment analysis.

[0003] In traditional multi-modal fusion sentiment analysis methods, usually, features of each modality are first extracted, then fused using an encoder, and finally, sentiment classification is performed through a classifier. However, this method is prone to problems such as insufficient information interaction between modalities, noise influence, and information loss, resulting in low reliability of multi-modal sentiment analysis. Summary of the Invention

[0004] The present invention provides a multi-modal fusion sentiment analysis method and device based on information exchange between modalities, which improves the technical problem of low reliability in sentiment analysis of traditional multi-modal fusion sentiment analysis methods.

[0005] A multi-modal fusion sentiment analysis method based on information exchange between modalities provided by the first aspect of the present invention includes:

[0006] Obtain multi-modal sentiment data to be analyzed, extract features from the multi-modal sentiment data to be analyzed, and generate original text features, original audio features, and original video features;

[0007] Use an sLSTM layer to extract temporal features from the original text features and output intermediate text features;

[0008] Based on a preset acoustic vocabulary and visual vocabulary, add index labels to the original audio features and the original video features frame by frame respectively to determine an audio index label sequence feature and a video index label sequence feature;

[0009] The audio index label sequence features and the video index label sequence features are respectively subjected to cross-modal interaction with the intermediate text features through a cross-modal embedder, and audio-text features and video-text features are output;

[0010] Add the [CLS] flag to the intermediate text features, the audio-text features, and the video-text features, and perform cross-modal information exchange through a multi-modal exchanger to determine standard audio-exchanged text features, standard text-exchanged audio features, standard video-exchanged text features, and standard text-exchanged video features;

[0011] According to the standard audio-exchanged text features, the standard text-exchanged audio features, the standard video-exchanged text features, and the standard text-exchanged video features, a feature fuser is used for feature fusion and then input into a sentiment analysis layer for sentiment classification, and a sentiment analysis result is output.

[0012] Optionally, the obtaining of the multi-modal sentiment data to be analyzed and the feature extraction of the multi-modal sentiment data to be analyzed to generate original text features, original audio features, and original video features include:

[0013] Obtain text sentiment data to be analyzed, audio sentiment data to be analyzed, and visual sentiment data to be analyzed;

[0014] Use a text extraction model to extract the original text features of the text sentiment data to be analyzed;

[0015] Based on an audio extraction model, extract the original audio features of the audio sentiment data to be analyzed;

[0016] Extract the original video features of the visual sentiment data to be analyzed through a video extraction model.

[0017] Optionally, the construction process of the acoustic vocabulary and the visual vocabulary includes:

[0018] Obtain a video training set, and extract a plurality of training sound frames and a plurality of training video frames from the video training set;

[0019] Use the k-means algorithm to cluster each of the training acoustic frames and each of the training video frames respectively to construct a plurality of acoustic clustering centers and a plurality of video clustering centers;

[0020] Use each of the acoustic clustering centers as acoustic vocabulary to construct an acoustic vocabulary, and use each of the video clustering centers as visual vocabulary to construct a visual vocabulary.

[0021] Optionally, the cross-modal interaction between the audio index label sequence feature and the video index label sequence feature and the intermediate text feature by the cross-modal embedder, and the output of the audio-text feature and the video-text feature, includes:

[0022] The embedding layer is used to project the audio index label sequence feature and the video index label sequence feature to the same dimension as the intermediate text feature respectively, and determine the intermediate video feature and the intermediate audio feature;

[0023] The intermediate video feature is mapped to a query matrix and subjected to a cross-modal attention mechanism operation with the intermediate text feature to determine the audio-text feature;

[0024] The intermediate video feature is mapped to a query matrix and subjected to a cross-modal attention mechanism operation with the intermediate text feature to determine the video-text feature.

[0025] Optionally, adding the [CLS] flag to the intermediate text feature, the audio-text feature and the video-text feature, and performing cross-modal information exchange through the multi-modal exchanger to determine the standard audio-exchanged text feature, the standard text-exchanged audio feature, the standard video-exchanged text feature and the standard text-exchanged video feature, includes:

[0026] Adding the [CLS] flag to the intermediate text feature, the audio-text feature and the video-text feature respectively to determine the embedded text feature, the embedded audio-text feature and the embedded video-text feature;

[0027] The global information of the embedded text feature, the embedded audio-text feature and the embedded video-text feature is extracted respectively through the multi-head attention layer with residual connection, and the global text feature, the global audio feature and the global video feature are output;

[0028] The layer normalization layer is used to normalize the global text feature, the global audio feature and the global video feature to generate the normalized text feature, the normalized audio feature and the normalized video feature;

[0029] The text feature element with the minimum attention score based on the [CLS] flag in the normalized text feature is exchanged with the audio feature element with the minimum attention score based on the [CLS] flag in the normalized audio feature to determine the audio-exchanged text feature and the text-exchanged audio feature;

[0030] The text feature element with the minimum attention score based on the [CLS] flag in the normalized text feature is exchanged with the video feature element with the minimum attention score based on the [CLS] flag in the normalized video feature to determine the video-exchanged text feature and the text-exchanged video feature;

[0031] The feedforward neural network layer with residual connections respectively performs feature transformation addition on the audio-exchanged text features, the text-exchanged audio features, the video-exchanged text features, and the text-exchanged video features, and inputs them into a layer normalization layer for standardization, outputting standard audio-exchanged text features, standard text-exchanged audio features, standard video-exchanged text features, and standard text-exchanged video features.

[0032] Optionally, according to the standard audio-exchanged text features, the standard text-exchanged audio features, the standard video-exchanged text features, and the standard text-exchanged video features, a feature fusion is performed using a feature fuser and then input into a sentiment analysis layer for sentiment classification, and a sentiment analysis result is output, including:

[0033] The standard audio-exchanged text features are concatenated with the standard text-exchanged audio features, and the standard video-exchanged text features are concatenated with the standard text-exchanged video features to generate text-audio concatenated features and text-video concatenated features;

[0034] Through a gating layer, a first gated weighted feature of the text-audio concatenated features and a second gated weighted feature of the text-video concatenated features are determined;

[0035] After the first gated weighted feature and the second gated weighted feature are added element-wise, they are input into a sentiment analysis layer for sentiment classification, and a sentiment analysis result is output.

[0036] A multi-modal fusion sentiment analysis device provided in the second aspect of the present invention includes:

[0037] A feature extraction module, configured to obtain multi-modal sentiment data to be analyzed, perform feature extraction on the multi-modal sentiment data to be analyzed, and generate original text features, original audio features, and original video features;

[0038] A text processing module, configured to perform temporal feature extraction on the original text features using an sLSTM layer and output intermediate text features;

[0039] An audio-video processing module, configured to add index labels frame by frame to the original audio features and the original video features respectively based on a preset acoustic vocabulary and a visual vocabulary, and determine audio index label sequence features and video index label sequence features;

[0040] A cross-modal embedding module, configured to perform cross-modal interaction on the audio index label sequence features and the video index label sequence features with the intermediate text features respectively through a cross-modal embedder, and output audio-text features and video-text features;

[0041] A multimodal exchange module is used to add [CLS] flags to the intermediate text features, the audio text features, and the video text features, and perform cross-modal information exchange through a multimodal exchanger to determine standard audio-exchanged text features, standard text-exchanged audio features, standard video-exchanged text features, and standard text-exchanged video features;

[0042] A fusion analysis module is used to perform feature fusion on the standard audio-exchanged text features, the standard text-exchanged audio features, the standard video-exchanged text features, and the standard text-exchanged video features by using a feature fuser, and then input the fused features into a sentiment analysis layer for sentiment classification to output a sentiment analysis result.

[0043] A computer device provided in the third aspect of the present invention includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the steps of the multimodal fusion sentiment analysis method based on cross-modal information exchange as described in any one of the above.

[0044] A computer-readable storage medium provided in the fourth aspect of the present invention stores a computer program thereon. When the computer program is executed, it implements the multimodal fusion sentiment analysis method based on cross-modal information exchange as described in any one of the above.

[0045] A computer program product provided in the fifth aspect of the present invention includes a computer program / instructions. When the computer program / instructions are executed by a processor, they implement the multimodal fusion sentiment analysis method based on cross-modal information exchange as described in any one of the above.

[0046] From the above technical solutions, it can be seen that the present invention has the following advantages:

[0047] The above solution of the present invention provides a multi-modal fusion sentiment analysis method based on inter-modal information exchange, including: obtaining multi-modal sentiment data to be analyzed, extracting features from the multi-modal sentiment data to be analyzed to generate original text features, original audio features, and original video features; using an sLSTM layer to extract temporal features from the original text features and output intermediate text features; adding index labels frame by frame to the original audio features and original video features based on a preset acoustic vocabulary and visual vocabulary respectively to determine an audio index label sequence feature and a video index label sequence feature; performing cross-modal interaction on the audio index label sequence feature and the video index label sequence feature with the intermediate text features through a cross-modal embedder respectively to output an audio-text feature and a video-text feature; adding a [CLS] flag to the intermediate text features, the audio-text feature, and the video-text feature, and performing cross-modal information exchange through a multi-modal exchanger to determine a standard audio-exchanged text feature, a standard text-exchanged audio feature, a standard video-exchanged text feature, and a standard text-exchanged video feature; according to the standard audio-exchanged text feature, the standard text-exchanged audio feature, the standard video-exchanged text feature, and the standard text-exchanged video feature, using a feature fuser to perform feature fusion and then inputting it into a sentiment analysis layer for sentiment classification to output a sentiment analysis result. Based on the above solution, the temporal expression of the text modality is enhanced through the sLSTM layer, the feature transformation of the audio modality and the video modality is performed based on the acoustic vocabulary and the visual vocabulary to reduce the initial distribution difference between heterogeneous modalities, the effectiveness of the video and audio modalities is enhanced through cross-modal interaction, and the relatively irrelevant information within the modality information is excluded based on the cross-modal information exchange based on the cls global information representation, further enhancing the expression of effective information, which overall helps to improve the reliability of multi-modal fusion sentiment analysis. Description of the Drawings

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0049] Figure 1 It is a flowchart of the steps of a multi-modal fusion sentiment analysis method based on inter-modal information exchange provided by an embodiment of the present invention;

[0050] Figure 2 It is a framework diagram of the sentiment analysis network provided by an embodiment of the present invention;

[0051] Figure 3 It is a schematic diagram of vocabulary index query provided by an embodiment of the present invention;

[0052] Figure 4 Schematic diagram of transmembrane state information exchange provided by an embodiment of the present invention;

[0053] Figure 5 Structural block diagram of a multi-modal fusion sentiment analysis device based on inter-modal information exchange provided by an embodiment of the present invention. Detailed implementation manners

[0054] An embodiment of the present invention provides a multi-modal fusion sentiment analysis method and device based on inter-modal information exchange, which are used to improve the technical problem of low reliability in sentiment analysis of traditional multi-modal fusion sentiment analysis methods.

[0055] In order to make the invention objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0056] Please refer to Figure 1 , Figure 1 Step flowchart of a multi-modal fusion sentiment analysis method based on inter-modal information exchange provided by an embodiment of the present invention.

[0057] A multi-modal fusion sentiment analysis method based on inter-modal information exchange provided by the present invention includes:

[0058] Step 101: Obtain the multi-modal sentiment data to be analyzed, extract features from the multi-modal sentiment data to be analyzed, and generate original text features, original audio features, and original video features.

[0059] Step 101 includes the following sub-steps:

[0060] Obtain the text sentiment data to be analyzed, the audio sentiment data to be analyzed, and the visual sentiment data to be analyzed;

[0061] Use a text extraction model to extract the original text features of the text sentiment data to be analyzed;

[0062] Based on an audio extraction model, extract the original audio features of the audio sentiment data to be analyzed;

[0063] Extract the original video features of the visual sentiment data to be analyzed through a video extraction model.

[0064] It should be noted that in multi-modal sentiment analysis, three main modalities are involved, namely text (L), video (V), and audio (A). Therefore, the multi-modal sentiment data to be analyzed in this embodiment can include text sentiment data to be analyzed, audio sentiment data to be analyzed, and visual sentiment data to be analyzed. The multi-modal sentiment data to be analyzed can be obtained from the CMU-MOSI dataset and the CMU-MOSEI dataset. The MOSI dataset proposed by A Zadeh (2016) and the MOSEI dataset proposed by Amir Zadeh (2018) are both obtained by annotating videos with too many people removed from YouTube. Among them, the MOSEI dataset has a total of 23,453 text data and 3,228 videos in total. The text data of the MOSI dataset contains approximately 2,191 sentences, and the video data contains approximately 2,219 video segments. Both datasets label the emotional intensity of the target with [-3, 3], and they are datasets applicable in the direction of multi-modal sentiment analysis.

[0065] In this embodiment, after the sentiment analysis network designed as shown in Figure 2 is trained, multi-modal sentiment analysis is performed, including a feature extractor, a feature processor, a cross-modal embedder, a multi-modal exchanger, a feature fuser, and a sentiment analysis layer. Among them, the feature extractor includes a text extraction model, an audio extraction model, and a video extraction model.

[0066] To process the data of the text modality, in this embodiment, the text extraction model extracts the key features of the text sentiment data to be analyzed to obtain the original text features. The tensor dimension of the original text features can be [50, 768], where "50" is the sequence length in the text and 768 is the feature dimension of a single word in the text. In a preferred implementation, the SentiLARE model can be used as the text extraction model. SentilARE is a language model designed specifically for sentiment analysis. By introducing a sentiment label-enhanced sentiment-aware embedding layer, the model combines sentiment information on the basis of each word to improve the ability to understand sentiment tendencies. At the same time, the model strengthens the consistent expression of sentiment features by adding sentiment information regularization in the embedding layer. In addition, SentilARE adopts a multi-task learning strategy to jointly train the language understanding and sentiment classification tasks, which not only improves the recognition of complex emotions but also enhances the generalization ability of the model and is applicable to scenarios such as social media sentiment monitoring and user comment analysis.

[0067] For the data in the audio modality, in this embodiment, an audio extraction model is used to extract features from the audio emotion data to be analyzed to obtain the original audio features. The sequence length and feature dimension of the original audio features can be [375, 5]; in a preferred implementation, the COVAREP tool is used as the audio extraction model. COVAREP can extract 74-dimensional audio features, including 12 Mel-frequency cepstral coefficients, pitch, voiced / unvoiced segmentation features, glottal source parameters, peak slope parameters, and maximum dispersion quotient and other acoustic features. All the extracted features are related to emotions and tones.

[0068] For the data in the video modality, in this embodiment, a video extraction model is used to extract features from the video emotion data to be analyzed to obtain the original video features. In the CMU-MOSI dataset, the tensor dimension of the original video features is [500, 20]; in a preferred implementation, the Facial Action Coding System (FACS) is used as the video extraction model to extract facial action units. The FACS system is a system widely used in facial expression analysis, which can accurately identify and encode facial action units. These facial action units reflect the facial expression changes and facial postures of people in different emotional states.

[0069] Step 102: Use the sLSTM layer to perform temporal feature extraction on the original text features and output intermediate text features.

[0070] It should be noted that in order to ensure the integrity of multi-modal data and the retention of temporal information, for the text modality, sequential neural networks such as LSTM are considered to learn the discourse-level representation, and the sLSTM layer can optimize the feature extraction and temporal modeling of the modality through a hierarchical structure. In this embodiment, the sLSTM layer is set in the feature processor to perform temporal feature extraction on the original text features to obtain intermediate text features.

[0071] Step 103: Based on the preset acoustic vocabulary and visual vocabulary, index labels are added to the original audio features and original video features frame by frame to determine the audio index label sequence features and video index label sequence features.

[0072] The acoustic vocabulary refers to a database or data table composed of multiple acoustic words. Each acoustic word is mapped and associated with multiple acoustic features. Each acoustic word can represent a typical acoustic pattern or feature combination through the corresponding combination of acoustic features.

[0073] The visual vocabulary refers to a database or data table composed of multiple visual words. Each visual word is mapped and associated with multiple visual features. Each visual word can represent a typical visual pattern or feature combination through the corresponding combination of visual features.

[0074] It should be noted that in this embodiment, the feature processor not only processes the features of the text modality, but also performs feature conversion processing on the features of the audio modality and the video modality. The purpose of this feature conversion is to convert non-text features into indexes to reduce the initial distribution differences between heterogeneous modalities, which will further narrow the distribution gap between text and non-text features during fusion; as Figure 3 shown, query the original audio features according to the preset acoustic vocabulary to obtain a new sequence composed of index labels for each corresponding frame, that is, the audio index label sequence feature. Similarly, query the original audio features through the preset visual vocabulary to obtain a new sequence composed of index labels for each corresponding frame, that is, the video index label sequence feature. The specific query process can be referred to the following formula:

[0075] ;

[0076] In the formula, is the non-text modality, that is, the video modality or the audio modality, is the sequence length of the non-text modality features, is the th frame of the non-text modality index label, is the number of words in the vocabulary, is the th frame of the feature element of the non-text modality, is the th word of the non-text modality.

[0077] In a specific embodiment, the construction process of the acoustic vocabulary and the visual vocabulary includes:

[0078] Obtain a video training set, and extract multiple training sound frames and multiple training video frames from the video training set;

[0079] Use the k-means algorithm to cluster each training acoustic frame and each training video frame respectively to construct multiple acoustic clustering centers and multiple video clustering centers;

[0080] Use each acoustic clustering center as an acoustic word to construct an acoustic vocabulary, and use each video clustering center as a visual word to construct a visual vocabulary.

[0081] Step 104: Through the cross-modal embedder, perform cross-modal interaction between the audio index label sequence feature and the video index label sequence feature and the intermediate text feature respectively, and output the audio-text feature and the video-text feature.

[0082] It should be noted that the cross-modal attention mechanism is a key technology in dealing with multi-modal tasks. It can establish associations between different modalities (such as text, audio, and images) and effectively fuse information. Its basic idea is to help the model focus on relevant information in other modalities when processing information in a certain modality through an attention mechanism, so as to achieve collaboration and information transmission between modalities. In this embodiment, the cross-modal embedder embeds the features of the text modality into the features of the audio modality and the video modality respectively based on the cross-modal attention mechanism to enhance the information effectiveness of the video modality and the audio modality.

[0083] Step 104 includes the following sub-steps:

[0084] The embedding layer is used to project the audio index label sequence features and the video index label sequence features to the same dimension as the intermediate text features respectively to determine the intermediate video features and the intermediate audio features;

[0085] The intermediate video features are mapped into a query matrix and subjected to cross-modal attention mechanism operation with the intermediate text features to determine the audio text features;

[0086] The intermediate video features are mapped into a query matrix and subjected to cross-modal attention mechanism operation with the intermediate text features to determine the video text features.

[0087] It should be noted that in this embodiment, the cross-modal embedder includes an embedding layer and a cross-modal attention layer;

[0088] First, through the embedding layer, the audio index label sequence features and the video index label sequence features are embedded with learnable parameters, and they are mapped to a high-dimensional space, so as to map the features of non-text modalities to the same dimension as the intermediate text features to enable cross-modal tasks with the intermediate text features: , where is the output of the embedding layer, that is, the intermediate video features or the intermediate audio features, is the embedding layer, is the audio index label sequence features or the video index label sequence features, is the feature dimension of non-text modalities;

[0089] Then, based on the transmembrane attention layer, cross-modal attention is used to capture long-range non-text sentiment information and generate text-based non-text embeddings, namely audio-text features and video-text features; in specific implementation, in the cross-modal attention mechanism, the input of each modality is usually regarded as a sequence vector, which are respectively used to represent the feature sequence of that modality (such as the word sequence of text or the regional features of an image), and the core of cross-modal attention lies in dynamically calculating the dependence relationship between different modalities through query (Q), key (K), and value (V) matrices, and these three matrices are all obtained by linearly transforming (fully connected layer) the input, that is , and , where is the weight of the fully connected layer of the query matrix, is the weight of the fully connected layer of the key matrix, is the weight of the fully connected layer of the value matrix, is the input tensor. Thus, after mapping the intermediate video features to the query matrix, the intermediate text features to the key matrix and value matrix, the transmembrane attention mechanism operation is performed to determine the audio-text features, and after mapping the intermediate video features to the query matrix, the intermediate text features to the key matrix and value matrix, the transmembrane attention mechanism operation is performed to determine the video-text features. The specific process can refer to , where is the output of the transmembrane attention mechanism operation, is the softmax activation function, is the query matrix of the non-text modality, is the transpose of the key matrix of the text modality, is the value matrix of the text modality, is the scaling parameter; it can be understood that in the initial training stage, since the text feature representation and the non-text feature representation are in two different feature spaces and the correlation between the text feature representation and the non-text feature representation is very small, in order to better learn the model parameters, the hyperparameter is used to scale the matrix before softmax processing.

[0090] Step 105: Add the [CLS] flag to the intermediate text features, audio-text features, and video-text features, and perform cross-modal information exchange through the multi-modal exchanger to determine the standard audio-exchanged text features, standard text-exchanged audio features, standard video-exchanged text features, and standard text-exchanged video features.

[0091] The [CLS] flag is a special flag that can serve as a global representation of the entire feature sequence.

[0092] It should be noted that after adding the [CLS] flag to the intermediate text features, audio text features, and video text features, the global information of the entire feature can be better captured in the multi-modal exchanger based on the [CLS] flag, and a cross-modal information exchange mechanism is introduced to perform information interaction between the features of the audio modality and the video modality and the features of the text modality respectively, thereby enhancing the dependence between modalities.

[0093] Step 105 includes the following sub-steps:

[0094] Add the [CLS] flag to the intermediate text features, audio text features, and video text features respectively to determine the embedded text features, embedded audio text features, and embedded video text features;

[0095] Extract the global information of the embedded text features, embedded audio text features, and embedded video text features respectively through the multi-head attention layer with residual connection, and output the global text features, global audio features, and global video features;

[0096] Use the layer normalization layer to normalize the global text features, global audio features, and global video features to generate the normalized text features, normalized audio features, and normalized video features;

[0097] Exchange the text feature elements with the minimum attention score based on the [CLS] flag in the normalized text features with the audio feature elements with the minimum attention score based on the [CLS] flag in the normalized audio features to determine the audio-exchanged text features and text-exchanged audio features;

[0098] Exchange the text feature elements with the minimum attention score based on the [CLS] flag in the normalized text features with the video feature elements with the minimum attention score based on the [CLS] flag in the normalized video features to determine the video-exchanged text features and text-exchanged video features;

[0099] Use the feed-forward neural network layer with residual connection to perform feature transformation addition on the audio-exchanged text features, text-exchanged audio features, video-exchanged text features, and text-exchanged video features respectively, and input them into the layer normalization layer for standardization, and output the standard audio-exchanged text features, standard text-exchanged audio features, standard video-exchanged text features, and standard text-exchanged video features.

[0100] It should be noted that in this embodiment, as Figure 2 and Figure 4 shown, the multi-modal exchanger includes a multi-head attention layer (Multi-Head Attention) with residual connection (Add), a layer normalization layer (LayerNorm, Norm), a feature exchange layer, and a feed-forward neural network layer (Feed Forward) with residual connection;

[0101] First, add the [CLS] flag to the intermediate text features, audio text features, and video text features respectively to determine the embedded text features, embedded audio text features, and embedded video text features. The three features change from [64, 50, 768] to [64, 51, 768], and their dimensions represent (batch_size, seq_length, hidden_size); in specific implementation, assume that the modal representation obtained after the cross-modal embedder is , the modality , the input of the multi-modal exchanger is , then the process of adding the [CLS] flag can be specifically referred to , represents the cls embedding generated by each modality;

[0102] Then, calculate the multi-head attention for the embedded text features, embedded audio text features, and embedded video text features respectively through the multi-head attention layer to learn the global context information of the CLS, and perform residual connection to generate the corresponding global text features, global audio features, and global video features, and input them into the layer normalization layer for normalization, corresponding to generating the normalized text features, normalized audio features, and normalized video features;

[0103] Next, select the minimum attention score based on the attention score of [CLS] for cross-modal information exchange in the feature exchange layer; it can be understood that when calculating the multi-head attention in the multi-head attention layer, the attention score between the feature elements at any two positions is calculated for each feature. Since [CLS] can represent the global information of the entire feature sequence, the attention score of other feature elements from the perspective of [CLS] is selected in the feature exchange layer, and the minimum attention score based on the [CLS] flag is selected of the feature elements for information exchange between the text modality and the video modality, and between the text modality and the audio modality, corresponding to obtaining the audio-exchanged text features, text-exchanged audio features, video-exchanged text features, and text-exchanged video features; assume that the feature elements of one of the modalities are the row of the text modality features, and the other modality is audio, is the row of the audio modality features, is the total number of rows, then the result of the exchange is:

[0104] ;

[0105] Finally, the features output by the feature exchange layer are respectively input into the feed-forward neural network layer (Feed Forward) with cascaded residual connections and the layer normalization layer for feature processing, and the standard audio-exchanged text features, standard text-exchanged audio features, standard video-exchanged text features, and standard text-exchanged video features are output. The residual connection can reduce the information loss caused by replacement;

[0106] Based on the above cross-modal information exchange mechanism, the feature weights between modalities can be dynamically adjusted. By information exchange, relatively irrelevant information within the modality can be effectively excluded, strengthening the transmission and expression of valuable emotional information. It can effectively solve the problems of insufficient information interaction, noise influence, and information loss between modalities, and can fill the gap in the multi-modal fusion field based on the replacement strategy. The purpose of this strategy is to improve the accuracy of emotional intensity, thereby assisting in enhancing the subsequent emotional classification performance.

[0107] Step 106: According to the standard audio-exchanged text features, standard text-exchanged audio features, standard video-exchanged text features, and standard text-exchanged video features, use a feature fusion device to perform feature fusion and then input it into the emotional analysis layer for emotional classification, and output the emotional analysis result.

[0108] Step 106 includes the following sub-steps:

[0109] Concatenate the standard audio-exchanged text features with the standard text-exchanged audio features, and the standard video-exchanged text features with the standard text-exchanged video features to generate text-audio concatenated features and text-video concatenated features;

[0110] Through the gating layer, determine the first gated weighted feature of the text-audio concatenated features and the second gated weighted feature of the text-video concatenated features;

[0111] After adding the first gated weighted feature and the second gated weighted feature element by element, input it into the emotional analysis layer for emotional classification, and output the emotional analysis result.

[0112] It should be noted that after inputting the multiple features output by the multi-modal exchanger into the feature fusion device for feature fusion, the emotional analysis result can be output by performing emotional classification through the emotional analysis layer;

[0113] In specific implementation, the feature fusion device of this embodiment includes a gating layer and a fusion layer, and the sentiment analysis layer may include a linear layer; first, the fusion layer splices the two pairs of features output by the multi-modal exchanger respectively, that is, splicing the standard audio-exchanged text features and the standard text-exchanged audio features, and splicing the standard video-exchanged text features and the standard text-exchanged video features, respectively obtaining text-audio spliced features and text-video spliced features; then, based on the gating method, the text-audio spliced features and the text-video spliced features are used to dynamically adjust the contributions of each modality to control and screen these spliced representations. Taking the text-audio spliced features as an example, first calculate the text-audio gating value of the text-audio spliced features of the text-audio spliced features , where is the text-audio gating weight, is the text-audio gating bias, is the sigmoid activation function, and both the weight and the bias are learnable parameters. The obtained text-audio gating value is a proportional coefficient within the range of 0 to 1. Multiplying the text-audio gating value by the text-audio spliced features gives the first gated weighted feature . Similar to the determination process of the first gated weighted feature, the second gated weighted feature of the text-video spliced features is constructed. By introducing the feature gating mechanism, the utilization efficiency of multi-modal information is further improved, thus helping to improve the accuracy and robustness of sentiment analysis. Finally, after using additive fusion in the fusion layer to fuse the first gated weighted feature and the second gated weighted feature, it is input into the sentiment analysis layer for sentiment classification to output the sentiment analysis result; it can be understood that in this embodiment, 7-class and 2-class sentiment intensities can be adopted. For the 7-class sentiment intensity, after the output of the feature fusion device, it is normalized by the softmax function and then a unique value is output by the Linear, and finally projected into the [-3, 3] interval for sentiment intensity judgment. For the binary classification task, the output result only needs to be positive or negative.

[0114] To verify the effectiveness of the above method, we conducted experiments on the MOSI and MOSEI datasets, and the experimental results are compared as shown in Table 1:

[0115] Table 1 Comparison of experimental results of different models on MOSI and MOSEI

[0116]

[0117] The data of TFN, MFN, SWAFN, and MULT in this table are from MMSA. According to the scores measured by our method on the MOSI dataset and the MOSEI dataset, our method has achieved certain results.

[0118] In the embodiments of the present invention, the temporal expression of the text modality is enhanced through the sLSTM layer. The audio modality and the video modality are subjected to feature transformation based on the acoustic vocabulary and the visual vocabulary to reduce the initial distribution difference between heterogeneous modalities. The effectiveness of the video and audio modalities is enhanced through cross-modal interaction. Based on the cross-modal information exchange based on the cls global information representation, relatively irrelevant information within the modality information is excluded, further enhancing the expression of effective information. Overall, it helps to improve the reliability of multi-modal fusion sentiment analysis. This novel cross-modal interaction method also provides a new thinking direction for the field of deep learning based on multi-modal fusion, and is of great significance for fields such as emotion recognition, user experience evaluation, and mental health monitoring.

[0119] Please refer to Figure 5 , Figure 5 which is a structural block diagram of a multi-modal fusion sentiment analysis device based on information exchange between modalities provided by the embodiments of the present invention.

[0120] A multi-modal fusion sentiment analysis device based on information exchange between modalities provided by the present invention includes:

[0121] A feature extraction module 501, configured to obtain multi-modal sentiment data to be analyzed, extract features from the multi-modal sentiment data to be analyzed, and generate original text features, original audio features, and original video features;

[0122] A text processing module 502, configured to perform temporal feature extraction on the original text features by using the sLSTM layer and output intermediate text features;

[0123] An audio-video processing module 503, configured to add index labels to the original audio features and the original video features frame by frame based on a preset acoustic vocabulary and visual vocabulary, and determine audio index label sequence features and video index label sequence features;

[0124] A cross-modal embedding module 504, configured to perform cross-modal interaction on the audio index label sequence features and the video index label sequence features with the intermediate text features respectively through a cross-modal embedder, and output audio-text features and video-text features;

[0125] A multi-modal exchange module 505, configured to add [CLS] flags to the intermediate text features, the audio-text features, and the video-text features, and perform cross-modal information exchange through a multi-modal exchanger to determine standard audio-exchanged text features, standard text-exchanged audio features, standard video-exchanged text features, and standard text-exchanged video features;

[0126] The fusion analysis module 506 is configured to perform feature fusion using a feature fuser based on standard audio-to-text features, standard text-to-audio features, standard video-to-text features, and standard text-to-video features, and then input the fused features into the sentiment analysis layer for sentiment classification, and output the sentiment analysis result.

[0127] Optionally, the feature extraction module 501 is specifically configured to:

[0128] Obtain the text sentiment data to be analyzed, the audio sentiment data to be analyzed, and the visual sentiment data to be analyzed;

[0129] Extract the original text features of the text sentiment data to be analyzed using a text extraction model;

[0130] Extract the original audio features of the audio sentiment data to be analyzed based on an audio extraction model;

[0131] Extract the original video features of the visual sentiment data to be analyzed through a video extraction model.

[0132] Optionally, the construction process of the acoustic vocabulary and the visual vocabulary includes:

[0133] Obtain a video training set, and extract multiple training sound frames and multiple training video frames from the video training set;

[0134] Use the k-means algorithm to cluster each training acoustic frame and each training video frame respectively, and construct multiple acoustic clustering centers and multiple video clustering centers;

[0135] Use each acoustic clustering center as an acoustic vocabulary to construct an acoustic vocabulary, and use each video clustering center as a visual vocabulary to construct a visual vocabulary.

[0136] Optionally, the cross-modal embedding module 504 is specifically configured to:

[0137] Use an embedding layer to project the audio index label sequence feature and the video index label sequence feature to the same dimension of the intermediate text feature respectively, and determine the intermediate video feature and the intermediate audio feature;

[0138] Map the intermediate video feature to a query matrix and perform a cross-modal attention mechanism operation with the intermediate text feature to determine the audio-text feature;

[0139] Map the intermediate video feature to a query matrix and perform a cross-modal attention mechanism operation with the intermediate text feature to determine the video-text feature.

[0140] Optionally, the multi-modal exchange module 505 is specifically configured to:

[0141] Add the [CLS] flag to the intermediate text features, audio text features, and video text features respectively to determine the embedded text features, embedded audio text features, and embedded video text features;

[0142] Use the multi-head attention layer with residual connection to extract global information from the embedded text features, embedded audio text features, and embedded video text features respectively, and output the global text features, global audio features, and global video features;

[0143] Use the layer normalization layer to normalize the global text features, global audio features, and global video features to generate the normalized text features, normalized audio features, and normalized video features;

[0144] Exchange the text feature elements with the minimum attention score based on the [CLS] flag in the normalized text features and the audio feature elements with the minimum attention score based on the [CLS] flag in the normalized audio features to determine the audio-exchanged text features and text-exchanged audio features;

[0145] Exchange the text feature elements with the minimum attention score based on the [CLS] flag in the normalized text features and the video feature elements with the minimum attention score based on the [CLS] flag in the normalized video features to determine the video-exchanged text features and text-exchanged video features;

[0146] Use the feed-forward neural network layer with residual connection to perform feature transformation and addition on the audio-exchanged text features, text-exchanged audio features, video-exchanged text features, and text-exchanged video features respectively, and input them into the layer normalization layer for standardization, and output the standard audio-exchanged text features, standard text-exchanged audio features, standard video-exchanged text features, and standard text-exchanged video features.

[0147] Optionally, the fusion analysis module 506 is specifically used for:

[0148] Concatenate the standard audio-exchanged text features with the standard text-exchanged audio features, and the standard video-exchanged text features with the standard text-exchanged video features to generate the text-audio concatenated features and text-video concatenated features;

[0149] Determine the first gated weighted feature of the text-audio concatenated features and the second gated weighted feature of the text-video concatenated features through the gated layer;

[0150] After adding the first gated weighted feature and the second gated weighted feature element by element, input them into the sentiment analysis layer for sentiment classification, and output the sentiment analysis result.

[0151] An embodiment of the present invention also provides a computer device, including a memory and a processor, where a computer program is stored in the memory; when the computer program is executed by the processor, the processor executes the steps of the multi-modal fusion sentiment analysis method based on inter-modal information exchange according to any one of the above embodiments.

[0152] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program / instruction is stored, and when the computer program / instruction is executed by a processor, the steps of the multi-modal fusion sentiment analysis method based on inter-modal information exchange according to any one of the above embodiments are implemented.

[0153] An embodiment of the present invention also provides a computer program product, including a computer program / instruction, and when the computer program / instruction is executed by a processor, the steps of the multi-modal fusion sentiment analysis method based on inter-modal information exchange according to any one of the above embodiments are implemented.

[0154] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0155] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0156] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0157] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0158] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0159] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A multimodal fusion sentiment analysis method based on inter-modal information exchange, characterized in that: include: Acquire multimodal emotion data to be analyzed, perform feature extraction on the multimodal emotion data to be analyzed, and generate original text features, original audio features, and original video features; The sLSTM layer is used to extract temporal features from the original text features and output intermediate text features; Based on a preset acoustic vocabulary and a preset visual vocabulary, index labels are added to the original audio features and the original video features frame by frame to determine audio index label sequence features and video index label sequence features; The audio index tag sequence feature and the video index tag sequence feature are cross-modally interacted with the intermediate text feature by a cross-modal embedder, and an audio text feature and a video text feature are output; Adding a [CLS] mark to the intermediate text feature, the audio text feature and the video text feature, and determining a standard audio exchange text feature, a standard text exchange audio feature, a standard video exchange text feature and a standard text exchange video feature through a multimodal exchange trans-membrane state information exchange; According to the standard audio-exchange text feature, the standard text-exchange audio feature, the standard video-exchange text feature and the standard text-exchange video feature, a feature fusion device is used to perform feature fusion, and then the features are input into a sentiment analysis layer for sentiment classification, and a sentiment analysis result is output.

2. The multimodal fusion sentiment analysis method based on inter-modal information exchange according to claim 1 is characterized in that: The step of obtaining the multimodal emotion data to be analyzed, performing feature extraction on the multimodal emotion data to be analyzed, and generating original text features, original audio features, and original video features includes: Obtain text emotion data to be analyzed, audio emotion data to be analyzed, and visual emotion data to be analyzed; Use the text extraction model to extract the original text features of the text sentiment data to be analyzed; Extracting original audio features of the audio emotion data to be analyzed based on the audio extraction model; The original video features of the visual emotion data to be analyzed are extracted through the video extraction model.

3. The multimodal fusion sentiment analysis method based on inter-modal information exchange according to claim 1 is characterized in that: The process of constructing the acoustic vocabulary and the visual vocabulary includes: Acquire a video training set, and extract a plurality of training sound frames and a plurality of training video frames from the video training set; Using a k-means algorithm to cluster the training acoustic frames and the training video frames respectively, and construct multiple acoustic clustering centers and multiple video clustering centers; An acoustic vocabulary is constructed using each of the acoustic cluster centers as an acoustic vocabulary, and a visual vocabulary is constructed using each of the video cluster centers as a visual vocabulary.

4. The multimodal fusion sentiment analysis method based on inter-modal information exchange according to claim 1 is characterized in that: The step of performing cross-modal interaction between the audio index tag sequence feature and the video index tag sequence feature and the intermediate text feature through a cross-modal embedder to output audio text features and video text features includes: Using an embedding layer to project the audio index label sequence features and the video index label sequence features to the same dimension as the intermediate text features, respectively, to determine intermediate video features and intermediate audio features; Mapping the intermediate video features into a query matrix and performing a cross-membrane attention mechanism operation on the intermediate text features to determine the audio text features; The intermediate video features are mapped into a query matrix and a cross-membrane attention mechanism operation is performed with the intermediate text features to determine the video text features.

5. The multimodal fusion sentiment analysis method based on inter-modal information exchange according to claim 1 is characterized in that: The step of adding a [CLS] mark to the intermediate text feature, the audio text feature, and the video text feature, and exchanging information across the multimodal exchanger to determine a standard audio exchange text feature, a standard text exchange audio feature, a standard video exchange text feature, and a standard text exchange video feature includes: Adding [CLS] marks to the intermediate text features, the audio text features, and the video text features, respectively, to determine embedded text features, embedded audio text features, and embedded video text features; Performing global information extraction on the embedded text features, the embedded audio text features, and the embedded video text features respectively through a residually connected multi-head attention layer, and outputting global text features, global audio features, and global video features; Normalizing the global text features, the global audio features, and the global video features using a layer normalization layer to generate normalized text features, normalized audio features, and normalized video features; The text feature element of the minimum attention score based on the [CLS] mark in the normalized text feature is exchanged with the audio feature element of the minimum attention score based on the [CLS] mark in the normalized audio feature to determine an audio-exchanged text feature and a text-exchanged audio feature; The text feature element with the minimum attention score based on the [CLS] mark in the normalized text feature is exchanged with the video feature element with the minimum attention score based on the [CLS] mark in the normalized video feature to determine a video exchange text feature and a text exchange video feature; A feedforward neural network layer using residual connection performs feature transformation and addition on the audio-exchange text features, the text-exchange audio features, the video-exchange text features and the text-exchange video features respectively, and inputs a normalization layer for standardization, and outputs standard audio-exchange text features, standard text-exchange audio features, standard video-exchange text features and standard text-exchange video features.

6. The multimodal fusion sentiment analysis method based on inter-modal information exchange according to claim 1 is characterized in that: The step of using a feature fusion device to perform feature fusion according to the standard audio-exchange text feature, the standard text-exchange audio feature, the standard video-exchange text feature, and the standard text-exchange video feature, and then inputting the feature into a sentiment analysis layer for sentiment classification, and outputting a sentiment analysis result, comprises: Splicing the standard audio-exchange text feature with the standard text-exchange audio feature, and the standard video-exchange text feature with the standard text-exchange video feature to generate a text-audio splicing feature and a text-video splicing feature; Determining, through a gating layer, a first gated weighted feature of the text-audio splicing feature and a second gated weighted feature of the text-video splicing feature; After adding the first gated weighted feature and the second gated weighted feature element by element, the feature is input into the sentiment analysis layer for sentiment classification, and the sentiment analysis result is output.

7. A multimodal fusion sentiment analysis device based on inter-modal information exchange, characterized in that: include: A feature extraction module is used to obtain multimodal emotion data to be analyzed, perform feature extraction on the multimodal emotion data to be analyzed, and generate original text features, original audio features, and original video features; The text processing module is used to extract temporal features from the original text features using the sLSTM layer and output intermediate text features; An audio and video processing module, used to add index tags to the original audio features and the original video features frame by frame based on a preset acoustic vocabulary and a preset visual vocabulary, and determine audio index tag sequence features and video index tag sequence features; A cross-modal embedding module, used to perform cross-modal interaction between the audio index tag sequence feature and the video index tag sequence feature and the intermediate text feature through a cross-modal embedder, and output audio text features and video text features; A multimodal exchange module, used for adding a [CLS] mark to the intermediate text feature, the audio text feature and the video text feature, and determining a standard audio exchange text feature, a standard text exchange audio feature, a standard video exchange text feature and a standard text exchange video feature through a multimodal exchange trans-membrane state information exchange; The fusion analysis module is used to use a feature fuser to perform feature fusion according to the standard audio-exchange text feature, the standard text-exchange audio feature, the standard video-exchange text feature and the standard text-exchange video feature, and then input the feature into the sentiment analysis layer for sentiment classification, and output the sentiment analysis result.

8. A computer device, characterized in that: It includes a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the multimodal fusion sentiment analysis method based on inter-modal information exchange as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by the processor, the steps of the multimodal fusion sentiment analysis method based on inter-modal information exchange as described in any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by the processor, the steps of the multimodal fusion sentiment analysis method based on inter-modal information exchange as described in any one of claims 1 to 6 are implemented.