A multimodal sentiment analysis method based on disentangled representation learning
Through the multimodal sentiment analysis method based on untangling representation learning, the problems of many noise and redundant information, difficulty in cross-modal learning, and underutilized language modality in multimodal sentiment analysis in the prior art are solved, and the deep fusion and sentiment analysis of audio, video and text modalities are achieved, which significantly improves the analysis effect.
Patent Information
- Application Number
- CN202510158857.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-13
AI Technical Summary
The existing multimodal sentiment analysis technology has problems such as a lot of noise and redundant information in audio and video features, difficulty in cross-modal learning, insufficient role of language modality in emotional expression, and low efficiency in feature extraction, fusion and emotional-related information recognition between data of different modality.
The multimodal sentiment analysis method based on untangling representation learning is adopted, and deep fusion and sentiment analysis of audio, video and text modal modalities are achieved through steps such as data collection, multimodal feature extraction, construction of untangling representation learning network, modal untangling and adversarial optimization, time smoothing constraints and feature fusion.
It significantly reduces redundant information in audio and video features, improves the consistency of time dimensions, makes full use of the importance of text features, improves the cross-modal fusion effect, and shows competitive performance on MOSI and MOSIE datasets.
Smart Images

Figure CN119622280B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of natural language processing, and in particular relates to a multimodal sentiment analysis method based on disentanglement representation learning. Background Art
[0002] With the rapid development of intelligent human-computer interaction, intelligent medical care and robotics, the application scope of multimodal data has been greatly expanded, providing a strong driving force for the research of Multimodal Sentiment Analysis (MSA). Human emotional expression is usually an organic fusion of natural language, facial expressions and vocal behavior. Emotion plays an important role in promoting human interaction, shaping communication methods and assisting decision-making. As an important research direction, multimodal sentiment analysis is committed to understanding human emotional information from multiple dimensions by parsing multimodal data such as text, audio and video, so that machines can more comprehensively grasp human emotional states. By accurately capturing these emotional signals, machines can interact with humans more effectively and produce more natural and deeper emotional resonance.
[0003] In view of the broad application prospects of multimodal sentiment analysis, current researchers have also proposed a large number of multimodal sentiment analysis models to try to solve the core problems in the field. However, existing multimodal sentiment analysis has many shortcomings: (1) Audio and video features often contain a lot of noise and redundant information, which makes it difficult to accurately reflect the core content of emotions. (2) There are significant differences between discrete text features and continuous audio and video features, which leads to huge challenges in cross-modal learning. (3) Existing research has not paid enough attention to the key role of language modality in emotional expression and has failed to fully explore its potential value. (4) The efficiency of feature extraction, fusion and recognition of emotion-related information between different modal data is low, which makes it difficult to accurately analyze the emotional polarity of each entity in multimodal data. Summary of the invention
[0004] Purpose of the invention: The technical problem to be solved by the present invention is to provide a multimodal sentiment analysis method based on disentangled representation learning in view of the shortcomings of the prior art, comprising the following steps:
[0005] Step 1, data collection: collect data sets for multimodal sentiment analysis. The selected data sets should be representative and diverse, cover multiple modalities such as audio, video and text, and be able to support subsequent training, evaluation and testing processes; the multimodalities include text modality, audio modality and video modality;
[0006] Step 2: Extract multimodal features: Extract text, audio, and video modal features for the input video data; use the Bert model to extract text model features , using Librosa tool to extract audio model features , use OpenFace to extract facial features and get video model features , for the extracted modal features , transformed to the same feature dimension through linear mapping, and the aligned modal features are obtained ; ;
[0007] Step 3: Construct a disentangled representation learning network, including: Generate modality-private features: Input to the modality-private encoder to generate the corresponding modality-private features , , ; Generate modal shared features: Modal features Input the shared encoder at the same time to generate modality shared features ; For modal shared features , , Perform full connection processing to obtain global shared features ;
[0008] Step 4: Modal disentanglement and adversarial optimization: Modal private features , , Shared features with modalities Input modality discriminator, optimize the parameters of modality discriminator, private encoder and shared encoder;
[0009] Step 5, temporal smoothness constraint: private features of audio modality , audio modality shared features and video modality private features , Video modality shared features Applying temporal smoothness constraints to reduce redundancy and noise in audio and video features over consecutive time periods;
[0010] Step 6, feature fusion: Deeply fuse audio and video modality features through text-guided mechanism, and finally generate fusion features for sentiment analysis ;
[0011] Step 7, model optimization and output: Fusion features Input the sentiment analysis module and combine it with sentiment classification or regression tasks to generate sentiment prediction results; at the same time, optimize the sentiment analysis module through the loss function of the training phase.
[0012] Step 2 includes:
[0013] For the input video data, use the Bert model to extract text model features , using Librosa tool to extract audio model features , use OpenFace to extract facial features and get video model features :
[0014] ,
[0015] ,
[0016] ,
[0017] Feature Set ;
[0018] For each modal feature Perform linear mapping to unify to the same feature dimension to obtain the aligned feature representation :
[0019] ,
[0020] Where m represents the mode, , For text, For audio, For video, and Represent the weight matrix and bias matrix of the linear mapping respectively.
[0021] Step 3 includes:
[0022] Modal characteristics are input into the modality-specific encoder to generate text, audio, and video modality-specific features respectively. , , , modal characteristics They are simultaneously input into the shared encoder to generate shared features for text, audio, and video modalities. , where the private encoder extracts features from each modal data through three independent encoders, and the shared encoder processes all modal data using the same encoder. The mathematical expression is:
[0023] ,
[0024] ,
[0025] in, , , They are text private encoders Parameters, audio private encoder Parameters of private video encoder Parameters, Is a shared encoder Parameters;
[0026] Shared Features , , Perform full connection processing to obtain global shared features :
[0027] ,
[0028] in, is the weight matrix, is the activation function.
[0029] Step 4 includes:
[0030] Modality Discriminator It is to classify and judge the feature representation of the input data and output the corresponding probability. The formula is:
[0031] ,
[0032] in, , is the weight matrix of the modality discriminator, T represents the transposed is the bias vector; is the input to the modality discriminator, is the size of the dimension, express dimensional real space, the output of the modal discriminator Represents and corresponds to three modes The probability of correlation;
[0033] The arccosine-based cross entropy loss function is used, which is defined as:
[0034] ,
[0035] ,
[0036] ,
[0037] in, is the target modality of the input sample, is the weight vector of mode m, is the target mode The weight vector of is the input feature vector, Represents the input feature vector and mode The angle of Represents the input feature vector and the target mode The angle of represents the cross entropy loss; is the scaling factor, is a marginal parameter; e is a natural constant;
[0038] Take the weighted sum for each label and calculate the overall adversarial loss:
[0039] ,
[0040] ,
[0041] in, is the batch size, Represents modal shared features With target mode The cross entropy loss between ; is the adversarial loss of shared features, is the adversarial loss of private features; Represents modal private features With target mode The cross entropy loss between .
[0042] Step 5 includes:
[0043] The temporal smoothness loss based on KL divergence measures the changing characteristics of audio and video modalities in continuous time. KL divergence is defined as:
[0044] ,
[0045] in , are two probability distributions, and They are , In the The value at the position; Indicates an event In distribution Probability and distribution under The difference in probability of the following;
[0046] For audio features, the KL divergence of three consecutive frames of audio features is calculated. The formula is:
[0047] ,
[0048] ,
[0049] ,
[0050] in, and Respectively represent Frame audio modality private features and Frame audio modality shared features, yes The processed results are averaged. Indicates the audio feature in The private characteristics of a frame are passed The average value after operation; Indicates the audio feature in The shared features of the frames are obtained by The average value after operation; Indicates the frame number. The audio feature is The sliding average of the private features of the frame, The audio feature is The sliding average of the shared features of the frames, , , is the temporal smoothness loss of audio features;
[0051] For video features, the KL divergence of three consecutive frames of audio features is calculated. The formula is:
[0052] ,
[0053] ,
[0054] ,
[0055] in, and Respectively represent Frame video modality private features and Frame video modality shared features, Indicates the video feature in The private characteristics of a frame are passed The average value after operation, Indicates the video feature in The shared features of the frames are obtained by The average value after operation; The video feature is The sliding average of the private features of the frame, The video feature is The sliding average of the shared features of the frames, is the temporal smoothness loss of video features;
[0056] Construct a consistency loss function, the formula is:
[0057] ,
[0058] ,
[0059] ,
[0060] in, and They represent the feature differences between consecutive frames of the video modality and the feature differences between consecutive frames of the audio modality, It is Frame and The average value of + , is the consistency loss;
[0061] The overall time smoothness loss is calculated by the following formula :
[0062] .
[0063] Step 6 includes: Initializing a learnable parameter tensor , and then the low-scale text features , audio modality private features , Video modality private features and The text-guided module includes a cross-modal attention mechanism for audio features and text-modal features. is used as the query vector Q, the private features of the audio modality As the key vector K and value vector V respectively, we get and The similarity matrix :
[0064] ,
[0065] in, and is the learnable parameter matrix, Represents the dimension of each attention head;
[0066] Calculating text features and video features The similarity matrix :
[0067] ,
[0068] in is a learnable parameter matrix;
[0069] Then generate new fusion features :
[0070] ,
[0071] in , and is the learnable parameter matrix;
[0072] Will Input into the first Transformer layer, and generate medium-scale text features through deep modeling :
[0073] ,
[0074] in, represents the first Transformer layer, are the parameters of the first Transformer layer;
[0075] Then the new fusion feature and mesoscale text features Input into the next text-guided module to obtain fusion features :
[0076] ,
[0077] in , , , , , and is a learnable parameter matrix;
[0078] Then the mesoscale text features Input to the second Transformer layer to extract high-scale text features ,then and It is input into the last text-guided module to obtain the final fusion features of audio features and video features. :
[0079] ,
[0080] ,
[0081] in, represents the second Transformer layer, are the parameters of the second Transformer layer, , , is a learnable parameter matrix;
[0082] Further modeling through the feedforward network FFN, the final fusion features are obtained:
[0083] ,
[0084] ,
[0085] ,
[0086] in, , , , is the learnable parameter matrix, is the output after calculation by the cross-modal self-attention mechanism, yes After normalization, the output Represent text features at different scales;
[0087] hour, , , is a low-scale text feature With fusion features The learnable parameter matrix in the cross-modal attention mechanism, yes and The output after calculation by the cross-modal self-attention mechanism, yes Output after normalization;
[0088] hour, , , is the mesoscale text feature With fusion features The learnable parameter matrix in the cross-modal attention mechanism, yes and The output after calculation by the cross-modal self-attention mechanism, yes Output after normalization;
[0089] hour, , , It is a high-scale text feature With fusion features The learnable parameter matrix in the cross-modal attention mechanism, yes and The output after calculation by the cross-modal self-attention mechanism, yes The output after normalization.
[0090] represents normalization, represents a feed-forward network; , , They represent the multimodal fusion features generated at low, medium, and high scales respectively;
[0091] Fusion features at different scales and shared modal features Weighted fusion via gating mechanism:
[0092] ,
[0093] ,
[0094] ,
[0095] ,
[0096] in, ,Right now , , , , and is a learnable weight matrix, yes Through the weight matrix The result after transformation, Refers to sum pooling, is the size of the pooling window, yes After summing and pooling, express The result after L2 normalization is: express The second norm of hour, yes The learnable weight matrix, yes Through the weight matrix The result after transformation, yes After summing and pooling, express The result after L2 normalization;
[0097] Enter the Multiply it element by element with the weight matrix, and then sum and pool the matrix after element-by-element multiplication to get , then L2 normalization is used to ensure the same scale, and finally, linear mapping weighted summation is performed to obtain the final fusion feature .
[0098] Step 7 includes:
[0099] The fusion features As input, it is passed to the multi-layer perceptron , output the sentiment prediction results through nonlinear transformation :
[0100] ,
[0101] The mean square error (MSE) is used as the loss function for task learning. , the calculation formula is:
[0102] ,
[0103] in, is the sentiment label of the zth sample, is the prediction result of the sentiment analysis module, is the number of samples in the dataset.
[0104] In step 7, the final total loss function is for:
[0105] ,
[0106] in, , and is a hyperparameter.
[0107] The present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the described method.
[0108] The present invention also provides a storage medium storing a computer program or instruction. When the computer program or instruction is run on a computer, the steps of the method described are executed.
[0109] Compared with the prior art, the present invention has the following beneficial effects: (1) It adopts temporal smoothness constraints to ensure the consistency of audio and video features in continuous time, significantly reduce redundant information, and improve the consistency of the temporal dimension.
[0110] (2) Introducing a text guidance module, which effectively integrates audio, video, and text modalities through guidance at different language scales, achieving deep collaboration and integration.
[0111] (3) Make full use of the importance of text features and integrate multimodal information of different scales through a deep cross-modal interaction layer, significantly improving the cross-modal fusion effect.
[0112] (4) The proposed method shows competitive performance on the MOSI and MOSIE datasets, verifying its effectiveness and advantages. BRIEF DESCRIPTION OF THE DRAWINGS
[0113] Figure 1 Flow chart of the method of the present invention.
[0114] Figure 2 It is a model framework diagram of the system of the present invention.
[0115] Figure 3 Flowchart of the text-guided module.
[0116] Figure 4 Flowchart of the cross-modal interaction layer. DETAILED DESCRIPTION
[0117] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.
[0118] This embodiment provides a multimodal sentiment analysis method based on disentanglement representation learning. The overall architecture of the system model provided in this embodiment is as follows: Figure 2 As shown in the figure, the model consists of three modules: (1) feature extraction module; (2) disentanglement representation learning module; and (3) fusion module.
[0119] like Figure 1 As shown, the method comprises the following steps:
[0120] Step 1, data collection: Collect training, evaluation and test datasets that can be used for multimodal sentiment analysis. The dataset contains data in three modalities: text, audio and video. Preprocess the dataset to meet the needs of subsequent feature extraction. The "multimodal" refers to text modality, audio modality and video modality.
[0121] Step 2: Extract multimodal features: Extract text, audio, and video modal features for the input video data. Text model features Using Bert model to extract audio model features Use Librosa tool to extract features and video model features Use OpenFace to extract facial features and modal features , transformed to the same feature dimension through linear mapping, and the aligned modal features are obtained .
[0122] Step 3: Construct a disentangled representation learning network: Modality private feature generation: Modality features Input to the modality-private encoder to generate the corresponding modality-private features , , . Modal shared feature generation: Modal features Input the shared encoder at the same time to generate modality shared features . Shared Features , , Perform full connection processing to obtain global shared features .
[0123] Step 4: Modal disentanglement and adversarial optimization: Generate modal private features , , Shared features with modalities Input modality discriminator. By optimizing the parameters of modality discriminator, private encoder and shared encoder, we ensure that modality private features only contain modality-specific information, while modality shared features learn shared information between modalities, thus achieving disentangled representation learning of modality features.
[0124] Step 5, time smoothness constraint: audio modal features , and video modality features , Temporal smoothness constraints are applied to reduce redundancy and noise in audio and video features over consecutive time segments.
[0125] Step 6, feature fusion: The present invention realizes the deep fusion of audio and video modality features through a text-guided mechanism.
[0126] Low-scale fusion: low-scale text features (Right now ) and audio and video modality features and their initialization Input into the text guidance module, generate new fusion features through cross-modal attention mechanism and full connection operation .
[0127] Mesoscale Fusion: Input the first Transformer layer to generate mesoscale text features through deep modeling . Then, and The data are input into the next text guidance module, and the fusion features are generated through similar cross-modal attention mechanism and full connection processing. .
[0128] High-scale fusion: Input to the second Transformer layer to extract high-scale text features . Then, and Input to the final text-guided module to obtain the final fusion features of audio and video features .
[0129] Cross-modal interaction: text features at different scales , , Final fusion features with audio and video Interactive fusion layer by layer to generate intermediate fusion features These features are further combined with the shared features Input to the gating mechanism for processing, thereby filtering redundant information and dynamically fusing multimodal effective features, and finally generating fusion features for sentiment analysis .
[0130] Step 7, model optimization and output: by fusion features The data is input into the subsequent sentiment analysis module and combined with sentiment classification or regression tasks to generate sentiment prediction results. At the same time, the prediction accuracy and generalization ability of the model are ensured through the optimization of the loss function in the training phase.
[0131] Wherein, step 1 comprises:
[0132] In the study of multimodal sentiment analysis, data collection is a key step in model training, verification and testing. In order to ensure the scientific nature of the research and the applicability of the data, this paper selects two representative and high-quality multimodal datasets: MOSI and MOSEI. The MOSI dataset was proposed by Amir Zadeh et al. in the journal IEEE Intelligent Systems in 2016 and described in detail in their paper “Amir Zadeh, Rowan Zellers, Eli Pincus, Louis-Philippe Morency, Multimodal sentiment intensity analysis in videos: Facialgestures and verbal messages, IEEE Intell. Syst. 31 (6) (2016) 82–88”. The MOSI dataset contains 93 video clips from YouTube, which cover a wide range of emotional expressions, including positive, negative and neutral emotions. The dataset contains a total of 2199 samples, as shown in Table 1. Each sample provides three modal features: text, audio, and video. At the same time, it is annotated in the range of [-3, +3] according to the intensity of emotion. The higher the value, the more positive the emotion, and the lower the value, the more negative the emotion. The MOSEI dataset is an extended version of the MOSI dataset, which was proposed by AmirAli Bagher Zadeh et al. in the 2018 56th Annual Meeting of the Association for Computational Linguistics paper "AmirAli BagherZadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, Louis-Philippe Morency, Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretabledynamic fusion graph, in: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Vol. 1, 2018, pp.2236–2246". MOSEI is currently the largest public multimodal sentiment analysis dataset, containing a total of 22,856 samples, each of which is also labeled in the range of [-3, +3] according to the intensity of the sentiment.
[0133] Table 1
[0134]
[0135] Step 2 includes:
[0136] In the process of processing the input video file, the feature extraction model uses Bert, Librosa, and OpenFace to extract modal features of text, audio, and video respectively, denoted as :
[0137] ,
[0138] ,
[0139] ,
[0140] Then, the features of each modal ( , For text, For audio, For video. ) to unify them to the same feature dimension, thus obtaining the aligned feature representation , , is the length of the sequence, is the length of the mode vector:
[0141] ,
[0142] in and are the weight matrix and bias matrix of the linear mapping.
[0143] Step 3 includes:
[0144] Each modal feature is input into the modality-private encoder to generate modality-private features , , , At the same time, they are input into the shared encoder to generate modality shared features , where the private encoder extracts features from each modality data through three independent encoders, while the shared encoder processes all modality data using the same encoder. The specific formula is:
[0145] ,
[0146] ,
[0147] in, , , They are text private encoders , audio private encoder , Video Private Encoder Parameters, Is a shared encoder Then the present invention will share the characteristics , , Perform full connection processing to obtain global shared features :
[0148] ,
[0149] in, and is the weight matrix, is the activation function.
[0150] Step 4 includes: During the disentanglement process, the modality discriminator processes the features extracted by the private encoder and the shared encoder to generate the probability of belonging to each modality. For example, the text private feature Input to the discriminator and get the same as text, audio, video (i.e. , , ) features. Modality Discriminator It is to classify and judge the feature representation of the input data and output the corresponding probability. Its definition formula is as follows:
[0151] ,
[0152] in, , is the weight matrix of the discriminator. yes The transposed matrix of is the bias vector, . is the input of the modality discriminator, and the output of the discriminator Represents and corresponds to three modes ( ) related probability. In order to enhance the discriminator's ability to distinguish similar categories, the cross entropy loss function based on arc cosine is used, which is specifically defined as:
[0153] ,
[0154] ,
[0155] ,
[0156] in, is the target modality of the input sample, is the weight vector of mode m, is the target mode The weight vector of is the input feature vector, The input eigenvectors and modes are calculated The angle of The input eigenvector and the target mode are calculated The angle of The cross entropy loss was calculated. is the scaling factor, is the marginal parameter. Then the weighted sum is taken for each label to calculate the overall adversarial loss:
[0157] ,
[0158] ,
[0159] in, is the batch size, For the modal The features of the target modality are calculated The cross entropy loss between . is the adversarial loss of shared features, is the adversarial loss of private features. The goal of shared features is opposite to that of the modality discriminator, that is, to generate consistent features, so the gradient of the shared encoder is reversed during back propagation. It remains fixed during forward propagation.
[0160] Step 5 includes:
[0161] The temporal smoothness loss based on KL divergence is introduced to measure the changing characteristics of audio and video modalities in continuous time, thereby reducing redundancy and noise. KL divergence is defined as:
[0162] ,
[0163] , are two probability distributions, and They are respectively The formula calculates the value of the event In distribution Probability and distribution under Taking audio features as an example, the KL divergence of three consecutive frames of audio features is calculated. The specific steps are as follows:
[0164] ,
[0165] ,
[0166] ,
[0167] in, and Respectively represent Frame audio modality private features and Frame audio modality shared features, yes The processed results are averaged. and Indicates the audio feature in The private and shared features of a frame are defined by The average value after the operation. Indicates the frame number. , The audio feature is The sliding average of the private and shared features of a frame, , , is the temporal smoothness loss of audio features;
[0168] For video features, the KL divergence of three consecutive frames of audio features is calculated. The formula is:
[0169] ,
[0170] ,
[0171] ,
[0172] in, and Respectively represent Frame video modality private features and Frame video modality shared features, and Indicates the video features in The private and shared features of a frame are defined by The average value after the operation. , The video feature is The sliding average of the private and shared features of a frame, is the temporal smoothness loss of the video feature. In order to enhance the coordination of audio and video modalities in continuous time, the present invention first evaluates the feature deviation of the two modalities in continuous time, and constructs a consistency loss function based on this. The specific formula is as follows:
[0173] ,
[0174] ,
[0175] ,
[0176] in, and Respectively represent the feature differences between consecutive frames of video and audio modalities, It is Frame and The average value of + , is the consistency loss.
[0177] Through the above steps, the temporal smoothness loss of audio and video features is calculated, and finally the following formula is obtained to express the overall temporal smoothness loss:
[0178] ,
[0179] Through temporal smoothness loss, the present invention can effectively reduce redundant information in audio and video features and ensure their consistency in the temporal dimension, thereby improving the performance of the model in multimodal tasks.
[0180] Step 6 includes:
[0181] The deep fusion of audio and video modalities is achieved through the text-guided mechanism, aiming to fully explore the semantic information contained in the text modality and use it as the core factor to guide the fusion of audio and video features, thereby enhancing the cross-modal synergy and semantic consistency. Figure 4 As shown, in the cross-modal interaction layer, the present invention uses text features of different scales and fusion features of audio and video to fuse layer by layer, and the obtained fusion features Further sent to the gate control mechanism and shared features to be processed.
[0182] Low-scale text information mainly focuses on the emotional features at the word or phrase level, medium-scale text information focuses more on the emotional expression at the sentence level, and high-scale text information focuses on the emotional trend in the global context. The present invention performs multi-level feature extraction on text data through the Transformer layer, generates text representations of different scales, and inputs these text features as guidance information into the text guidance module to further assist the initial fusion of audio and video modalities.
[0183] Step 6.1, low-scale fusion: The present invention first initializes a learnable parameter tensor , and then the low-scale text features ( ), audio features , Video Features and are input into the text guidance module. Figure 3 As shown in Figure 2, the core of the text-guided module lies in the cross-modal attention mechanism. In the cross-modal attention calculation, taking audio features as an example, text modality features is used as the query vector Q, the audio modality features As the key vector K and value vector V respectively, we get and The similarity matrix :
[0184] ,
[0185] in, and is the learnable parameter matrix, Represents the dimension of each attention head. Text features and video features The similarity matrix The calculation of The calculation is similar:
[0186] ,
[0187] After cross-modal attention calculation, a fused intermediate feature representation is generated, which is then combined with the fully connected features of audio and video. Further fusion, finally generating new fusion features :
[0188] ,
[0189] in, and are all learnable parameter matrices.
[0190] Step 6.2, mesoscale fusion: In order to further mine the mesoscale text information of the text modality, Input into the first Transformer layer, and generate mesoscale text features through deep modeling :
[0191] ,
[0192] in, represents the first Transformer layer, is the parameter of the first Transformer layer. and mesoscale text features is input into the next text-guided module. Similar to the previous text-guided module, through cross-modal attention and full connection, we get :
[0193] ,
[0194] in, , , , , and are all learnable parameter matrices.
[0195] Step 6.3, high-scale fusion: Then Input to the second Transformer layer to extract high-scale text features ,then and It is input into the last text-guided module to obtain the final fusion features of audio features and video features. :
[0196] ,
[0197] ,
[0198] in, represents the second Transformer layer, are the parameters of the second Transformer layer, , , is a learnable parameter matrix;
[0199] Step 6.4, cross-modal interaction: The features extracted from the text modality at low, medium, and high scales are used as query vectors Q, and the fusion features of audio and video are As the key vector K and the value vector V. Based on this, the fusion feature is calculated through the cross-modal attention mechanism, and the layer normalization is performed on the basis of the fusion feature, and then further modeled through the feedforward network (FFN), and the final fusion feature is obtained. The calculation is as follows:
[0200] ,
[0201] ,
[0202] ,
[0203] in, , , , is the learnable parameter matrix, is the output after calculation by the cross-modal self-attention mechanism, yes After normalization, the output Represent text features at different scales;
[0204] hour, , , is a low-scale text feature With fusion features The learnable parameter matrix in the cross-modal attention mechanism, yes and The output after calculation by the cross-modal self-attention mechanism, yes Output after normalization;
[0205] hour, , , is the mesoscale text feature With fusion features The learnable parameter matrix in the cross-modal attention mechanism, yes and The output after calculation by the cross-modal self-attention mechanism, yes Output after normalization;
[0206] hour, , , It is a high-scale text feature With fusion features The learnable parameter matrix in the cross-modal attention mechanism, yes and The output after calculation by the cross-modal self-attention mechanism, yes The output after normalization.
[0207] represents normalization, represents a feed-forward network; , , Represent the multimodal fusion features generated at low, medium and high scales respectively. The present invention introduces a gating mechanism to filter redundant information and dynamically screen the effective information between different modalities. Specifically, the fusion features at different scales are and shared modal features Weighted fusion via gating mechanism:
[0208] ,
[0209] ,
[0210] ,
[0211] ,
[0212] in, ,Right now , , , , and is a learnable weight matrix, Through the weight matrix The result after transformation, refers to sum pooling, is the size of the pooling window, yes After summing and pooling, express The result after L2 normalization is: express The second norm of . hour, yes The learnable weight matrix, yes Through the weight matrix The result after transformation, yes After summing and pooling, express The result after L2 normalization. The present invention firstly transforms the input Multiply it element by element with the weight matrix, and then sum and pool the matrix after element-by-element multiplication to get , and then L2 normalization is performed to ensure the same scale. Finally, a linear mapping weighted sum is performed to obtain the final fusion feature This feature integrates the information of multimodal features at different scales and significantly improves the effect of cross-modal fusion.
[0213] Step 7 includes:
[0214] The fused multimodal features As input, it is passed to the multi-layer perceptron and outputs the sentiment prediction result through nonlinear transformation. :
[0215] ,
[0216] In the process of model training, in order to make the prediction results as close as possible to the real emotional label, this paper uses mean square error (MSE) as the main loss function for task learning. Its calculation formula is as follows:
[0217] ,
[0218] in, is the sentiment label of the nth sample, is the prediction of the model of the present invention, is the number of samples in the dataset. , the model can better fit the sentiment label and improve the accuracy of sentiment prediction. So the final total loss function is:
[0219] ,
[0220] in, , and is a hyperparameter used to balance the weight between task loss and regularization loss. By optimizing the above-mentioned loss functions, the model can more fully explore the synergistic relationship and temporal consistency between multimodal features, thereby achieving accurate modeling of sentiment prediction tasks and further improving the generalization ability and practical application effect of the model.
[0221] Comparison Baseline: This example uses the following models as baselines to compare model performance. (1) TFN: A high-order tensor is generated by combining the features of text, audio, and video modalities through Cartesian products. Each dimension of this tensor captures the interaction information between unimodal, bimodal, and trimodal modalities. (2) LMF: Efficient multimodal fusion is performed by decomposing high-rank weight tensors into modality-specific low-rank factors. (3) MISA: Disentangled representation learning is used to separate modality-invariant common features and modality-specific unique features. Through this mechanism, MISA can simultaneously maintain the common information between modalities and the uniqueness of each modality. (4) FDMER: A modality discriminator is introduced, and adversarial learning is used to guide the parameter learning of public and private encoders. At the same time, the model designs a customized loss function to achieve modality consistency and parallax constraints. (5) PS-Mixer: A polar vector and an intensity vector are designed to judge the polarity and intensity of emotions respectively. Combined with a multi-layer perceptron (MLP) communication module consisting of multiple fully connected layers and activation functions. (6) TETFN: Obtain a unified and effective multimodal representation by learning text-oriented pairwise cross-modal mapping. (7) TCAN: Adopts a text query cross-attention mechanism to interact information between visual and acoustic modalities. (8) JTUM: Combining a unimodal label generation module with a cross-modal converter, it can generate an effective joint representation in multimodal data. (9) DEVA: Adopts a sentiment description generator to convert raw audio and visual data into textual sentiment descriptions. Subsequently, these sentiment descriptions are fused with the source data to enhance the sentiment information expressed by the features. (10) DLF: Introduces four geometric metrics to refine the disentanglement process, combined with a language-guided cross-attention mechanism, and uses supplementary specific modality information to enhance the effectiveness of language representation.
[0222] Table 2
[0223]
[0224] It can be seen from Table 2 and Table 3 that the method proposed in the present invention is substantially superior to the performance of all the compared baseline models on the two public data sets, thereby verifying the effectiveness of the present invention.
[0225] Table 3
[0226]
[0227] The present invention provides a multimodal sentiment analysis method based on disentangled representation learning. There are many methods and ways to implement the technical solution. The above is only a preferred implementation of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention. All components not specified in this embodiment can be implemented by existing technologies.
Claims
1. A multimodal sentiment analysis method based on disentangled representation learning, characterized in that: The following steps are involved: Step 1, data collection: collect data sets for multimodal sentiment analysis. The selected data sets should be representative and diverse, cover multiple modalities such as audio, video and text, and be able to support subsequent training, evaluation and testing processes; the multimodalities include text modality, audio modality and video modality; Step 2: Extract multimodal features: Extract text, audio and video modal features for the input video data respectively; use the Bert model to extract the text model feature U t , use Librosa tool to extract audio model features U a , use OpenFace to extract facial features and get the video model features U v , for the extracted modal feature U m , transformed to the same feature dimension through linear mapping, and the aligned modal features I are obtained m ; m∈{t,a,v}; Step 3: Construct a disentangled representation learning network, including: Generate modality-private features: transform modality feature I m Input to the modality private encoder to generate the corresponding modality private features P t , P a , P v ; Generate modal shared features: Modal feature I m Input the shared encoder at the same time to generate modality shared features C t , C a , C v ; For the modal shared feature C t , C a , C v Perform full connection processing to obtain the global shared feature F c ; Step 4: Modal disentanglement and adversarial optimization: transform the modal private feature P t , P a , P v and modality shared features C t , C a , C v Input modality discriminator, optimize the parameters of modality discriminator, private encoder and shared encoder; Step 5, time smoothness constraint: private feature P of the audio modality a , audio modality shared features C a and video modality private features P v , video modality shared features C v Applying temporal smoothness constraints to reduce redundancy and noise in audio and video features over consecutive time periods; Step 6, feature fusion: Deep fusion of audio and video modality features is achieved through the text-guided mechanism, and finally the fusion feature Y for sentiment analysis is generated; Step 6 includes: initializing a learnable parameter tensor S 1 , and then the low-scale text features Audio modality private feature P a , video modality private feature P v and S 1 The text-guided module includes a cross-modal attention mechanism for audio features and text-modal features. is used as the query vector Q, the audio modality private features P a As the key vector K and value vector V respectively, we get and P a The similarity matrix γ is: in, and is the learnable parameter matrix, d k Represents the dimension of each attention head; Calculating text features and video feature P v The similarity matrix β of in and is a learnable parameter matrix; Then generate a new fusion feature S 2 : Where S 2 , and is the learnable parameter matrix; Will Input into the first Transformer layer, and generate medium-scale text features through deep modeling Among them, E 1 represents the first Transformer layer, are the parameters of the first Transformer layer; Then the new fusion feature S 2 and mesoscale text features Input into the next text-guided module to obtain the fusion feature S 3 : in and is a learnable parameter matrix; Then the mesoscale text features Input to the second Transformer layer to extract high-scale text features then and S 3 It is input into the last text-guided module to obtain the final fusion feature S of audio features and video features. 4 : Among them, E 2 represents the second Transformer layer, are the parameters of the second Transformer layer, is a learnable parameter matrix; Further modeling through the feedforward network FFN, the final fusion features are obtained: where x∈{1,2,3}, is the learnable parameter matrix, is the output after calculation by the cross-modal self-attention mechanism, yes After normalization, the output Represent text features at different scales; When x=1, is a low-scale text feature And the fusion feature S 4 The learnable parameter matrix in the cross-modal attention mechanism, yes With S 4 The output after calculation by the cross-modal self-attention mechanism, yes Output after normalization; When x=2, is the mesoscale text feature And the fusion feature S 4 The learnable parameter matrix in the cross-modal attention mechanism, yes With S 4 The output after calculation by the cross-modal self-attention mechanism, yes Output after normalization; When x=3, It is a high-scale text feature And the fusion feature S 4 The learnable parameter matrix in the cross-modal attention mechanism, yes With S 4 The output after calculation by the cross-modal self-attention mechanism, yes Output after normalization; LayerNorm means normalization, FFN means feed-forward network; F 1 、F 2 、F 3 They represent the multimodal fusion features generated at low, medium, and high scales respectively; The fusion features F at different scales x and the shared modal features F c Weighted fusion via gating mechanism: G j =W j ·F j , Among them, j∈{1,2,3,c}, that is, F j ∈{F 1 ,F 2 ,F 3 ,F c },W j , and is a learnable weight matrix, G j Yes F j After the weight matrix W j The result after transformation, SumPool refers to sum pooling, k is the size of the pooling window, It's G j After summing and pooling, express The result after L2 normalization is: express The second norm of; when j = 1, W1 is F 1 The learnable weight matrix, G 1 Yes F 1 The result after transformation by weight matrix W1 is: It's G 1 After summing and pooling, express The result after L2 normalization; Enter F j Multiply it element by element with the weight matrix, and then sum and pool the matrix after element-by-element multiplication to get Then L2 normalization is performed to ensure the same scale. Finally, linear mapping weighted summation is performed to obtain the final fusion feature Y. Step 7, model optimization and output: Input the fusion feature Y into the sentiment analysis module, combine it with the sentiment classification or regression task, and generate the sentiment prediction result; at the same time, optimize the sentiment analysis module through the loss function of the training phase.
2. The method according to claim 1, characterized in that Step 2 includes: For the input video data, the Bert model is used to extract the text model feature U t , use Librosa tool to extract audio model features U a , use OpenFace to extract facial features and get the video model features U v : U t =BERT(T), The a =Book(A), U t =OpenFace(V), Denoted as feature set D = {U t ,U a ,U v }; For each modal feature U m Perform linear mapping to unify to the same feature dimension, thereby obtaining the aligned feature representation I m : I m =W m ·U m +b m , Where m represents the modality, m∈{t,a,v}, t is text, a is audio, v is video, W m and b m Represent the weight matrix and bias matrix of the linear mapping respectively.
3. The method according to claim 2, characterized in that Step 3 includes: Modal Characteristics I m are input into the modality-private encoder to generate text, audio, and video modality-private features P respectively. t , P a , P v , modal characteristics I m They are simultaneously input into the shared encoder to generate shared features C for text, audio, and video modalities respectively. t , C a , C v , where the private encoder extracts features from each modal data through three independent encoders, and the shared encoder processes all modal data using the same encoder. The mathematical expression is: P t =H t (I t ,θ t ),P a =H a (I a ,θ a ),P v =H v (I v ,θ v ), C m =H C (I m ,θ C ), Among them, θ t ,θ a ,θ v They are text private encoder H t Parameters of audio private encoder H a Parameters of the video private encoder H v The parameter θ C is the shared encoder H C Parameters; The shared feature C t , C a , C v Perform full connection processing to obtain the global shared feature F c : in, is the weight matrix and σ(·) is the activation function.
4. The method according to claim 3, characterized in that Step 4 includes: Modality discriminator D(h;θ D ) is to classify and judge the feature representation of the input data and output the corresponding probability. The formula is: Among them, W D , W F is the weight matrix of the modality discriminator, T represents the transpose, b F is the bias vector; is the input of the modality discriminator, d m is the size of the dimension, Indicates d m ×1-dimensional real space, the output D(h; θ D ) represents the probability associated with the corresponding three modes t, a, and v; The arccosine-based cross entropy loss function is used, which is defined as: Among them, y m is the target modality of the input sample, W m is the weight vector of mode m, is the target mode y m The weight vector of is the input feature vector, θ m represents the angle between the input eigenvector and mode m, Represents the input feature vector and the target mode y m The angle, e am represents the cross entropy loss; α is the scaling factor, τ is the margin parameter; e is a natural constant; Take the weighted sum for each label and calculate the overall adversarial loss: Where n is the batch size, l am (C m ,y m ) represents the modal shared feature C m With the target mode y m The cross entropy loss between Cam is the adversarial loss of shared features, l Pam is the adversarial loss of private features; l am (P m ,y m ) represents the modal private feature P m With the target mode y m The cross entropy loss between .
5. The method according to claim 4, characterized in that Step 5 includes: The temporal smoothness loss based on KL divergence measures the changing characteristics of audio and video modalities in continuous time. KL divergence is defined as: Where p and q are two probability distributions, p g and q g are the values of p and q at the gth position respectively; KL(p‖q) represents the event p g The difference between the probability under distribution p and the probability under distribution q; For audio features, the KL divergence of three consecutive frames of audio features is calculated. The formula is: Among them, P a [i] and C a [i] represents the private features of the audio modality of the i-th frame and the shared features of the audio modality of the i-th frame, respectively. Mean is the average operation of the result after softmax processing. Represents the average value of the private features of the audio feature in the i-th frame after the softmax operation; represents the average value of the shared features of the audio features in the i-th frame after the softmax operation; b represents the number of frames, is the sliding average of the private features of the audio feature at the i-th frame, is the sliding average of the shared features of the audio features in the i-th frame, e kla is the temporal smoothness loss of audio features; For the KL divergence of video features, the formula is: Among them, P v [i] and C v [i] represents the private features of the i-th frame video modality and the shared features of the i-th frame video modality, respectively. It represents the average value of the private features of the video feature in the i-th frame after the softmax operation. Represents the average value of the shared features of the video features in the i-th frame after the softmax operation; is the sliding average of the private features of the video feature at the i-th frame, is the sliding average of the shared features of the video features in the i-th frame, l klv is the temporal smoothness loss of video features; Construct a consistency loss function, the formula is: in, and They represent the feature differences between consecutive frames of the video modality and the feature differences between consecutive frames of the audio modality, ΔM i is the i-th frame and The average value of l klav is the consistency loss; The overall time smoothness loss l is calculated by the following formula tsc : l tsc =l kla +l klv +l klav 。 6. The method according to claim 5, characterized in that Step 7 includes: The fused feature Y is taken as input and passed to the multi-layer perceptron MLP, which outputs the sentiment prediction result through nonlinear transformation. The mean square error MSE is used as the loss function for task learning. task , the calculation formula is: Among them, y z is the sentiment label of the zth sample, is the prediction result of the sentiment analysis module, N t is the number of samples in the dataset.
7. The method according to claim 6, characterized in that In step 7, the final total loss function l all for: l all =l task +λl Cam +δl Pam +ψl tsc , Among them, λ, δ and ψ are hyperparameters.
8. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 7.
9. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 7 are executed.
Citation Information
Patent Citations
Multimodal aspect emotion joint extraction method integrated with entity knowledge
CN118504570A
Multi-modal sentiment analysis method based on multi-agent cooperation
CN118673406A