Sentiment Analysis Method and System Based on Multimodal Prefix and Cross-Modal Attention
Through multimodal prefix and cross-modal attention methods, combined with speech, text and image data for sentiment analysis, the problems of insufficient utilization of single-modal data and insufficient fusion of cross-modals are solved, and the training speed of the model and the accuracy of sentiment analysis are improved.
Patent Information
- Application Number
- CN202311621681.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-11-29
AI Technical Summary
In the prior art, only single-modal data is considered in the emotional analysis process, resulting in the data set being underutilized, the model training accuracy is not high enough, and the cross-modal data cannot be effectively fused, resulting in the emotional analysis results being inaccurate enough.
The multimodal prefix and cross-modal attention method are used to obtain three modal data: speech, text and image, and encode and prefix generation, respectively, splice into a modified text marker sequence, and use a pre-trained language model for emotional classification.
The training speed and generalization ability of the model are improved, the data imbalance problem is solved, the advantages of text modes are retained, and the interactions between multimodals are captured through cross-modal attention, which improves the accuracy of sentiment analysis.
Smart Images

Figure CN117609882B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sentiment analysis, and particularly to a sentiment analysis method and system based on multi-modal prefix and cross-modal attention. Background Art
[0002] The statements in this section merely mention the background art related to the present invention and do not necessarily constitute prior art.
[0003] Sentiment analysis aims to extract the potential attitudes and opinions towards an entity. Traditional sentiment analysis methods mainly target text. However, with the development of technology, people have started to express their opinions and emotions in the form of audio, images, and videos. Therefore, multi-modal sentiment analysis (MSA) models have been widely proposed. Compared with single-modal sentiment analysis, multi-modal sentiment analysis can achieve better prediction results because the complementary relationship between multiple modal data can analyze emotions more deeply.
[0004] In the process of implementing the present invention, the inventors found the following technical problems in the prior art:
[0005] In the prior art, only single-modal data is often considered in the process of sentiment analysis, and the data set cannot be fully utilized, resulting in low accuracy after model training;
[0006] In the prior art, cross-modal data fusion is not considered in the process of sentiment analysis, resulting in inaccurate sentiment analysis results. Summary of the Invention
[0007] To solve the deficiencies of the prior art, the present invention provides a sentiment analysis method and system based on multi-modal prefix and cross-modal attention;
[0008] On the one hand, a sentiment analysis method based on multi-modal prefix and cross-modal attention is provided, including:
[0009] Obtain a video segment to be sentiment analyzed; the video segment to be sentiment analyzed includes three types of modal data: speech, text, and image;
[0010] Input the video segment to be sentiment analyzed into a trained multi-modal sentiment analysis model, and output a multi-modal sentiment analysis result;
[0011] Among them, the trained multi-modal sentiment analysis model is used to: respectively encode the speech, text, and image in the video clip to obtain speech features, text token sequences, and visual features; respectively perform prefix generation on the speech features and visual features to obtain speech prefix tokens and visual prefix tokens; concatenate the speech prefix tokens and visual prefix tokens, and add the concatenation result to the text token sequence to obtain a corrected text token sequence; perform sentiment classification on the corrected text token sequence to obtain a sentiment classification label.
[0012] On the other hand, a sentiment analysis system based on multi-modal prefix and cross-modal attention is provided, including:
[0013] An acquisition module, which is configured to: acquire a video clip to be sentiment-analyzed; the video clip to be sentiment-analyzed includes three types of modal data: speech, text, and image;
[0014] An output module, which is configured to: input the video clip to be sentiment-analyzed into the trained multi-modal sentiment analysis model and output a multi-modal sentiment analysis result;
[0015] Among them, the trained multi-modal sentiment analysis model is used to: respectively encode the speech, text, and image in the video clip to obtain speech features, text token sequences, and visual features; respectively perform prefix generation on the speech features and visual features to obtain speech prefix tokens and visual prefix tokens; concatenate the speech prefix tokens and visual prefix tokens, and add the concatenation result to the text token sequence to obtain a corrected text token sequence; perform sentiment classification on the corrected text token sequence to obtain a sentiment classification label.
[0016] On yet another aspect, an electronic device is further provided, including:
[0017] A memory for non-temporarily storing computer-readable instructions; and
[0018] A processor for running the computer-readable instructions,
[0019] Among them, when the computer-readable instructions are run by the processor, the method described in the first aspect above is executed.
[0020] On yet another aspect, a storage medium is further provided, which non-temporarily stores computer-readable instructions, and when the non-temporary computer-readable instructions are executed by a computer, the instructions for executing the method described in the first aspect are executed.
[0021] On yet another aspect, a computer program product is further provided, including a computer program, and when the computer program runs on one or more processors, it is used to implement the method described in the first aspect above.
[0022] The above technical solution has the following advantages or beneficial effects:
[0023] 1. To address the problem of underutilization of the dataset, the present invention selects to use the original data for pre-training, which can fully utilize the data information, improve the training speed and generalization ability of the model, and solve the problem of data imbalance that may occur in the target task dataset.
[0024] 2. To address the problem of how to maintain the advantages of the text modality, the present invention introduces A_LAA and proposes V_LAA. These two modules can encode speech and visual features into prefix tokens. Feeding the prefix tokens into the pre-trained language model can retain the advantages of the large-scale language model and achieve good results in the MSA task.
[0025] 3. To address the problem of interaction between multiple modalities, the present invention can expand the function of the RoBERTa model by adding prefix tokens to RoBERTa, enabling it to have the ability to learn cross-modal attention, which is conducive to capturing the interaction between multiple modalities. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0027] Figure 1 is the overall architecture diagram of Embodiment 1;
[0028] Figure 2 is the audio classification head of Embodiment 1;
[0029] Figure 3 is the visual classification head of Embodiment 1;
[0030] Figure 4 is the representative example in the case study of Embodiment 1 and its predicted and true scores;
[0031] Figures 5(a) and 5(b) are the trend charts of the loss index and evaluation index during the training process of Embodiment 1;
[0032] Figure 6 is the schematic diagram of cross-modal attention visualization of Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0034] Term Explanation:
[0035] MSA (Multimodal Sentiment Analysis)
[0036] LSTM (long short - term memory)
[0037] BERT (Bidirectional Encoder Representation from Transformers)
[0038] RoBERTa (Robustly Optimized BERT)
[0039] A_LAA (Audio Lightweight Attentive Aggregation)
[0040] V_LAA (Visual Lightweight Attentive Aggregation)
[0041] In recent years, the research on multimodal sentiment analysis has mainly focused on the development of complex fusion methods. Traditional fusion methods usually adopt relatively simple model structures and operations. Chen et al. proposed gated multimodal embedded LSTM that can effectively filter noise. Wang et al. designed SAL - CNN to improve the generalization of neural networks by selecting an additive algorithm. Later, the attention mechanism was incorporated into the fusion methods. Zadeh A et al. proposed a multi - attention recurrent network to discover the interactions between modalities through multi - attention blocks. Zadeh A et al. used "Delta - memory attention" and "Multi - View Gated Memory" to simultaneously capture temporal and inter - modal interactions. There are also some fusion methods using tensors. Zadeh A et al. used the triple Cartesian product of modal embeddings to model unimodal, bimodal, and trimodal interactions. Ming Hou et al. proposed polynomial tensor pooling to integrate multimodal information through high - order matrices. Some methods optimize the fusion results by enhancing multimodal representations. Hazarika et al. learned more effective modal representations by partitioning into two sub - spaces of modality - invariant and modality - specific representations. Yu et al. proposed a label generation module based on self - supervised learning strategies to learn unimodal representations. Han et al. optimized multimodal representations by hierarchically maximizing the mutual information between unimodal input pairs and between multimodal fusion and unimodal inputs.
[0042] Influenced by Transformer, pre-trained language models have developed rapidly and been successfully extended to multi-modal related tasks. Radford et al. proposed to use text information to supervise the training of the CLIP model for the visual modality, which demonstrated powerful performance in image-text related tasks. Li et al. proposed the BLIP model by adding the Multimodal mixture of Encoder-Decoder structure and the Captioning and Filtering data cleaning method on the basis of ALBEF. Due to the excellent performance of fine-tuning large-scale pre-trained models based on Transformer in various downstream tasks, researchers have started to utilize this advantage in multi-modal sentiment analysis tasks. Rahman et al. designed the Multimodal adaptive Gate module to help BERT process multi-modal data. Arjmand et al. proposed a speech prefix language model based on Transformer, which processes the speech modality as dynamic prefix labels and trains them together with the text modality. Yu et al. proposed the first speech-text dialog pre-trained model and performed well in the MSA task.
[0043] Previous studies have also shown that in most cases, the text modality is superior to other modalities because it provides more explicit semantics and context relationships. Compared with the sound and visual modalities, the text modality also has a more mature language model pre-trained on large-scale corpora. Therefore, inspired by Rahman et al., the present invention converts both the speech and visual modalities into tokens and adds them to the discrete sequence of the pre-trained language model, while ensuring the dominant position of the text, and utilizes the pre-trained language model to learn cross-modal interactions. Moreover, most pre-training methods often appear in pairs such as speech and text, vision and text, and in the MSA method, there are few methods using Extra Training Data for pre-training. Therefore, the present invention attempts to use an external data set containing three modalities for pre-training and then fine-tune it to downstream tasks. For the selection of pre-training data, since manual annotation often consumes a large amount of cost and may lead to the loss of key information of the modality, the present invention uses multi-modal raw data as pre-training data.
[0044] The present invention proposes a multi-modal language model for multi-modal sentiment analysis, and its core function is to use the pre-trained language model to learn cross-modal attention. Specifically, the present invention uses a lightweight encoding module to encode the audio and visual modes into prefix tokens, and then provides them to the pre-trained language model based on Transformer. Finally, the present invention fine-tunes the pre-trained model to adapt it to the task of multi-modal sentiment analysis.
[0045] Embodiment 1
[0046] This embodiment provides an emotion analysis method based on multi-modal prefix and cross-modal attention;
[0047] An emotion analysis method based on multi-modal prefix and cross-modal attention includes:
[0048] S101: Obtain a video clip to be subjected to emotion analysis; the video clip to be subjected to emotion analysis includes three types of modal data: speech, text, and image;
[0049] S102: Input the video clip to be subjected to emotion analysis into a trained multi-modal emotion analysis model, and output a multi-modal emotion analysis result;
[0050] Among them, the trained multi-modal emotion analysis model is used for:
[0051] Encode the speech, text, and image in the video clip respectively to obtain speech features, text token sequences, and visual features;
[0052] Perform prefix generation on the speech features and visual features respectively to obtain speech prefix tokens and visual prefix tokens;
[0053] Concatenate the speech prefix tokens and visual prefix tokens, and add the concatenation result to the text token sequence to obtain a corrected text token sequence;
[0054] Perform emotion classification on the corrected text token sequence to obtain an emotion classification label.
[0055] Furthermore, the network structure of the multi-modal emotion analysis model includes:
[0056] A speech feature encoding module, a text feature encoding module, and a visual feature encoding module;
[0057] The output end of the speech feature encoding module is connected to the input end of the audio prefix generation module;
[0058] The output end of the visual feature encoding module is connected to the input end of the visual prefix generation module;
[0059] The output ends of the audio prefix generation module, the visual prefix generation module, and the text feature encoding module are all connected to the input end of the concatenation module;
[0060] The output end of the concatenation module is connected to the input end of the pre-trained language model, and the output end of the pre-trained language model outputs an emotion classification prediction result.
[0061] Furthermore, the speech feature encoding module is implemented by using a Wav2Vec 2.0 audio encoder.
[0062] Exemplarily, for the speech feature encoder module, Wav2Vec 2.0 is selected as the audio encoder, and X t is input into the audio encoder to obtain the sequence representation h a .
[0063]
[0064] Wherein, represents the pre-trained parameters from Wav2Vec 2.0, represents the speech latent feature vector, and i represents the time step of the audio modality.
[0065] Furthermore, the text feature encoding module is implemented by using the RoBERTa text encoder.
[0066] Exemplarily, for the text feature encoding module: due to the successful application of the pre-trained language model in downstream tasks, 12-layer RoBERTa is used as the text encoder, and then the text input X t is input into the tokenizer of RoBERTa to decompose the sentence into a more discrete sequence of tokens H t .
[0067] H t = tokenizer(X t ) = {[CLS], t1, t2,..., t i , [SEP]}
[0068] Wherein, represents the byte-level token, and i represents the time step in the text modality. [CLS] and [SEP] represent the start and end of the sequence respectively.
[0069] It should be understood that tokenizer represents the tokenizer, which is a tool for splitting text into words or bytes, and it can convert text into a form that is easier for the model to understand.
[0070] Furthermore, the visual feature encoding module is implemented by using the Vision Transformer image encoder.
[0071] Exemplarily, for the visual feature encoding module, Vision Transformer is used as the encoder for images. VIT can learn the internal representation of images and extract the required features for downstream tasks. It converts the input image into a series of image patches and performs embeddings, and then sends them into the multi-layer Transformers encoder for processing. In the present invention, X v is input into the visual encoder to obtain the sequence representation h v .
[0072]
[0073] Among them, represents the pre-trained parameters from VIT.
[0074] Furthermore, the internal structures of the audio prefix generation module and the visual prefix generation module are the same. The audio prefix generation module includes:
[0075] An input end, a BiGRU module, an attention mechanism module, and a layer normalization module connected in sequence;
[0076] The attention mechanism module includes: a first linear layer, a first activation function layer, a second linear layer, and a second activation function layer connected in sequence;
[0077] The input end of the first linear layer is connected to the output end of the BiGRU module;
[0078] The output end of the second activation function layer is connected to the input end of the layer normalization module.
[0079] It should be understood that the linear layer refers to the Linear Layer, which can also be called the fully connected layer. Its function is to map the input features to the output space (obtain a new dimension) by adjusting the weights and biases.
[0080] Furthermore, the working processes of the audio prefix generation module and the visual prefix generation module are the same. Among them, the working process of the audio prefix generation module includes:
[0081] Extract context information from the audio features to obtain audio context features;
[0082] Aggregate the audio context features and the audio features through the attention mechanism to obtain the final audio prefix.
[0083] Furthermore, the aggregation of the audio context features and the audio features through the attention mechanism to obtain the final audio prefix specifically includes:
[0084] First, the feature representation learned by BiGRU undergoes a linear transformation and then the Sigmoid function is applied to obtain an intermediate representation of the attention weights, and the intermediate representation of the attention weights can affect the information focused by the model;
[0085] Then, a second linear transformation is performed and the Softmax function is applied to obtain the final attention weight coefficients;
[0086] Finally, according to the final attention weight coefficients, weighted summation is performed on the audio context features and the audio features, and finally concatenation is performed to obtain the final audio prefix Ca 。
[0087] Furthermore, the working processes of the audio prefix generation module and the visual prefix generation module are the same. Among them, the working process of the audio prefix generation module includes:
[0088]
[0089]
[0090]
[0091]
[0092]
[0093]
[0094] where m ∈ {a, v}. and represent the forward GRU and backward GRU calculations, T m represents the time step. δ represents the sigmoid activation function, W represents the weight matrix, b represents the bias, represents the concatenation operation.
[0095] It should be understood that the core of the model of the present invention is to combine the audio and visual modalities to learn cross-modal interactions on the basis of retaining the main position of the text modality. Although the audio and visual have been encoded by the encoder before, in order to better meet the understanding requirements of RoBERTa, the present invention adds A-LAA and V-LAA at the tails of the two encoders respectively as small encoders for secondary encoding, aiming to convert h a and h v into the prefix tokens C a and C v 。
[0096] The present invention extends the visual encoder V-LAA on the basis of the speech context encoder A-LAA, and the structure is as Figure 1 shown. After extracting the audio and visual features, they are respectively fed into the BiGRU, so that richer context information can be captured and more compact feature representations can be output. Then, these feature representations are aggregated through the attention mechanism to generate the final output.
[0097] The attention mechanism is mainly implemented through linear layers and activation functions. Specifically, using the attention mechanism, the module can dynamically focus on information in different parts of the input sequence during the aggregation stage, which can help the model better understand and process the input data.
[0098] Furthermore, the pre-trained language model is implemented using 12-layer RoBERTa.
[0099] Furthermore, the training process of the pre-trained language model includes:
[0100] Construct a first training set, where the first training set is a corrected text token sequence with known sentiment classification labels; the corrected text token sequence is obtained by encoding the speech, text, and images in the video clip respectively to obtain speech features, text token sequences, and visual features; perform prefix generation on the speech features and visual features respectively to obtain speech prefix tokens and visual prefix tokens; concatenate the speech prefix tokens and visual prefix tokens, and add the concatenation result to the text token sequence to obtain the corrected text token sequence;
[0101] Input the first training set into the language model and train the language model. Stop training when the loss function value of the language model no longer decreases to obtain the trained language model.
[0102] Furthermore, the loss function L of the language model mlm , is specifically expressed as:
[0103]
[0104] where y mask represents the true label probability distribution, represents the probability distribution predicted by the model, and H represents the cross-entropy function.
[0105] Exemplarily, the training process of the pre-trained language model includes:
[0106] Adjust the sequence input to the pre-trained RoBERTa to:
[0107] {[CLS],C a ,C v ,t1,t2,...,[MASK],...,t i ,[SEP]}
[0108] where [MASK] is a special token for the masked language model task (MLM). In the MLM task, some words in the input sequence are randomly masked. In the present invention, MLM aims to predict these [MASK] words using the unmasked words and the prefix labels of the audio and visual modalities.
[0109] The MaskedLMOutput is the output object of MLM. In the present invention, the loss attribute of MaskedLMOutput is used as the optimization target, and the cross-entropy loss is used as the loss function. The module parameters are adjusted by minimizing the loss value, thereby optimizing the prediction ability of the model in downstream tasks.
[0110] Furthermore, for the trained multi-modal sentiment analysis model, its training process includes:
[0111] Construct a second training set, which is a video file with known sentiment classification labels;
[0112] Input the second training set into the multi-modal sentiment analysis model to train the multi-modal sentiment analysis model. When the loss function value of the multi-modal sentiment analysis model no longer decreases, stop the training to obtain the trained multi-modal sentiment analysis model;
[0113] During the process of training the multi-modal sentiment analysis model, keep the parameters of the pre-trained language model unchanged, and the parameters of other modules except the pre-trained language model change.
[0114] Furthermore, the loss function L of the multi-modal sentiment analysis model MSE , is specifically expressed as:
[0115]
[0116]
[0117] where n represents the number of samples, θ H represents the parameters of the classification head, represents the predicted value of the i-th sample, y represents the true label of the i-th sample, and H last represents the hidden state of the last layer of the language model.
[0118] It should be understood that the present invention is fine-tuned on the multi-modal sentiment analysis task, and only the sub-models except the language encoder are fine-tuned, while keeping the weights of the language encoder unchanged. The advantage of doing this is that it can maintain the training representation ability of the language model, focus on optimizing specific modules, and make the model more effectively adapt to downstream tasks.
[0119] To obtain the final prediction, the present invention uses the hidden state of the last layer of the language model, denoted as H last , and then passes it through the classification head RobertaClassificationHead of the RoBERTa model. The role of the classification head is to map the hidden state of RoBERTa to a specific output space, and then obtain the prediction result 。The present invention selects MSE Loss as the loss function, calculates the gradient of the loss function with respect to the model parameters through backpropagation, and finally uses the gradient descent algorithm to decrease the value of the loss function. The smaller the value of MSE Loss, the smaller the difference between the predicted value and the true label, and the more accurate the prediction of the model.
[0120] The overall structure of the model of the present invention is as Figure 1 shown. The main process can be divided into three parts: data preparation, modality encoding, and model training.
[0121] First, the present invention performs data transformation on the original data of the pre-training dataset so that they can adapt to the processing requirements of the model.
[0122] In the subsequent encoding stage, the present invention uses Wav2vec2.0 and Vision Transformer (VIT) to extract the features of the audio and visual modalities respectively, and then uses A-LAA (Audio Lightweight Attentive Aggregation) and V-LAA (Visual Lightweight Attentive Aggregation) to process the two modality features into a prefix form suitable for the encoding of the pre-training language model, and forms a new sequence with the text tokens for training.
[0123] The objective of the present invention is to conduct research and analysis on the multi-modal sentiment analysis task. Therefore, in the final fine-tuning part, the present invention transfers the pre-trained model to the multi-modal sentiment analysis dataset CMU-MOSI for testing.
[0124] Given a video clip, which includes three modalities: text, audio, and visual, the work of the present invention is to analyze the emotional intensity in the video through these three modalities. The input of the model is X m . Where m ∈ {t, v, a}, and t, v, a represent the three modalities of text, visual, and audio respectively. The output of the model is a result reflecting the emotional intensity , and is used for the final result prediction.
[0125] Before specifically introducing the model, pre-training data is prepared for the model. The dataset selected by the present invention is the raw data of CMU-MOSEI, which consists of text sentences, video clips, and audio clips. However, the publicly available original CMU-MOSEI dataset also has the following problems: 1) sample missing; 2) misalignment in some samples. To solve these two problems, the present invention decides to preprocess the data.
[0126] For the audio and text modalities, the present invention uses the original audio data and text. For the visual modality, the present invention extracts frames from video clips to obtain image data. Since a single picture is not representative of a video, the present invention intercepts multiple pictures for each video and then takes the mean of their feature representations as the feature representation of the video clip. This not only makes up for the lack of samples but also optimizes the feature representation of the visual modality.
[0127] Currently, the two most widely used multi-modal sentiment analysis datasets are CMU-MOSEI and CMU-MOSI respectively. Compared with CMU-MOSI, CMU-MOSEI is larger in scale and has a more diverse range of topics. Therefore, the present invention selects CMU-MOSEI as the pre-training dataset and CMU-MOSI as the fine-tuning dataset. CMU-MOSEI contains 5,000 monologue videos from 1,000 different speakers on a video website and 250 topics, which are divided into 22,856 annotated video clips, and the sentiment scores are labeled from -3 to +3, where scores below zero indicate the intensity of negative and scores above zero indicate the intensity of positive. The present invention uses the original data of CMU-MOSEI and the corresponding text content and performs data cleaning on it.
[0128] The CMU-MOSI dataset collects 93 monologue videos mainly focused on movie reviews from a video website. These videos are from 89 different speakers, and each speaker expressed their views on a certain movie or topic in the video. The length of each video varies from 2 to 5 minutes, and the videos are divided into 2,199 segments in total. Each segment is manually annotated, and the sentiment scores are labeled from -3 to +3.
[0129] The specific division of the dataset and the proportion of each part are shown in Table 1.
[0130] Table 1: Dataset division.
[0131]
[0132] The present invention compares the model with competitive baselines in the MSA task. Considering the composition of the model, the present invention finally selects the method based on the Transformer language model and the currently relatively advanced method as the baselines.
[0133] TFN: Tensor Fusion Network uses the tensor fusion method to model the relationships between modalities and uses the triple Cartesian product to simulate the interactions of single-modal, bi-modal, and tri-modal.
[0134] LMF: Low-rank Multimodal Fusion performs effective multimodal fusion by decomposing a high-order weight tensor into modality-specific low-rank factors.
[0135] ICCN: Interaction Canonical Correlation Network generates multimodal embedding features by learning the correlations between multimodals through deep canonical correlation analysis.
[0136] MISA: Modality-Invariant and-Specific Representations learns the commonalities between multimodals and the characteristics of individual modalities by partitioning into two subspaces, and better learns modality representations to assist fusion.
[0137] Self-MM: Self-Supervised Multi-task Multimodal sentiment analysis network trains unimodal tasks through a label generation module based on self-supervised learning to obtain independent unimodal supervision.
[0138] MAG-BERT: Multimodal Adaptation Gate BERT is an improvement on RAVEN, which can be used for BERT and XLNet models to fine-tune using multimodal information.
[0139] MMIM: MultiModal InfoMax hierarchically maximizes the mutual information between cross-modalities and between unimodal inputs and multimodal fusion results, better preserving the key information of modalities in the fusion result.
[0140] TEASEL: Transformer-Based Speech-Prefixed Language Model combines speech information into the pre-trained language model by converting it into tokens, achieving better results without training the entire Transformer.
[0141] EMT-DLFR: Efficient Multimodal Transformer with Dual-Level Feature Restoration uses EMT and DLFR to achieve more efficient multimodal fusion and increases the robustness of the model.
[0142] UniMSE: Unified Multimodal Sentiment Analysis and Emotion Recognition fully explores the complementary relationship between sentiment and emotion, and unifies the MSA and ERC tasks.
[0143] In the experiments of the present invention, the hardware environment for model operation is a single A100-SXM4-80GB GPU, and the software environment is Ubuntu, CUDA, Pytorch 1.10.1. The present invention also sets fixed random seeds in the code to ensure the reproducibility of the model. Table 2 presents the introduction of specific parameter settings.
[0144] Since multimodal sentiment analysis involves regression and classification tasks, the commonly used evaluation metrics usually include Pearson correlation coefficient (Corr), mean absolute error (MAE), Accuracy, and F1-Score.
[0145] Specifically, Corr measures the degree of linear relationship between the predicted value and the true value, and its range is between -1 and 1. MAE is used to quantify the average absolute error between the predicted value and the true value. Accuracy is used to evaluate the accuracy of classification, while F1-Score provides a comprehensive measure of precision and recall in binary classification tasks. To comprehensively evaluate the model, the present invention not only adopts binary accuracy (Acc-2), but also introduces multi-level accuracy metrics (Acc-5, Acc-7) as evaluation criteria. The calculation processes of the above metrics are as follows:
[0146] Table 2: Hyperparameter settings in each dataset.
[0147]
[0148] Table 3 shows the performance comparison between the model of the present invention and the baseline models, measured by the evaluation metrics introduced in the implementation details section. Due to the differences in evaluation settings among the baseline papers, the results presented in the present invention mainly rely on the state-of-the-art papers (with comparability). It should be noted that since TEASEL does not provide official code, the present invention compares the results provided by the third-party re-implemented code. To more comprehensively evaluate the model of the present invention, the present invention also compares Acc-5. However, since most papers lack Acc-5 results, the present invention includes the EMT-DLFR model that evaluates Acc-5.
[0149] From the results in Table 3, the present invention can observe that the model of the present invention is superior to the existing state-of-the-art model (UniMSE) in most metrics. Specifically, it is 0.16% / 1.15% higher than the best-performing UniMSE in Acc2, 0.047 in MAE, and 0.031 in Corr, setting a new record in Acc-5 and achieving the second-best result in Acc-7. For F1-Score, the present invention did not achieve good results. The present invention speculates that this may be due to class imbalance in the dataset. Compared with other metrics, F1-Score is more sensitive to the balance of positive and negative instances.
[0150] In summary, the above results can prove the effectiveness of the method proposed by the present invention in multi-modal sentiment analysis.
[0151] Table 3: Experimental results of different models on the CMU-MOSI dataset.
[0152]
[0153] (B): Linguistic features are based on BERT. (RB): Linguistic features are based on RoBERTa. ◇: Results from UniMSE. *: Results from EMT-DLFR. For the missing results of Acc-7 and Acc-5 in ◇, the data of EMT-DLFR is used. R: Reproduction results of TejasMahajan for TEASEL. (↑) / (↓): Higher / lower values indicate better model evaluation results. For Acc-2 and F1-score, the left and right sides of / represent the evaluation results of the negative / non-negative and negative / positive groups respectively. The bolded part represents the best result, and the underlined part represents the second-best result.
[0154] The present invention conducted a series of ablation experiments on CMU-MOSI, and the experimental results are shown in Table 4.
[0155]
[0156] The present invention verified the impact of modalities on model performance by removing one or two modalities. First, the present invention found that the multi-modal combination provided the best performance, and removing the visual or audio modality would reduce the performance of the model. This shows the necessity of visual and audio modalities in multi-modal sentiment analysis and also proves that the model of the present invention can learn the interaction features between modalities. Next, the present invention observed that when both the visual and audio modalities were removed simultaneously, the average performance of the model was the worst. This shows that for the model of the present invention, three modalities are better than two modalities, and two modalities are better than a single modality.
[0157] To further observe the performance of the model and existing problems, the present invention analyzes several typical cases in the CMU-MOSI dataset. As Figure 3 shown, it includes predicted values, true values, and corresponding input data (for visual and audio modalities, the present invention also adds text descriptions). Figure 2 : Representative examples in the case study, their predictions, and true scores. The example data is taken from the CMU-MOSI dataset.
[0158] In case (A), the true value has a relatively positive sentiment, and the model provides a prediction very close to the truth. The present invention analyzes that this may be due to the dominance of the phrase "pretty good" in the text, supplemented by a nodding gesture visually and a rising intonation audibly.
[0159] In case (B), the true value is a neutral sentiment, while the prediction shows a slightly negative sentiment. The present invention analyzes that this may be affected by the word "know" in the text. "Know" indicates that the speaker has realized or understood something, and when combined with a glance gesture visually, it may lead to a negative prediction.
[0160] In case (C), the true value shows a slightly negative sentiment, but the prediction shows a more negative sentiment. The present invention analyzes that this may be due to the appearance of the word "not" in the text, combined with frowning and head-shaking gestures visually, resulting in more negative predictions. Therefore, the present invention concludes that audio and visual modalities do affect the predictions of multimodal sentiment analysis, but appropriate noise reduction may be required to prevent the introduction of redundant information.
[0161] To gain a deeper understanding of how the model of the present invention operates. The present invention visualizes the loss changes, evaluation metric changes, and cross-modal attention. As Figure 3 、 Figure 4 shown.
[0162] Loss can measure the performance of the model and the quality of model training. Therefore, the present invention tracks the loss during the training process of the training, validation, and test sets. As Figure 4 shown, as the number of steps increases, all three losses decrease. This indicates that the model of the present invention does learn as expected.
[0163] Figure 4 It also shows the changing trends of evaluation metrics during training. The present invention observes that Acc-2, F1, and Corr all increase as the time steps increase, indicating that the performance of the model gradually improves during training, and the linear relationship between the prediction and the true value becomes stronger. In addition, since a lower MAE indicates that the prediction is closer to the true value, the trend of MAE is also consistent with the ideal training state.
[0164] To verify whether the model learns the prefix information of audio and vision during the fine-tuning process, the present invention randomly selects the sample "2iD-tVS8NPw_28" from CMU-MOSI and performs cross-modal attention visualization, as shown in Figures 5(a) and 5(b). The speaker said, "So it's kind of interesting to see where they'll go" in a vivid tone and expression, and got a positive score of (+1.6000). It can be observed from Figures 5(a) and 5(b) that different word tokens simultaneously focus on acoustic and visual prefixes, indicating that the learning of the model is effective. Figure 6 Schematic diagram of cross-modal attention visualization for Embodiment 1.
[0165] Embodiment 2
[0166] This embodiment provides an emotion analysis system based on multi-modal prefixes and cross-modal attention;
[0167] The emotion analysis system based on multi-modal prefixes and cross-modal attention includes:
[0168] An acquisition module, which is configured to: acquire a video segment to be subjected to emotion analysis; the video segment to be subjected to emotion analysis includes three types of modal data: speech, text, and image;
[0169] An output module, which is configured to: input the video segment to be subjected to emotion analysis into a trained multi-modal emotion analysis model and output a multi-modal emotion analysis result;
[0170] Among them, the trained multi-modal emotion analysis model is used to: respectively encode the speech, text, and image in the video segment to obtain speech features, a text token sequence, and visual features; respectively generate prefixes for the speech features and visual features to obtain speech prefix tokens and visual prefix tokens; concatenate the speech prefix tokens and visual prefix tokens, and add the concatenation result to the text token sequence to obtain a corrected text token sequence; perform emotion classification on the corrected text token sequence to obtain an emotion classification label.
[0171] It should be noted here that the above acquisition module and output module correspond to steps S101 to S102 in Embodiment 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.
[0172] In the above embodiments, the descriptions of each embodiment have their own focuses. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0173] The proposed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the above modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0174] Embodiment III
[0175] This embodiment also provides an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the above one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes the method described in Embodiment I above.
[0176] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0177] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random memory. For example, the memory may also store information about the device type.
[0178] In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software.
[0179] The method in Embodiment I can be directly embodied as being executed by the hardware processor, or executed by the combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0180] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0181] Embodiment Four
[0182] This embodiment also provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in Embodiment One is completed.
[0183] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A sentiment analysis method based on multi-modal prefix and cross-modal attention, characterized in that, Including: Obtaining a video clip to be sentiment - analyzed; The video clip to be sentiment - analyzed includes three - modality data: speech, text, and image; Inputting the video clip to be sentiment - analyzed into a trained multi - modality sentiment analysis model to output a multi - modality sentiment analysis result. The network structure of the multi - modality sentiment analysis model includes: A speech feature encoding module, a text feature encoding module, and a visual feature encoding module; The output end of the speech feature encoding module is connected to the input end of the audio prefix generation module; The output end of the visual feature encoding module is connected to the input end of the visual prefix generation module; The output ends of the audio prefix generation module, the visual prefix generation module, and the text feature encoding module are all connected to the input end of the concatenation module; The output end of the concatenation module is connected to the input end of the pre - trained language model, and the output end of the pre - trained language model outputs a sentiment classification prediction result; Among them, the trained multi - modality sentiment analysis model is used for: encoding the speech, text, and image in the video clip respectively to obtain speech features, text token sequences, and visual features; generating prefixes for the speech features and visual features respectively to obtain speech prefix tokens and visual prefix tokens; concatenating the speech prefix tokens and visual prefix tokens in series and adding the concatenation result to the text token sequence to obtain a corrected text token sequence; performing sentiment classification on the corrected text token sequence to obtain a sentiment classification label. The internal structures of the audio prefix generation module and the visual prefix generation module are the same. The audio prefix generation module includes: An input end, a BiGRU module, an attention mechanism module, and a layer normalization module connected in sequence; The attention mechanism module includes: a first linear layer, a first activation function layer, a second linear layer, and a second activation function layer connected in sequence; The input end of the first linear layer is connected to the output end of the BiGRU module; The output end of the second activation function layer is connected to the input end of the layer normalization module; The working processes of the audio prefix generation module and the visual prefix generation module are the same. Among them, the working process of the audio prefix generation module includes: where m ∈ {a, v}, and denote the forward GRU and backward GRU calculations, T m denotes the time step, δ denotes the sigmoid activation function, W denotes the weight matrix, b denotes the bias, denotes the concatenation operation.
2. The sentiment analysis method based on multi-modal prefix and cross-modal attention according to claim 1, characterized in that, The working processes of the audio prefix generation module and the visual prefix generation module are the same. Among them, the working process of the audio prefix generation module includes: Extracting context information from audio features to obtain audio context features; Aggregating the audio context features and audio features through an attention mechanism to obtain the final audio prefix; Among them, the aggregating the audio context features and audio features through an attention mechanism to obtain the final audio prefix specifically includes: First, the feature representation learned by BiGRU undergoes a linear transformation and then applies the Sigmoid function to obtain an intermediate representation of the attention weight, and the intermediate representation of the attention weight can affect the information focused by the model; Then, perform a second linear transformation and apply the Softmax function to obtain the final attention weight coefficient; Finally, according to the final attention weight coefficients, the audio context features and the audio features are weighted and summed, and finally concatenated to obtain the final audio prefix C a .
3. The sentiment analysis method based on multi-modal prefix and cross-modal attention according to claim 1, characterized in that The training process of the pre - trained language model includes: Construct a first training set, where the first training set is a corrected text token sequence with known sentiment classification labels; the corrected text token sequence is obtained by encoding the speech, text, and images in the video clip respectively to obtain speech features, text token sequences, and visual features; prefix generation is performed on the speech features and visual features respectively to obtain speech prefix tokens and visual prefix tokens; the speech prefix tokens and visual prefix tokens are concatenated, and the concatenation result is added to the text token sequence to obtain the corrected text token sequence; Input the first training set into the language model and train the language model. Stop training when the loss function value of the language model no longer decreases to obtain the trained language model; The loss function L of the language model mlm , which is specifically expressed as: Among them, y mask represents the true label probability distribution, represents the probability distribution predicted by the model, and H represents the cross-entropy function.
4. The sentiment analysis method based on multi-modal prefix and cross-modal attention as described in claim 1, wherein The training process of the trained multi-modal sentiment analysis model includes: Construct a second training set, where the second training set is a video file with known sentiment classification labels; Input the second training set into the multi-modal sentiment analysis model and train the multi-modal sentiment analysis model. Stop training when the loss function value of the multi-modal sentiment analysis model no longer decreases to obtain the trained multi-modal sentiment analysis model; During the training process of the multi-modal sentiment analysis model, keep the parameters of the pre-trained language model unchanged, and the parameters of other modules except the pre-trained language model change; The loss function L of the multi-modal sentiment analysis model MSE , which is specifically expressed as: Among them, n represents the number of samples, and θ H represents the parameters of the classification head, represents the predicted value of the i-th sample, y represents the true label of the i-th sample, and H last represents the hidden state of the last layer of the language model.
5. A sentiment analysis system based on multi-modal prefix and cross-modal attention, characterized in that, It includes: An acquisition module, which is configured to: acquire a video clip to be sentiment analyzed; The video clip to be sentiment analyzed includes three types of modal data: speech, text, and images; An output module, which is configured to: input the video clip to be sentiment analyzed into the trained multi-modal sentiment analysis model and output the multi-modal sentiment analysis result. The network structure of the multi-modal sentiment analysis model includes: A speech feature encoding module, a text feature encoding module, and a visual feature encoding module; The output end of the speech feature encoding module is connected to the input end of the audio prefix generation module; The output end of the visual feature encoding module is connected to the input end of the visual prefix generation module; The output ends of the audio prefix generation module, the visual prefix generation module, and the text feature encoding module are all connected to the input end of the concatenation module; The output end of the concatenation module is connected to the input end of the pre-trained language model, and the output end of the pre-trained language model outputs the sentiment classification prediction result; Among them, the trained multi-modal sentiment analysis model is used to: encode the speech, text, and images in the video clip respectively to obtain speech features, text token sequences, and visual features; perform prefix generation on the speech features and visual features respectively to obtain speech prefix tokens and visual prefix tokens; concatenate the speech prefix tokens and visual prefix tokens, and add the concatenation result to the text token sequence to obtain the corrected text token sequence; perform sentiment classification on the corrected text token sequence to obtain sentiment classification labels. The internal structures of the audio prefix generation module and the visual prefix generation module are the same. The audio prefix generation module includes: An input end, a BiGRU module, an attention mechanism module, and a layer normalization module connected in sequence; The attention mechanism module includes: a first linear layer, a first activation function layer, a second linear layer, and a second activation function layer connected in sequence; The input end of the first linear layer is connected to the output end of the BiGRU module; The output end of the second activation function layer is connected to the input end of the layer normalization module; The working processes of the audio prefix generation module and the visual prefix generation module are the same. Among them, the working process of the audio prefix generation module includes: where m ∈ {a, v}, and represent the forward GRU and backward GRU calculations, T m represents the time step, δ represents the sigmoid activation function, W represents the weight matrix, b represents the bias, represents the concatenation operation.
6. An electronic device, characterized in that it includes: A memory for non-temporarily storing computer-readable instructions; And A processor for running the computer-readable instructions, Wherein, when the computer-readable instructions are run by the processor, the method described in any one of claims 1-4 above is executed.
7. A storage medium, characterized in that it is non-transitory Store computer-readable instructions, wherein when the non-temporary computer-readable instructions are executed by a computer, the instructions for executing the method described in any one of claims 1-4 are executed.