Multi-modal sentiment analysis method based on pre-training and text modal guidance
By adopting pre-trained multimodal emotion analysis method in emotion analysis, combined with the characteristics of video, audio and text modalities, the problem that emotion analysis based on text modality cannot accurately identify user emotions is solved, and more accurate emotional tendency recognition and analysis effects are achieved.
Patent Information
- Application Number
- CN202510700248.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Sentiment analysis based on text modality cannot accurately identify user emotions, especially in cases where emotions conflict between single modalities.
The pre-training multimodal sentiment analysis method is used to obtain the initial features of video, audio and text modalities and distinguish them into similar features and different features using a projector. Then, the data sampler is used to calculate the similarity between samples, build positive and negative feature pairs, and update the parameters of the feature extraction module and projector by backpropagating the comparison loss and single-modal prediction errors. Finally, the pre-training parameters are frozen, and the text-guided weight module and the cross-attention mechanism are fused to obtain supermodal features and conduct sentiment analysis.
It realizes more accurate identification of emotional tendencies in complex scenarios, inhibits the influence of redundant audio and video information, and improves the accuracy of sentiment analysis.
Smart Images

Figure CN120217309A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of sentiment analysis, and particularly relates to a multi-modal sentiment analysis method based on pre-training and text modality guidance. Background Art
[0002] Sentiment analysis is an important research direction in the fields of natural language processing (NLP) and artificial intelligence (AI), which involves analyzing, understanding, and interpreting human emotions. Multi-modal deep learning interprets and analyzes multi-modal signals together. Since each modality can carry different information, the goal of multi-modal analysis is to integrate these different information sources to obtain a richer and more accurate analysis result than a single modality. Although the information carried by multi-modal video data is relatively rich, the facial expressions (including head orientation, head posture, etc.), speech emphasis of the voice (including volume, pitch), and semantic meaning of the language corresponding to the three modalities of video, audio, and text all reflect the emotional state of the expresser. However, the text content includes a large amount of semantic information, which can directly reflect the speaker's thoughts, intentions, and emotional attitudes. Through text, we can accurately grasp the intuitive meaning and deep meaning of language. For example, by using different words, sentence structures, and language styles, we can express a rich variety of emotions and attitudes. Moreover, compared with other modalities such as sound and video, text data is easy to obtain, process, and analyze, and its processing technology is relatively mature. For example, text tokenization, semantic analysis, and even sentiment tendency determination. In addition, the interpretability of text data is strong, and the analysis results are easy to be understood and accepted by people. However, sentiment analysis mainly based on the text modality also has certain limitations. When there are conflicts in the emotions expressed between single modalities, such as the text content is "You can continue playing. You will definitely get full marks in the exam tomorrow." and "Wow, you are really on time. I almost thought I arrived an hour early.", we cannot judge whether the emotion is positive or negative only from the text modality. In fact, many ironic conversations are presented by non-textual modalities. Therefore, on the basis of taking the text as the guiding modality, distinguishing the consistency and dissimilarity of modality expressions is also the key challenge to make sentiment analysis more accurate. Summary of the Invention
[0003] The purpose of the present invention is to provide a multi-modal sentiment analysis method based on pre-training and text modality guidance, aiming to solve the problem that the user's emotion cannot be recognized based on the text modality and the sentiment analysis result is inaccurate.
[0004] The present invention is implemented as follows. A multi-modal sentiment analysis method based on pre-training and text modality guidance, the method includes: Pre-training stage: Obtain the initial features of video, audio, and text modalities; The initial features of each modality are divided into similar features and dissimilar features by a projector, which consists of a regularization layer, a linear layer with a Tanh activation function, and a Dropout layer; A data sampler is used to calculate the similarity between samples based on similar features and dissimilar features, and positive and negative feature pairs within and between samples are constructed. The similarity between samples is characterized by a cosine similarity score; The contrastive loss and the unimodal prediction error are calculated, and the sum of the two is used for backpropagation to update the parameters of the feature extraction module and the projector; Weight fusion stage: All learnable parameters of the feature extraction module and the projector in the pre-training stage are frozen to form a pre-trained feature extraction tool; The short video is input into the pre-trained feature extraction tool to obtain the similar features and dissimilar features of each modality; The similar features of the text modality and the similar features of the audio and video modalities are input into a text-guided weight module. The weight module learns high-scale text features through two Transformer layers and updates the cross-modal features based on the multi-head attention mechanism and learnable parameters; The updated cross-modal features and the similar features of the text modality are fused through a cross self-attention mechanism and input into a classifier to obtain a prediction result.
[0005] Preferably, the step of using a data sampler to calculate the similarity between samples based on similar features and dissimilar features and constructing positive and negative feature pairs within and between samples includes: Calculate the cosine similarity scores of sample pairs in the dataset, and the cosine similarity scores are based on the concatenated vectors of the initial features of each modality; Sort the same-label samples and different-label samples respectively according to the cosine similarity scores to construct a candidate similar sample set and a candidate dissimilar sample set; Select samples with cosine similarity scores higher than a preset value from the candidate similar sample set to form positive pairs between samples, and select samples with different cosine similarity scores from the candidate dissimilar sample set to form negative pairs between samples.
[0006] Preferably, in the step of calculating the contrastive loss and the unimodal prediction error and using the sum of the two for backpropagation to update the parameters of the feature extraction module and the projector, the calculation of the contrastive loss includes the contrast between positive and negative pairs within samples and positive and negative pairs between samples. The positive and negative pairs within samples use the similar features of the text modality as anchors to constrain the visual and auditory similar features to approach it and the dissimilar features of each modality to move away from each other; the positive and negative pairs between samples are based on the positive and negative sample pairs constructed by the data sampler.
[0007] Preferably, in the step of calculating the contrastive loss and the unimodal prediction error and performing backpropagation using the sum of the two to update the parameters of the feature extraction module and the projector, the calculation method of the unimodal prediction error is as follows: The similar features and dissimilar features of each modality are input into an MLP classifier with shared weights. The similar features are used to predict the multimodal label, and the dissimilar features are used to predict the unimodal label or the multimodal label. The cross-entropy loss is calculated based on the prediction results and the true labels and accumulated.
[0008] Preferably, the text-guided weight module includes: Two Transformer layers for taking the text-modal similar features as low-scale text features and learning to generate medium-scale and high-scale text features; Three weight layers for calculating the similarity weight matrices of text with audio and text with video modalities respectively based on the high-scale text features through the multi-head attention mechanism and updating the super-modal features.
[0009] Preferably, the update formula of the super-modal features is: ; where α and β are the similarity weight matrices of text with audio and text with video respectively, and respectively represent the th text-guided weight layer and the corresponding output super-modal features, and are learnable parameters, and respectively represent the values in the weight layer.
[0010] Preferably, in the step of fusing the updated super-modal features with the text-modal similar features through the cross self-attention mechanism and inputting them into the classifier to obtain the prediction results, the text-modal similar features are used as the query Query, the super-modal features are used as the key Key and the value Value, and cross-modal feature fusion is achieved through self-attention calculation.
[0011] Preferably, the calculation formula of the cosine similarity score is: ; where, represents the cosine similarity between vector w and vector v, represents the concatenation operation of vectors, and are two different samples, , , respectively represent the initial representations of each modality under the sample, respectively represent The initial representations of each modality under the samples is the matrix transpose.
[0012] Another object of the present invention is to provide a multi-modal sentiment analysis system based on pre-training and text modality guidance, and the system includes: A feature extraction module, configured to obtain initial features of video, audio, and text modalities, where the text modality features are extracted by Bert, and the video and audio modality features are encoded by a Transformer encoder; A projector module, which consists of a regularization layer, a linear layer with a Tanh activation function, and a Dropout layer, and is configured to divide the initial features of each modality into similar features and dissimilar features; A data sampler module, configured to construct positive and negative feature pairs within and between samples based on cosine similarity scores; A pre-training module, configured to calculate a contrastive loss and a unimodal prediction error, and update the parameters of the feature extraction module and the projector module through backpropagation; A weight fusion module, including a text-guided weight module and a cross self-attention mechanism fusion layer, and is configured to update the super-modal features based on high-scale text features and perform cross-modal fusion classification after freezing the pre-training parameters.
[0013] Preferably, the text-guided weight module includes two Transformer layers and three weight layers. The Transformer layers are used to learn high-scale text features, and the weight layers realize the adaptive update of super-modal features based on the multi-head attention mechanism and learnable parameters.
[0014] The multi-modal sentiment analysis method based on pre-training and text modality guidance provided by the present invention uses a text-guided weight module to suppress redundant audio-visual information; through a data sampler and a feature contrast learning method, it realizes multi-modal fusion with the text modality as the anchor point, can more accurately identify the sentiment tendency in complex scenarios, and has wide value in the fields of sentiment classification, social media content recommendation, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is the overall flowchart of the multi-modal sentiment analysis method based on pre-training and text modality guidance provided by the embodiment of the present invention; Figure 2 is the architecture diagram of the multi-modal sentiment analysis system based on pre-training and text modality guidance provided by the embodiment of the present invention; Figure 3 is the state diagram of the pre-training stage provided by the embodiment of the present invention; Figure 4 is the state diagram of the weight fusion stage provided by the embodiment of the present invention; Figure 5 It is the architecture diagram of the weight module provided by the embodiment of the present invention; Figure 6 It is the data sampling flow chart provided by the embodiment of the present invention. Detailed implementation manners
[0016] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, but not to limit the present invention.
[0017] As Figure 1 shown, it is the overall flowchart of the multi-modal sentiment analysis method based on pre-training and text modality guidance provided by the embodiment of the present invention. The method includes: As Figure 3 shown, pre-training stage: S1. Obtain the initial features of video, audio and text modalities.
[0018] In this step, the features of the video modality, audio modality and text modality of the social network short video are obtained from a short social video with sound. The original data includes audio clips, video clips and dialogue texts. The parameters for extracting multi-modal features are learned through pre-training. Based on the foregoing extraction of the three modalities of text, video and audio, the text modality extracts text features through Bert, and the video and audio modalities are respectively encoded by two Transformer encoders to obtain the initial features T, V and A.
[0019] Among them, Bert is a deep bidirectional language representation model based on the Transformer architecture, which realizes the dynamic encoding of text context information through a multi-layer self-attention mechanism. Different from traditional unidirectional language models, Bert adopts a joint training strategy of masked language model and next sentence prediction, enabling it to capture the bidirectional semantic associations before and after the text sequence at the same time. In the multi-modal sentiment analysis system described in the present invention, the Bert module is specifically used to extract the deep semantic features of the video text modality: 1) Capture the context dependence of sentiment keywords through a word-level attention weight matrix; 2) Realize sentence-level sentiment tendency encoding by using the aggregated representation of the [CLS] special token.
[0020] S2. Use a projector to distinguish the initial features of each modality into similar features and dissimilar features. The projector consists of a regularization layer, a linear layer with a Tanh activation function, and a Dropout layer.
[0021] In this step, the feature encoding is distinguished into similar features and dissimilar features by the projectors of similar features and dissimilar features, and the unimodal sentiment is predicted by the similar features and dissimilar features of each modality, where the dimensionality sizes of the similar and dissimilar features of each modality are unified.
[0022] According to the initial features of the three modalities described above, they are respectively put into three different projectors to divide the encoding of each modality into similar features and dissimilar features. (T s and T d represent the similar features and dissimilar features of the text modality, V s and V d represent the similar features and dissimilar features of the video modality, A s and A d represent the similar features and dissimilar features of the audio modality) Among them, the projector is composed of a regularization layer, a linear layer with a Tanh activation function, and a dropout layer. The role of this design is to project the features into a canonical space through a non-linear mapping, and to make the similar modality features and dissimilar features under the same modality have a comparable scale by constraining the feature amplitude range. The purpose of this design is to let the similarity features capture the consistent information shared between different modalities through the overall multi-modal labels of the samples, while the dissimilarity features can retain the mode-specific information represented by the unimodal labels.
[0023] S3. Use the data sampler to calculate the similarity between samples based on the similar features and dissimilar features, and construct positive and negative feature pairs within and between samples. The similarity between samples is characterized by the cosine similarity score.
[0024] S4. Calculate the contrast loss and the unimodal prediction error, and use the sum of the two for backpropagation to update the parameters of the feature extraction module and the projector.
[0025] In this step, the data sampler designed is used to calculate the similarity between samples of each batch using the similar and dissimilar features of the three modalities. Based on the similarity calculation, positive and negative samples are distinguished, and the positive and negative feature pairs within and between samples are distinguished with the text as the anchor point. At the same time, the contrast loss and the unimodal prediction error are calculated.
[0026] According to the feature projector described above, the features of the three modalities are divided into similar features and dissimilar features. Based on the premise of the feature contrast loss function, the similarity of a batch of samples is judged. Therefore, the present invention designs a data sampler.
[0027] The data sampler retrieves similar samples of a given sample through the multi-modal features and labels of the samples, and further performs supervised contrast learning between samples. The sampling process includes: According to the above construction of the positive and negative pairs of features within and between samples, the construction method is derived from the data sampler designed by the present invention. The data sampler retrieves similar samples of a given sample based on multimodal features and multimodal labels, thereby performing supervised comparative learning between samples. The sampling process is as follows: like Figure 6 As shown, first, the cosine similarity scores of sample pairs are calculated in a dataset D containing |D| samples.
[0028] ; in represents the cosine similarity score between vector w and vector v, Represents the concatenation operation of vectors, and For two different samples, , , Respectively The initial representation of each mode under the sample, Respectively The initial representation of each mode under the sample, Transpose the matrix.
[0029] Then, we retrieve the candidate similar and dissimilar sample sets for each sample. , according to the similarity score from high to low, the same multimodal label The sample ranking is used as the candidate similar sample set , label The sample ranking is used as the candidate different sample set From the candidate similar sample set Randomly select two approximate samples with higher cosine similarity scores from the sample i to form a positive pair between samples, denoted as ; From the candidate dissimilar sample set Randomly select four different samples to form negative pairs between samples, denoted as (Two of the samples With lower cosine similarity scores, the other two samples has a higher cosine similarity score). Samples and samples in have different labels and similar semantic information, so The samples in the sample are difficult to compare with the sample Therefore, The samples in are also added to In the process of contrast learning, Samples and samples in Distinguish.
[0030] Based on the construction of positive and negative pairs of samples by the above data sampler, the positive and negative pairs of each modal feature of the samples are further constructed, mainly including the construction of positive and negative pairs within the samples and the construction of positive and negative pairs between the samples.
[0031] Among them, the construction of positive and negative pairs within the samples uses six decomposed features to form positive and negative pairs within the samples. Select the text similarity feature as the anchor point, so that the visual and auditory similarity features and move closer to , and at the same time move the dissimilarity features in all modalities , , away from .
[0032] ; is the positive sample pair within the sample, is the negative sample pair within the sample, and represent the similar samples and dissimilar samples of the sample respectively, and the role of is to expand the comparison range.
[0033] Among them, the construction of positive and negative pairs between the samples is to construct positive / negative pairs between the samples according to the and sampled by the data sampler.
[0034] ; Among them is the positive sample pair within the sample, is the negative sample pair within the sample, is the positive sample of the sample , and k is the negative sample of the sample .
[0035] is composed of the common features of different modalities in one piece of information. is composed of the common features of the text modality in one piece of information and the unique features of all modalities. Then it is composed of the common features between different pieces of information. In the formula, represents the set of positive examples of the th piece of information, represents the set of negative examples of the th piece of information.
[0036] Calculate the contrast loss and single-modal prediction error according to the above features: Contrastive loss: The contrast loss is a simple joint contrast loss that compares similar samples with different samples in a batch and similar features with different features in a sample.
[0037] ; in represents the contrast loss of sample i; The feature vector representing a pair of decomposed features (into similar features and different features) can be the feature vector within the sample, such as , or it can be a feature vector between different samples, such as ,Notice ; Represents the set of positive pairs, including the positive pairs within the sample Positive alignment between samples , ; Represents the negative pair set, including negative pairs within the sample and negative pairs between samples , ,Contrastive loss aims to shorten the cross-modal feature distance of similar samples and push away the feature distance of heterogeneous samples; is the temperature parameter, which regulates the learning intensity of the model for similar and different features.
[0038] Unimodal prediction error : For each sample , the 6 decomposed features are input into a weight-sharing MLP classifier respectively, and 6 prediction results are obtained. , and Predicting multimodal labels via MLP mapping , different characteristics , and Predicting modality-specific labels via MLP mapping , and , but when the modality-specific label is not available, the different features are also used to predict the multimodal label The principle is to allow similar features to capture the consistent information shared between different modalities through the overall multimodal label of the sample, and to allow different features to retain the specific modality information represented by the unimodal label.
[0039] ; in is the predicted label, It is the true label.
[0040] In this step, backpropagation is performed using the sum of the contrastive loss and the unimodal prediction error to update the relevant parameters of the feature extraction and projector. Among them is the unimodal loss, is the contrastive loss.
[0041] Backpropagation and parameter update: Based on the total loss function The backpropagation algorithm is adopted to calculate the gradients of the total loss function with respect to the weight parameters of each unit in the feature extraction module (video feature extraction unit, audio feature extraction unit, text feature extraction unit), and the mapping matrix parameters of the projector; Using optimizers such as stochastic gradient descent or adaptive moment estimation (Adam), the above parameters are updated according to the calculated gradients, so that the total loss function gradually decreases until the preset convergence condition is reached.
[0042] As Figure 4 shown, the weight fusion stage: S5. Freeze all the learnable parameters of the feature extraction module and the projector in the pre-training stage to form a pre-trained feature extraction tool; input the short video into the pre-trained feature extraction tool to obtain the similar features and dissimilar features of each modality.
[0043] In this step, freeze all the learnable parameters of the above-mentioned feature extraction and projector, regard the first stage (pre-training stage) as a trained feature extraction tool, input a short video containing video modality, audio modality and text modality to the pre-trained feature extraction tool, and obtain the similar features and dissimilar features of video modality, audio modality and text modality.
[0044] S6. Input the similar features of the text modality and the similar features of the audio and video modalities into the text-guided weight module. The weight module learns high-scale text features through two Transformer layers and updates the super-modal features based on the multi-head attention mechanism and learnable parameters.
[0045] In this step, as Figure 5 shown, based on the aforementioned text-guided weight module, which includes two Transformer layers and three weight layers, aims to learn language features of different scales and adaptively learn super-modal features from visual and audio modalities under the guidance of language features.
[0046] In the two Transformer layers of the weight module, according to the similar representation of the text modality , regarded as low-scale text features . By inputting low-scale text features into the two introduced Transformer layers to learn the medium-scale and high-scale text features, namely and . In this stage, the text features are modeled as: ; wherein, , is the text feature of different scales (low scale, medium scale, and high scale) with a size of , and represent the ir-th Transformer layer and the corresponding parameters for text feature learning.
[0047] On the other hand, in the weight layer, according to the language features of the different scales , a super-modal feature is first initialized, and then the super-modal feature is updated by calculating the relationship between the obtained language features and using the remaining two modalities with multi-head attention. The specific update method is as follows: is used as Query, is used as Key, and the similarity weight matrix ɑ between the text feature and the audio feature is obtained, ; where softmax represents the weight normalization operation, is the representation of used as Query, is the similar feature of the audio modality used as Key, and are learnable parameters, is the similar feature of the audio modality, is the dimension of each attention head. (In actual operation, 8-head attention is used, and is set to 16).
[0048] Similar to ɑ, β represents the similarity weight matrix between the language modality and the visual modality, ; where is the representation of used as Query, is the similar feature of the video modality used as Key, and are learnable parameters, For the similar features of the video modality, is the dimension of each attention head. (In actual operation, 8-head attention is used and is set to 16) Based on the above method for updating the hyper-modal features, the hyper-modal features can be updated to the said hyper-modal features by weighted audio features and weighted video features. The update formula is: ; where α and β are the similarity weight matrices between text and audio, and text and video respectively, and respectively represent the th text-guided weight layer and the corresponding output hyper-modal features, and are learnable parameters, and respectively represent the Value values in the weight layer.
[0049] This module learns different-scale language features through initial feature extraction and two Transformer layers, and learns an adaptive hyper-modal representation under this guidance. The features of the text modality are always set as the Query of the weight layer because the language information is relatively clean and can provide more emotion-related information for effective multi-modal sentiment analysis tasks. The information from the visual and audio modalities that is irrelevant and conflicting with the true emotion may limit the prediction accuracy of the model. This module effectively suppresses the adverse effects of redundant information in the visual and audio modalities. However, the visual and audio modalities cannot be removed either because the visual and audio modalities can also provide more complementary information.
[0050] S7. Fuse the updated hyper-modal features with the similar features of the text modality through the cross self-attention mechanism, and input them into the classifier to obtain the prediction result.
[0051] According to the and output by the text modality-guided weight module above, they are respectively connected with the initialized sequence representation to obtain the new text modality features and the new hyper-modal features . Then, a joint multi-modal representation H is obtained through the cross-modal Transformer layer, and finally the joint multi-modal representation H is input into a classifier to obtain the final prediction output .
[0052] ; where is the high-scale language feature passing through two Transformer layers, is the hypermodal feature output by the third weight layer. The cross-modal Transformer layer fuses the text feature with the new hypermodal feature , and is used as the Q (query) of the Transformer, and the new hypermodal feature is used as K and V (key and value).
[0053] Based on the above overall training process, the present invention sets the structure of the loss function as follows, Multimodal prediction loss :(Mean Absolute Error) ; where n is the number of samples in batch B, is the multimodal prediction label, is the multimodal true label.
[0054] As Figure 2 shown, the present invention also provides a multimodal sentiment analysis system based on pre-training and text modality guidance. The system includes: A feature extraction module for obtaining the initial features of video, audio, and text modalities. The text modality features are extracted by Bert, and the video and audio modality features are encoded by a Transformer encoder; A projector module composed of a regularization layer, a linear layer with a Tanh activation function, and a Dropout layer for dividing the initial features of each modality into similar features and dissimilar features; A data sampler module for constructing positive and negative feature pairs within and between samples based on cosine similarity scores; A pre-training module for calculating the contrast loss and the unimodal prediction error, and updating the parameters of the feature extraction module and the projector module through backpropagation; A weight fusion module including a text-guided weight module and a cross-self-attention mechanism fusion layer for updating the hypermodal features and performing cross-modal fusion classification based on high-scale text features after freezing the pre-training parameters.
[0055] The text-guided weight module contains two Transformer layers and three weight layers. The Transformer layers are used to learn high-scale text features, and the weight layers realize the adaptive update of hypermodal features based on the multi-head attention mechanism and learnable parameters.
[0056] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-modal sentiment analysis method based on pre-training and text modality guidance, characterized in that The method includes: Pre-training stage: Obtain initial features of video, audio, and text modalities; Use a projector to distinguish the initial features of each modality into similar features and dissimilar features; Use a data sampler to calculate the similarity between samples based on similar features and dissimilar features, and construct positive and negative feature pairs within and between samples. The similarity between samples is characterized by a cosine similarity score; Calculate the contrastive loss and the unimodal prediction error, and use the sum of the two for backpropagation to update the parameters of the feature extraction module and the projector; Weight fusion stage: Freeze all learnable parameters of the feature extraction module and the projector in the pre-training stage to form a pre-trained feature extraction tool; Input the short video into the pre-trained feature extraction tool to obtain similar features and dissimilar features of each modality; Input the similar features of the text modality and the similar features of the audio and video modalities into a text-guided weight module. The weight module learns high-scale text features through two Transformer layers and updates the super-modal features based on the multi-head attention mechanism and learnable parameters; Fuse the updated super-modal features and the similar features of the text modality through the cross self-attention mechanism, and input them into a classifier to obtain the prediction result.
2. The multimodal sentiment analysis method based on pre-training and text modality guidance according to claim 1, wherein The step of using a data sampler to calculate the similarity between samples based on similar features and dissimilar features and construct positive and negative feature pairs within and between samples includes: Calculate the cosine similarity score of sample pairs in the dataset. The cosine similarity score is based on the concatenated vector of the initial features of each modality; Sort the same-label samples and different-label samples respectively according to the cosine similarity score to construct a candidate similar sample set and a candidate dissimilar sample set; Select samples with a cosine similarity score higher than a preset value from the candidate similar sample set to form positive pairs between samples, and select samples with different cosine similarity scores from the candidate dissimilar sample set to form negative pairs between samples.
3. The multimodal sentiment analysis method based on pre-training and text modality guidance according to claim 1, characterized in that In the step of calculating the contrastive loss and the unimodal prediction error, and using the sum of the two for backpropagation to update the parameters of the feature extraction module and the projector, the calculation of the contrastive loss includes the comparison of positive and negative pairs within and between samples. The positive and negative pairs within samples use the similar features of the text modality as the anchor point to constrain the visual and auditory similar features to approach it and the dissimilar features of each modality to move away from each other. The positive and negative pairs between samples are based on the positive and negative sample pairs constructed by the data sampler.
4. A multimodal sentiment analysis method based on pre-training and text modality guidance according to claim 1, characterized in that, In the step of calculating the contrastive loss and the unimodal prediction error, and using the sum of the two for backpropagation to update the parameters of the feature extraction module and the projector, the calculation method of the unimodal prediction error is: input the similar features and dissimilar features of each modality into a weight-sharing MLP classifier. The similar features predict the multi-modal label, and the dissimilar features predict the unimodal label or the multi-modal label. Calculate the cross-entropy loss based on the prediction result and the true label and accumulate it.
5. A multimodal sentiment analysis method based on pre-training and text modality guidance according to claim 1, characterized in that, The text-guided weight module includes: Two Transformer layers, which are used to use the similar features of the text modality as low-scale text features to learn and generate medium-scale and high-scale text features; Three weight layers, based on the high-scale text features, respectively calculate the similarity weight matrices of text and audio, and text and video modalities through the multi-head attention mechanism, and update the hyper-modal features.
6. A multimodal sentiment analysis method based on pre-training and text modality guidance according to claim 1, characterized in that The super-modal feature has the following update formula: ; Among them, α and β are the similarity weight matrices of text-audio and text-video respectively, and respectively represent the th text-guided weight layer and the corresponding output cross-modal features, and are learnable parameters, and respectively represent the values in the weight layer.
7. A multimodal sentiment analysis method based on pre-training and text modality guidance according to claim 1, characterized in that In the step of fusing the updated hyper-modal features with the text-modal similarity features through the cross self-attention mechanism and inputting them into the classifier to obtain the prediction results, the text-modal similarity features are used as the query Query, and the hyper-modal features are used as the key Key and value Value, and cross-modal feature fusion is achieved through self-attention calculation.
8. A multimodal sentiment analysis method based on pre-training and text modality guidance according to claim 1, characterized in that The calculation formula of the cosine similarity score is: ; Among them, represents the cosine similarity between vector w and vector v, represents the concatenation operation of vectors, and are two different samples, 、 、 respectively represent the initial representations of each modality under the sample, respectively represent the initial representations of each modality under the sample, is the matrix transpose.
9. A multimodal sentiment analysis method based on pre-training and text modality guidance according to claim 1, characterized in that The projector consists of a regularization layer, a linear layer with a Tanh activation function, and a Dropout layer.
10. A multimodal sentiment analysis method based on pre-training and text modality guidance according to claim 1, characterized in that The weight module learns high-scale text features through two Transformer layers, and updates the hyper-modal features based on the multi-head attention mechanism and learnable parameters.
Citation Information
Patent Citations
Multi-modal sentiment analysis method, system and equipment and storage medium
CN115809438A
Multi-modal sentiment analysis method based on feature decoupling and graph attention network
CN119691552A
Multimodal sentiment analysis method based on comparative decomposition feature multitask network
CN119848500A
Cited By
Multi-modal emotion intelligent analysis method based on pre-training model
CN120873450A
Multi-modal depression sentiment analysis system based on sensitive attribute tag constraint
CN121723403A
Vascular deformation prediction method in minimally invasive surgery video
CN121904048A