Multistage cross-modal alignment method based on comparative learning
Through the global and local cross-modal alignment method combined with the multi-task learning framework, the semantic gap and misalignment problems between modes in multi-modal sentiment analysis are solved, and more accurate user sentiment analysis is achieved, especially the recognition of aspect words and emotional polarity.
Patent Information
- Application Number
- CN202510476484.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art cannot effectively use multimodal data for user sentiment analysis, especially the inability to accurately obtain users' emotional judgments on specific aspects, and the semantic gap and misalignment problems between modals have not been fully solved, resulting in insufficient efficiency and accuracy of sentiment analysis.
The multi-level cross-modal alignment method based on contrast learning is adopted, and the text and image representations are aligned through the global cross-modal alignment module and the local cross-modal alignment module, and sequence label prediction is carried out in combination with the multi-task learning framework and conditional random fields to achieve coarse and fine-grained alignment of text and image.
The accuracy and efficiency of multimodal sentiment analysis are improved, especially the recognition ability in terms of terms and affective polarity in terms of recognition, which significantly improves the performance of the model.
Smart Images

Figure CN120372545A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of natural language processing and computer vision. Specifically, it relates to a multi-level cross-modal alignment method based on contrastive learning for improving the accuracy and efficiency of multi-modal sentiment analysis. Background Art
[0002] The rapid development of social networks not only brings convenience and speed to interpersonal communication, enabling people to share life details with those thousands of miles away without leaving home, but also makes emotional expression closer to reality. Different from the early days of social networks when the content was limited to pure text information, nowadays, to facilitate users to convey their thoughts more appropriately, the expression method has evolved from a single text description to the current combination of multi-modal content such as kaomoji, emoticons, pictures, and videos. Just as the emotional transmission during face-to-face communication depends more on factors such as body language and tone of voice, the diversification of content on many social platforms led by WeChat Moments, Weibo, and Twitter (Twitter) has also made the classification of users' emotional tendencies at the current stage complex and unable to achieve high efficiency and accuracy. The previous single-modal text sentiment analysis can no longer effectively utilize the diversified content in the big data era.
[0003] Multi-modal sentiment analysis is an important part of moving towards fully using multiple modal information for sentiment-based representation learning. Enterprises can also use this to understand the key concerns of users and the trends of social hotspots, so as to better adjust company policies to improve user satisfaction. However, most current multi-modal sentiment analyses treat multi-modal data together to judge the emotional polarity of users, but cannot obtain the emotional judgment of users on specific aspects. Therefore, focusing on aspect words can consider aspect-level emotional information on the basis of grasping the overall situation. However, it is very difficult to extract the corresponding emotional information of images on the basis of extracting aspect words.
[0004] Therefore, some researchers consider proposing to jointly extract aspect words and classify their corresponding emotions on the basis of text-image relationship detection in the annotated dataset. However, this method mainly focuses on the global interaction across modalities and cannot well solve the fine-grained correspondence relationship between images and texts. Although the existing technology can fuse text and visual information to a certain extent by extracting feature information from text and images through text and visual encoders and then mapping them to the same-dimensional mapping space for interaction. However, since images and texts are two modalities with a large semantic gap, irrelevant visual objects often have a negative impact on the fusion of text and visual modalities during the interaction process. Therefore, the limitation of these fusion methods lies in that they do not align the semantics of visual objects and text content before fusion; due to the misalignment between their modalities, the interaction problem has not been fully solved.
[0005] Therefore, in terms of cross-modal alignment, although pre-training can be carried out based on a large amount of labeled data, it consumes a large amount of human and computing resources. Moreover, few studies have explored bridging the semantic gap between modalities and effectively using visual information for cross-modal fusion from coarse-grained to fine-grained alignment. The multi-level cross-modal alignment (CMCA) method based on contrastive learning proposed in this study is precisely to solve this problem. The CMCA method realizes the alignment of text and image at the coarse-grained and fine-grained levels by designing two modules: global cross-modal alignment (GCMA) and local cross-modal alignment (LCMA). In the GCMA module, positive and negative examples are generated through contrastive learning, and a multi-layer perceptron (MLP) is used to project the text and image representations into the same space, and the contrastive loss function is minimized to maximize the similarity of positive examples and minimize the similarity of negative examples. In the LCMA module, a cross-attention mechanism is adopted for fine-grained alignment to identify and correlate smaller and more specific semantic units in the image and text, so as to more accurately capture the correspondence between the two modalities. In addition, a multi-task learning framework is introduced, and sequence label prediction is performed through a conditional random field (CRF) to further improve the performance of the model. Through this method, multi-modal information can be effectively aligned and fused without adding a large amount of labeled data, improving the accuracy and efficiency of sentiment analysis. Summary of the Invention
[0006] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a multi-level cross-modal alignment method based on contrastive learning, which is used to solve the problems in the prior art such as single-modal sentiment analysis being unable to effectively utilize multi-modal data, multi-modal sentiment analysis methods being unable to obtain users' sentiment judgments on specific aspects, lack of fine-grained correspondence, limitations of global interaction, semantic gap problems, cross-modal misalignment problems, and dependence on a large amount of labeled data.
[0007] To achieve the above purpose and other related purposes, the present invention provides a multi-level cross-modal alignment method based on contrastive learning, which includes the following steps:
[0008] Step 1: Use the pre-trained RoBERTa model as a text encoder to encode the input text to obtain a text representation; use the pre-trained Vision Transformer model as an image encoder to encode the input image to obtain an image representation;
[0009] Step 2: Through the global cross-modal alignment module, adopt contrastive learning technology to align the text and image representations to enhance the representational consistency between the two;
[0010] Step 3: Through the local cross-modal alignment module, use the cross-attention mechanism to perform fine-grained alignment on the text representation and the image representation;
[0011] Step 4. Integrate cross-modal information from text and images using a multi-task learning framework;
[0012] Step 5. Perform sequence label prediction through a conditional random field to identify and classify aspect terms and sentiment.
[0013] Further, the specific steps of the said Step 1 are as follows:
[0014] Step 1-1. Generate positive and negative examples from a batch of N (T s , V g ) input pairs. For the a-th pair in the batch, is the text representation, is the image representation. Assume that the positive example is the text and image representations from the same input pair, i.e., while the negative example is composed of representations from different input pairs, i.e., For each pair of inputs in the batch, we can obtain one positive example and N - 1 negative examples. The effect of contrastive learning is mainly affected by the number of negative examples, and the number of negative examples is positively correlated with the effect. The influence of a small number of mismatched pairs that may appear in the positive examples is negligible;
[0015] Step 1-2. For each example apply two different multi-layer perceptrons respectively, each MLP containing one hidden layer, and apply them to and to obtain the projected text representation and the image representation It is found that this MLP projection can help the encoder learn better representations;
[0016] Step 1-3. Maximize the similarity of positive examples and minimize the similarity of negative examples by minimizing two contrastive loss functions, which are the image-to-text contrastive loss and the text-to-image contrastive loss respectively. The image-to-text contrastive loss function is defined for the i-th positive projection pair in the batch as follows:
[0017]
[0018] The text-to-image contrastive loss is defined for the i-th positive projection pair as follows:
[0019]
[0020] Finally, sum the two losses of all positive projection pairs in the batch to obtain the total loss function:
[0021]
[0022] Among them, λ c ∈ [0,1] is a hyperparameter. By minimizing the loss function, the representations of the text encoder and the image encoder will be more consistent.
[0023] Furthermore, step 2 includes the following steps:
[0024] Step 2-1: For each pair of text and image representations, generate positive sample pairs and negative sample pairs;
[0025] Step 2-2: Use different multi-layer perceptrons to map the text and image representations to a common feature space;
[0026] Step 2-3: By minimizing the image-to-text contrast loss and the text-to-image contrast loss, maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs;
[0027] Step 2-4: By dynamically adjusting the weight parameters in the contrast loss function, adaptively optimize the cross-modal consistency according to the similarity between the text and the image.
[0028] Furthermore, step 3 includes the following steps:
[0029] Step 3-1: Take the text representation as the query, the image representation as the key and value, and calculate the attention weights between the text and the image through the multi-head cross-attention mechanism, where the text representation while the image representation is The formula is as follows:
[0030]
[0031] Among them are the parameter matrices of the query, key, and value respectively. Through the multi-head cross-attention mechanism, the final visually-perceived text representation can be obtained: The calculation is as follows:
[0032]
[0033] where W p is the weight matrix of the multi-head cross-attention;
[0034] Step 3-2: Optimize the feature fusion process through layer normalization and residual connection to enhance the model's ability to capture local features;
[0035] Step 3-3: Use a conditional random field to perform sequence labeling on the fused text representation to identify aspect terms and their sentiment polarities in the text. The calculation is as follows:
[0036]
[0037] Among them, A is the state transition matrix, and its element A i,j represents the transition score from label i to label j, is the weight vector, and the loss function adopts negative log probability, which is specifically expressed as:
[0038]
[0039] Furthermore, the multi-task learning framework in step 4 includes two sub-tasks to integrate cross-modal information from text and images and perform sequence label prediction through conditional random fields. The two sub-tasks are respectively:
[0040] ①Aspect term extraction, responsible for identifying all aspect terms from the text;
[0041] ②Aspect sentiment classification, predicting the sentiment polarity associated with each identified aspect term;
[0042] Through the shared representation layer and feature fusion mechanism, information interaction and collaborative optimization between the two sub-tasks are achieved.
[0043] Even further, the following steps are also included:
[0044] Step 6-1, during the training process, use unaligned large-scale text data to pre-train the text encoder to improve the model's understanding ability of text semantics;
[0045] Step 6-2, in the fine-tuning stage, jointly fine-tune the pre-trained text encoder and the image encoder to learn cross-modal feature representations;
[0046] Step 6-3, through contrastive learning and cross-modal attention mechanism, enhance the model's recognition ability of aspect terms and sentiment polarities in multi-modal data;
[0047] Step 6-4, in the test stage, use the trained model to perform aspect term extraction and sentiment polarity analysis on new multi-modal data, and output aspect terms and their corresponding sentiment labels.
[0048] As described above, the multi-level cross-modal alignment method based on contrastive learning of the present invention has the following beneficial effects:
[0049] (1) Through two modules of global cross-modal alignment (GCMA) and local cross-modal alignment (LCMA), text and image alignment is achieved at the coarse-grained and fine-grained levels, effectively solving the problem of large semantic gaps between modalities and improving the fusion effect of multi-modal information.
[0050] (2) Introduce a multi-task learning framework and Conditional Random Field (CRF) for sequence label prediction, which significantly improves the accuracy and efficiency of sentiment analysis, especially in identifying aspect terms and sentiment polarities in multi-modal data. Description of the Drawings
[0051] Figure 1 is the overall architecture diagram of the multi-level cross-modal alignment model based on contrastive learning provided by the present invention.
[0052] Figure 2 is the application flowchart of the multi-level cross-modal alignment method based on contrastive learning of the present invention in multi-modal sentiment analysis. Detailed Embodiments
[0053] The present invention will be further described below with reference to the accompanying drawings by way of examples, but the scope of the present invention is not limited in any way.
[0054] The present invention provides a multi-task framework that uses only the text modality to assist multi-modal prediction, enhances the alignment of multi-modal features, and obtains better performance. A multi-level cross-modal alignment method based on contrastive learning is proposed for sentiment analysis in multi-modal aspects.
[0055] Both ViT and RoBERTa are initialized from pre-trained checkpoints and are used to encode visual and language modalities.
[0056] To align the learned features, the modality adopts a multi-task learning architecture containing two sub-tasks to integrate cross-modal information of text and images. In addition to the cross-attention module, contrastive learning is introduced to align multi-modal features, thereby improving the efficiency of MABSA task prediction.
[0057] The network structure diagram and flowchart of the method of the present invention are as shown in the attached Figure 1 and the attached Figure 2 and are as follows. Specifically, when implemented, the following steps are included:
[0058] Step 1, Feature extraction: It is further divided into feature extraction for images and feature extraction for text;
[0059] ① Feature extraction for images: First, use the Vision Transformer (ViT) to process the input image. The image is segmented into multiple patches, and each patch is mapped to a feature space through linear projection. After being processed by ViT, these features generate a series of image feature vectors V1, V2,..., V max .
[0060] ② Feature extraction in terms of text: First, use the RoBERTa model to process the input text. The text is segmented into multiple tokens through WordPiece Tokenization. After being processed by RoBERTa, these tokens generate a series of text feature vectors. T1, T2, ..., T max .
[0061] Step 2: Input the image feature vectors and text feature vectors into the Cross-Modal Alignment Module. This module generates positive and negative sample pairs by using the contrastive learning method and projects them into a new space through MLP to obtain T c and V c . Define the contrastive loss function from image to text as:
[0062]
[0063] where sim(·) represents the cosine similarity, τ is the temperature parameter, which is a hyperparameter. Define the contrastive loss function from text to image as:
[0064]
[0065] The total contrastive loss function is:
[0066]
[0067] where λ c is the hyperparameter for balancing the two losses.
[0068] Step 3: Use the multi-head cross-attention mechanism to align the image and text features, where the text representation is set as the query, and the image representation serves as the key and value. The formula is as follows:
[0069]
[0070] where are the parameter matrices of the query, key, and value respectively. Through the multi-head cross-attention mechanism, the final visual perception text representation can be obtained: The calculation is as follows:
[0071]
[0072] where W p is the weight matrix of the multi-head cross-attention.
[0073] Step 4: Use LayerNorm to normalize the fused features to generate the final multi-modal representation.
[0074] Step 5: Image Enhanced Prediction. Use Conditional Random Field (CRF) to process the multi-modal representation to generate the final image prediction result. At the same time, process the aligned text features through a linear layer (Linear) to output multiple labels (such as B-POS, I, O, B-ENU, B-NEG).
[0075] Through multi-level cross-modal alignment, this model first aligns and fuses image and text features at the global and local levels, and then makes predictions for images and texts respectively. This architecture can effectively utilize the complementary information of images and texts, improving the accuracy and robustness of predictions. Through global and local alignment, the model can better understand and integrate information from different modalities, thus achieving better performance in tasks such as multi-modal sentiment analysis.
[0076] In the global alignment stage, the model uses the method of contrastive learning and utilizes positive and negative sample pairs to enhance the consistency between text and image features. In the local alignment stage, the model achieves a finer-grained alignment through the cross-attention mechanism, which helps the model more accurately capture the semantic relationship between images and texts. Finally, through the CRF layer and the linear layer, the model can effectively make predictions for images and texts and output the corresponding labels.
[0077] We compared our experimental model with other baseline models and fully demonstrated the excellent results achieved by the proposed CMCA model. The results of different models are shown in Table 1 below:
[0078]
[0079] Table 1
[0080] Compared with the above baseline models, our model exceeded the AoM in 2023 in the three evaluation metrics (F1, P, R) of the two datasets. Among them, in terms of the P metric, compared with the previously best-performing Aom model, our model improved by 0.7% and 1.7% respectively; this fully demonstrates the effectiveness of our multi-level alignment method. In addition, the F1 metrics of the two datasets increased by 0.4% and 2.1% respectively. We analyzed and speculated that this is because the above models may have ignored the simultaneous attention to global and local information, which may lead to misjudgments of the number and position of aspect words, thus significantly affecting the comprehensive evaluation metric F1. This also shows the effectiveness of the proposed CMCA model in integrating multi-granularity information.
[0081] This study proposes a multi-level cross-modal alignment architecture. By contrastive learning, it learns the common features between modalities to achieve global-grained alignment of text and images, and introduces a cross-modal attention mechanism to achieve finer-grained text-image alignment. At the same time, a multi-task architecture is adopted, using a single text modality to supervise multi-modal prediction. Experimental results show that this model improves the performance of MABSA on the Twitter-2015 and Twitter-2017 datasets.
Claims
1. A multi-level cross-modal alignment method based on contrastive learning, characterized in that The method includes the following steps: Step 1: Use a pre-trained RoBERTa model as a text encoder to encode the input text to obtain a text representation; use a pre-trained Vision Transformer model as an image encoder to encode the input image to obtain an image representation; Step 2: Through the global cross-modal alignment module, adopt contrastive learning technology to align the text and image representations to enhance the representational consistency between the two; Step 3: Through the local cross-modal alignment module, use the cross-attention mechanism to perform fine-grained alignment on the text representation and the image representation; Step 4: Utilize a multi-task learning framework to integrate cross-modal information from text and images; Step 5: Perform sequence label prediction through a conditional random field to identify and classify aspect terms and sentiments.
2. The multi-level cross-modal alignment method based on contrastive learning according to claim 1, wherein The said Step 1 includes the following steps: Step 1-1. Generate positive and negative examples from a batch of (T s , V g ) input pairs. For the a-th pair in the batch, is the text representation, is the image representation. Assume that the positive example is the text and image representations from the same input pair, i.e., while the negative example is composed of representations from different input pairs, i.e., For each pair of inputs in the batch, we can obtain one positive example and N - 1 negative examples. The effect of contrastive learning is mainly affected by the number of negative examples, and the number of negative examples is positively correlated with the effect. The influence of a small number of mismatched pairs that may appear in the positive examples is negligible; Step 1-2: For each example Two different multi-layer perceptrons are respectively adopted, and each MLP contains one hidden layer, which are respectively applied to and to obtain the projected text representation and image representation It is found that this MLP projection can help the encoder learn better representations; Step 1-3: Maximize the similarity of positive examples and minimize the similarity of negative examples by minimizing two contrastive loss functions, which are the image-to-text contrastive loss and the text-to-image contrastive loss respectively. The image-to-text contrastive loss function is defined as follows for the i-th positive projection pair in the batch: The text-to-image contrastive loss is defined as follows for the i-th positive projection pair: Finally, sum the two losses of all positive projection pairs in the batch to obtain the total loss function: Among them, λ c ∈ [0, 1] is a hyperparameter. By minimizing the loss function, the representations of the text encoder and the image encoder will be more consistent.
3. The multi-level cross-modal alignment method based on contrastive learning according to claim 1, wherein The said Step 2 includes the following steps: Step 2-1: For each pair of text and image representations, generate positive sample pairs and negative sample pairs; Step 2-2: Use different multi-layer perceptrons to map the text and image representations to a common feature space; Step 2-3: Maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs by minimizing the image-to-text contrastive loss and the text-to-image contrastive loss; Step 2-4: Dynamically adjust the weight parameters in the contrastive loss function to adaptively optimize cross-modal consistency according to the similarity between text and images.
4. The multi-level cross-modal alignment method based on contrastive learning according to claim 1, wherein The said Step 3 includes the following steps: Step 3-1: Use the text representation as the query, the image representation as the key and value, and calculate the attention weights between the text and the image through the multi-head cross-attention mechanism, where the text representation while the image representation is The formula is as follows: Among them are the parameter matrices for query, key, and value respectively. Through the multi-head cross-attention mechanism, the final visual perception text representation can be obtained: The calculation is as follows: Among them, W p is the weight matrix of the multi-head cross-attention; Step 3-2: Optimize the feature fusion process through layer normalization and residual connections to enhance the model's ability to capture local features; Step 3-3: Use a conditional random field to perform sequence labeling on the fused text representation to identify aspect terms and their sentiment polarities in the text, and the calculation is as follows: where A is the state transition matrix, and its element A i,j represents the transition score from label i to label j, is the weight vector, and the loss function uses negative log probability, which is specifically expressed as:
5. The multi-level cross-modal alignment method based on contrastive learning according to claim 1, wherein The multi-task learning framework in the said Step 4 includes two sub-tasks to integrate cross-modal information from text and images and perform sequence label prediction through a conditional random field. The two sub-tasks are respectively: ①Aspect term extraction, responsible for identifying all aspect terms from the text; ②Aspect sentiment classification, predicting the sentiment polarity associated with each identified aspect term; Through a shared representation layer and a feature fusion mechanism, information interaction and collaborative optimization between the two sub-tasks are achieved.
6. The multi-level cross-modal alignment method based on contrastive learning according to any one of claims 1-5, characterized in that, It also includes the following steps: Step 6-1: During the training process, use unaligned large-scale text data to pre-train the text encoder to improve the model's ability to understand text semantics; Step 6-2: In the fine-tuning stage, jointly fine-tune the pre-trained text encoder and image encoder to learn cross-modal feature representations; Step 6-3: Through contrastive learning and cross-modal attention mechanisms, enhance the model's ability to recognize aspect terms and sentiment polarities in multi-modal data; Step 6-4: In the testing stage, use the trained model to perform aspect term extraction and sentiment polarity analysis on new multi-modal data, and output aspect terms and their corresponding sentiment labels.
Citation Information
Cited By
Intelligent agent self-adaptive decision-making method and device based on multi-modal semantic alignment
CN120542470A
Training reasoning method based on voice-text-image multi-mode contrast learning
CN121562829A
CAD drawing information extraction method for constructing multiple agents based on multi-modal large model
CN121765486A