Prompt-based double-layer cross-modal distillation learning multi-modal aspect sentiment analysis method and system, electronic equipment and readable storage medium
Through the hint-based double-layer cross-modal distillation learning method, the problem of insufficient adaptability of multimodal sentiment analysis in noise and multi-scenarios is solved, and the emotion classification effect with high accuracy and robustness is achieved.
Patent Information
- Application Number
- CN202510068400.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-16
AI Technical Summary
The existing multimodal sentiment analysis methods perform poorly when dealing with noise scenarios, are difficult to adapt in multiple scenarios, and have shortcomings in the alignment of graphics and text information.
A two-layer cross-modal distillation learning method based on hints is proposed. By generating text templates related to the aspect word as prompts, combining pre-trained language model and multi-modal sentiment analysis model, a two-layer aspect characterization construction and a gate mechanism filtering noise is carried out to achieve cross-modal distillation learning.
It improves the accuracy of emotion classification, shows strong robustness and adaptability in noise and multiple scenarios, and solves the problem of poor robustness of existing methods in scenarios without visual or visual interference.
Smart Images

Figure CN119990101A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multimodal aspect sentiment analysis method, system, electronic device, and readable storage medium based on prompt-based double-layer cross-modal distillation learning, and belongs to the technical field of multimodal sentiment analysis. Background Art
[0002] In the digital society, social media platforms have become an important channel for the public to express emotions and share opinions, generating a large amount of multimodal data (including text, pictures, etc.). However, traditional pure text aspect-level sentiment analysis can no longer meet the increasingly complex information needs. Multimodal aspect-based sentiment analysis (MABSA) significantly improves the accuracy and comprehensiveness of sentiment analysis by integrating text and image information, making up for the shortcomings of a single modality in fine-grained sentiment analysis. Although multimodal sentiment analysis can cope with multimodal social media environments such as pictures and texts, the interference of noisy data in real social media environments is inevitable. Existing multimodal sentiment analysis methods are usually based on idealized data sets and are difficult to adapt to complex noise scenarios. Therefore, how to make accurate and robust sentiment judgments under noise interference has become an important research challenge.
[0003] Early research focused on designing effective models to capture the relationship between aspect information and sentence context in text. For example, Li and Song et al. applied graph convolutional networks to extract syntactic dependency features from parsed syntactic dependency trees, and obtained aspect-aware syntactic representations through attention mechanisms and aspect masks. Recently, MABSA research has focused on improving the alignment between vision and text. For example, CMMT proposed multi-aspect and sentiment detection tasks for cross-modal interactive learning. HIMT and VLP-MABSA use object detection methods to extract objects from visual information and exclude background noise. However, most MABSA methods perform poorly when dealing with visual noise. For this, Li et al. reconstructed missing semantic features through adaptive feature optimization and knowledge integration self-distillation methods; Han et al. designed a dynamic learner in which a context sliding window with adjustable size is set to supplement missing modal features. However, these methods can only solve part of the visual noise problem. At present, there are still some shortcomings in multimodal aspect sentiment analysis methods: 1. These models usually infer aspect sentiment polarity in text-only or multimodal environments, and are difficult to adapt to mixed noise scenes. 2. Due to the modal differences and information redundancy between images and text, it is difficult to accurately align aspect words with image and text information.
[0004] In response to the above problems, this paper proposes a multimodal aspect sentiment analysis method based on double-layer cross-modal distillation learning with hints, aiming to address the challenges of multi-scene adaptability and noise robustness in multimodal aspect sentiment analysis (MABSA). Cross-modal visual distillation is a potential way to integrate aspect-specific visual details to improve the performance of MABSA in multiple scenes, thereby achieving multi-scene adaptability. By learning cross-modal distillation with text-image aligned data, this strategy effectively solves the challenges of multi-scene adaptability and robust representation, and can also be extended to text-only scenes. However, it is often challenging to train MABSA models and perform cross-modal distillation from scratch to achieve text-image semantic alignment. Since these two tasks are closely related, they may negatively affect each other, making it difficult to learn robust cross-modal distillation capabilities. Based on the latest MABSA model, the hint-guided distillation method can focus on cross-modal visual distillation for multiple scenes. This approach ensures that the existing text-image alignment in MABSA is not affected, while cross-modal visual distillation learning is performed through hints. Inspired by this, we propose a cue-based two-layer distillation learning strategy based on the latest MABSA model to learn the ability to extract valuable visual information in noisy or even imageless scenes through cross-modal distillation.
[0005] In addition, in the current research work on fine-grained sentiment analysis, the core of the task is to determine the sentiment polarity of aspect words in image and text information. Effectively focusing on aspect information can help the model understand the context more accurately and improve the accuracy of sentiment classification. Based on this, we constructed a two-layer aspect representation based on hints to enhance the model's attention to aspect semantics.
[0006] In summary, the present invention proposes a multimodal aspect sentiment analysis method based on prompt-based two-layer cross-modal distillation learning. Summary of the invention
[0007] The technical problem to be solved by the present invention is: the present invention provides a multimodal sentiment analysis method, system, electronic device, and readable storage medium based on prompt-based double-layer cross-modal distillation learning, so as to solve the problems of poor multi-scene adaptability and noise robustness in multimodal sentiment analysis. The present invention improves the accuracy of sentiment classification.
[0008] The technical solution of the present invention is: a multimodal aspect sentiment analysis method based on prompt-based double-layer cross-modal distillation learning, the method comprising:
[0009] Step 1, prompt generation: construct prompts through text templates with aspect words as the core, and regard the text templates as the extension of aspect words to form prompts;
[0010] Step 2: Embedding text and visual representations: Use a pre-trained language model (such as BERT) to generate embedded representations for the input text and prompts; use an existing multimodal aspect sentiment analysis model (DQPSA) to extract visual aspect representations for the input visual information and prompts;
[0011] Step 3: Construction of shallow and deep aspect representations: Calculate shallow aspect representations of text and prompts using shallow attention mechanism; Use multi-layer Transformer to further extract deep aspect representations;
[0012] Step 4, gating mechanism to filter noise: by filtering out several tokens with the highest correlation between visual aspect representation, shallow aspect representation, deep aspect representation and CLS tag, and integrating them, the noise is filtered out, and the visual gating representation, shallow gating representation and deep gating representation are respectively selected;
[0013] Step 5, two-layer cross-modal distillation: calculate the shallow distillation loss for the shallow gated representation and the visual gated representation; calculate the deep distillation loss for the deep gated representation and the visual gated representation;
[0014] Step 6, model training and optimization: By jointly optimizing classification loss, shallow distillation loss and deep distillation loss, the model's sentiment classification ability in multimodal scenarios is trained.
[0015] Furthermore, in the Step 1, the constructing of the prompt includes: constructing the prompt Q using a predefined text template containing a fixed context structure; used to enhance the model's attention to the semantics related to aspect words; the input sentence S and the generated prompt Q will serve as the basis for subsequent representation.
[0016] Furthermore, the Step 2 includes:
[0017] Step 2.1. For sentence S and prompt Q, get sentence embedding T through the pre-trained language model BERT S and prompts embedded in T Q ; Sentence Embedding T S and prompts embedded in T Q It is expressed as:
[0018] T S =BERT emb (S) (1)
[0019] T Q =BERT emb (Q) (2)
[0020] Where T S ∈R n×d , T Q ∈Rm×d , n represents the sentence sequence length, m represents the prompt sequence length, and d is the dimension of the BERT embedding layer;
[0021] Step 2.2: Use the pre-trained PDQ module in the DQPSA model to extract the visual representation H by inputting the noisy image V and the prompt Q V ; Visual Representation H V It is expressed as:
[0022] H V =PDQ_module(V,Q) (3)
[0023] Among them, H V ∈R m×d Representation of visual aspects extracted by the representation model.
[0024] Furthermore, the Step 3 includes:
[0025] Step 3.1, Obtain shallow aspect representation: For a given sentence embedding T S and prompts embedded in T Q , first calculate the attention score α through the interactive attention mechanism; the calculation formula of the attention score α is:
[0026] α=softmax(T Q W1T S T ) (4)
[0027] where W1∈R d×d is the shallow learning weight matrix, α∈R m×k ;
[0028] Based on this attention score, we obtain the cue-guided shallow aspect representation H L , shallow aspect representation H L It is expressed as:
[0029] H L =relu(α)T S (5)
[0030] Among them, H L ∈R m×d It is a superficial aspect representation;
[0031] Step 3.2, obtain deep aspect representation: shallow aspect representation H through multi-layer Transformer structure L Further extraction is performed and the deep aspect representation is calculated; the deep aspect representation is:
[0032] H H =TransformweEncoder(HL ) (6)
[0033] Among them, H H ∈R m×d Represents the deep aspects.
[0034] Furthermore, the Step 4 includes:
[0035] Step 4.1. Represent H in terms of vision V , by calculating the cosine similarity between each token and the CLS tag, the k tokens with the highest correlation are selected and integrated to filter out the noise and screen out the key representations. Among them, the CLS tag is the first feature in the visual representation, that is, the global representation;
[0036] The similarity score calculation formula is as follows:
[0037]
[0038] Among them, γ i Indicates the i-th token With CLS tag The similarity score between
[0039] Select the top k tokens in the similarity score ranking as the new visual gating representation
[0040]
[0041] in, k is set as the gating coefficient and participates in training as a hyperparameter;
[0042] Step 4.2: Characterize H in terms of shallow aspects L and deep aspect representation H H , the gating mechanism is also applied to perform Step 4.1 to improve the noise robustness of the aspect representation and generate a shallow gated representation after the gating mechanism is processed and deep gating representation
[0043] Furthermore, the Step 5 includes:
[0044] Step 5.1: Deep gated representation after gating mechanism processing Shallow gating representation and visual gating representation Converted into probability distributions: shallow gated representation Deep gated characterization Visual gating representation The probability distribution PL , P H , P V It is expressed as:
[0045]
[0046]
[0047] Where τ is a temperature hyperparameter that controls the smoothness of softmax;
[0048] Step 5.2, calculate the shallow distillation loss and deep distillation loss to measure the transfer learning effect of the model;
[0049] Shallow distillation loss L LV It is used to measure the KL divergence between the shallow gated representation and the visual representation after the gated mechanism, and the shallow distillation loss L LV for:
[0050]
[0051] Deep distillation loss L HV It is used to measure the KL divergence between the deep gated representation and the visual aspect representation after the gating mechanism, and the deep distillation loss L HV for:
[0052]
[0053] Where P V j , P L j , P H j Respectively represent P L , P H , P V The jth element in .
[0054] Furthermore, the Step 6 includes:
[0055] In the training phase, by minimizing the shallow distillation loss and the deep distillation loss L LV , L HV , so that the model can fully learn the visual representation; then the deep and shallow representations of the model are combined and sent to the classifier for sentiment prediction; the classifier processing process is as follows:
[0056] First, the deep and shallow aspect representations are added and pooled to obtain a fused representation:
[0057]
[0058] The fused representation is then processed through the relu activation function, and the logarithmic probability distribution is calculated through the log_softmax function:
[0059] Y=log_softmax(w1relu(w0η)) (13)
[0060] in Represents the addition of two matrices, η∈R 1×d represents the fusion representation finally fed into the classifier, Y∈R num Represents the prediction result, num represents the number of classifications, w0∈R d×0.5d , w1∈R 0.5d×num is the learned weight matrix;
[0061] The cross entropy loss between the predicted result Y and the true label M is used as the prediction loss:
[0062]
[0063] M i represents the i-th true label, Y i Represents the i-th prediction result;
[0064] Finally, the total loss is as follows:
[0065] L=λ pre L pre +λ LV L LV +λ HV L HV (15)
[0066] The three hyperparameters λ are pre , LV , HV They are used to control the contribution rates of the three losses respectively;
[0067] During the training phase, the model accepts text, visual information, and prompts as input;
[0068] In the inference phase, the model only needs text and prompts to complete the sentiment classification task.
[0069] The present invention also provides a multimodal aspect sentiment analysis system based on prompt-based two-layer cross-modal distillation learning. The system includes: a module for executing the above-mentioned multimodal aspect sentiment analysis method based on prompt-based two-layer cross-modal distillation learning.
[0070] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the multimodal aspect sentiment analysis method based on prompt-based two-layer cross-modal distillation learning is implemented.
[0071] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multimodal aspect sentiment analysis method based on the above-mentioned prompt-based two-layer cross-modal distillation learning.
[0072] The beneficial effects of the present invention are:
[0073] 1. The multimodal aspect sentiment analysis method based on prompt-based two-layer cross-modal distillation learning proposed in this paper performs well in multimodal sentiment analysis;
[0074] 2. The method of the present invention shows strong noise robustness and scene adaptability in scenes such as visual noise;
[0075] 3. When only text input is used during reasoning, the model first guides the focus on aspect word information through prompts, then effectively obtains the multimodal information of the existing MABSA model through distillation learning, and filters redundant noise through the gating mechanism. These strategies are combined to make the method of the present invention have SOTA high performance;
[0076] 4. The method of the present invention solves the problem that the existing methods are poorly robust or even inapplicable in scenarios with no vision or visual interference;
[0077] 5. The present invention demonstrates extremely high accuracy and robustness in sentiment analysis tasks involving noisy vision. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 It is a framework diagram of the method in the present invention. DETAILED DESCRIPTION
[0079] Example 1: The following method proposed in this example is implemented on three multimodal sentiment analysis datasets (Twitter2015, Twitter2017 and MASAD);
[0080] The present invention uses three benchmark datasets, namely Twitter2015, Twitter2017 and MASAD, to perform sentiment analysis tasks. The two Twitter datasets collect user posts from 2014-2015 and 2016-2017, respectively. MASAD is a newly released large-scale multimodal dataset that covers 57 aspects in seven fields: food, goods, architecture, animals, humans, plants and scenery. In the three datasets, each sample contains an aspect word, a sentence and a corresponding image. In the three datasets, the sentiment polarity of aspect words in the Twitter dataset is divided into three categories: positive, neutral and negative, while the sentiment polarity of aspect words in the MASAD dataset is only divided into two categories: positive and negative. The statistical information of these datasets is shown in Table 1.
[0081] Table 1 Dataset statistics
[0082]
[0083] like Figure 1 As shown, a multimodal aspect sentiment analysis method based on prompt-based two-layer cross-modal distillation learning, the method comprising:
[0084] Step 1, prompt generation: construct prompts through text templates with aspect words as the core, and regard the text templates as the extension of aspect words to form prompts;
[0085] Furthermore, in the Step 1, constructing the prompt includes: using a predefined text template "sentiment of Aspect is [positive neutral negative]" containing a fixed context structure, where Aspect represents an aspect word, replacing Aspect with the aspect word in each sample in the data, thereby obtaining a prompt Q corresponding to each sample, which is used to enhance the model's attention to the semantics related to the aspect word; the input sentence S and the generated prompt Q will serve as the basis for subsequent representation.
[0086] Step 2: Embedding text and visual representations: Use a pre-trained language model (such as BERT) to generate embedded representations for the input text and prompts; use an existing multimodal aspect sentiment analysis model (DQPSA) to extract visual aspect representations for the input visual information and prompts;
[0087] Furthermore, the Step 2 includes:
[0088] Step 2.1. For sentence S and prompt Q, get sentence embedding T through the pre-trained language model BERT S and prompts embedded in T Q ; Sentence Embedding T S and prompts embedded in TQ It is expressed as:
[0089] T S =BERT emb (S) (1)
[0090] T Q =BERT emb (Q) (2)
[0091] Where T S ∈R n×d , T Q ∈R m×d , n represents the sentence sequence length, m represents the prompt sequence length, and d is the dimension of the BERT embedding layer;
[0092] Step 2.2: Use the pre-trained PDQ module in the DQPSA model to extract the visual representation H by inputting the noisy image V and the prompt Q V ; Visual Representation H V It is expressed as:
[0093] H V =PDQ_module(V,Q) (3)
[0094] Among them, H V ∈R m×d Representation of visual aspects extracted by the representation model.
[0095] Step 3: Construction of shallow and deep aspect representations: Calculate shallow aspect representations of text and prompts using shallow attention mechanism; Use multi-layer Transformer to further extract deep aspect representations;
[0096] Furthermore, the Step 3 includes:
[0097] Step 3.1, Obtain shallow aspect representation: For a given sentence embedding T S and prompts embedded in T Q , first calculate the attention score α through the interactive attention mechanism; the calculation formula of the attention score α is:
[0098] α=softmax(T Q W1T S T ) (4)
[0099] where W1∈R d×d is the shallow learning weight matrix, α∈R m×k ;
[0100] Based on this attention score, we obtain the cue-guided shallow aspect representation H L , shallow aspect representation HL It is expressed as:
[0101] H L =relu(α)T S (5)
[0102] Among them, H L ∈R m×d It is a superficial aspect representation;
[0103] Step 3.2, obtain deep aspect representation: shallow aspect representation H through multi-layer Transformer structure L Further extraction is performed and the deep aspect representation is calculated; the deep aspect representation is:
[0104] H H =TransformerEncoder(H L ) (6)
[0105] Among them, H H ∈R m×d Represents the deep aspects.
[0106] Step 4, gating mechanism to filter noise: by filtering out several tokens with the highest correlation between visual aspect representation, shallow aspect representation, deep aspect representation and CLS tag, and integrating them, the noise is filtered out, and the visual gating representation, shallow gating representation and deep gating representation are respectively selected;
[0107] Furthermore, the Step 4 includes:
[0108] Step 4.1. Represent H in terms of vision V , by calculating the cosine similarity between each token and the CLS tag, the k tokens with the highest correlation are selected and integrated to filter out the noise and screen out the key representations. Among them, the CLS tag is the first feature in the visual representation, that is, the global representation;
[0109] The similarity score calculation formula is as follows:
[0110]
[0111] Among them, γ i Indicates the i-th token With CLS tag The similarity score between
[0112] Select the top k tokens in the similarity score ranking as the new visual gating representation
[0113]
[0114] in, k is set as the gating coefficient and participates in training as a hyperparameter;
[0115] Step 4.2: Characterize H in terms of shallow aspects L and deep aspect representation H H , the gating mechanism is also applied to perform Step 4.1 to improve the noise robustness of the aspect representation and generate a shallow gated representation after the gating mechanism is processed and deep gating representation
[0116] Step 5, two-layer cross-modal distillation: calculate the shallow distillation loss for the shallow gated representation and the visual gated representation; calculate the deep distillation loss for the deep gated representation and the visual gated representation;
[0117] Furthermore, the Step 5 includes:
[0118] Step 5.1: Deep gated representation after gating mechanism processing Shallow gating representation and visual gating representation Converted into probability distributions: shallow gated representation Deep gated characterization Visual gating representation The probability distribution P L , P H , P V It is expressed as:
[0119]
[0120] Where τ is a temperature hyperparameter that controls the smoothness of softmax;
[0121] Step 5.2, calculate the shallow distillation loss and deep distillation loss to measure the transfer learning effect of the model;
[0122] Shallow distillation loss L LV It is used to measure the KL divergence between the shallow gated representation and the visual representation after the gated mechanism, and the shallow distillation loss L LV for:
[0123]
[0124] Deep distillation loss L HV It is used to measure the KL divergence between the deep gated representation and the visual aspect representation after the gating mechanism, and the deep distillation loss L HV for:
[0125]
[0126] Where P V j , P L j , P H j Respectively represent P L , P H , P V The jth element in .
[0127] Step 6, model training and optimization: By jointly optimizing classification loss, shallow distillation loss and deep distillation loss, the model's sentiment classification ability in multimodal scenarios is trained.
[0128] Furthermore, the Step 6 includes:
[0129] In the training phase, by minimizing the shallow distillation loss and the deep distillation loss L LV , L HV , so that the model can fully learn the visual representation; then the deep and shallow representations of the model are combined and sent to the classifier for sentiment prediction; the classifier processing process is as follows:
[0130] First, the deep and shallow aspect representations are added and pooled to obtain a fused representation:
[0131]
[0132] The fused representation is then processed through the relu activation function, and the logarithmic probability distribution is calculated through the log_softmax function:
[0133] Y=log_softmax(w1relu(w0η)) (13)
[0134] in Represents the addition of two matrices, η∈R 1×d represents the fusion representation finally fed into the classifier, Y∈R num Represents the prediction result, num represents the number of classifications, w0∈R d×0.5d , w1∈R 0.5d×num is the learned weight matrix;
[0135] The cross entropy loss between the predicted result Y and the true label M is used as the prediction loss:
[0136]
[0137] M i represents the i-th true label, Y i Represents the i-th prediction result;
[0138] Finally, the total loss is as follows:
[0139] L=λ pre L pre +λ LV L LV +λ HV L HV (15)
[0140] The three hyperparameters λ are pre , LV , HV They are used to control the contribution rates of the three losses respectively;
[0141] During the training phase, the model accepts text, visual information, and prompts as input;
[0142] In the inference phase, the model only needs text and prompts to complete the sentiment classification task.
[0143] The present invention also provides a multimodal aspect sentiment analysis system based on prompt-based two-layer cross-modal distillation learning, the system comprising:
[0144] a prompt generation module, for constructing prompts through text templates with aspect words as the core, and treating the text templates as extensions of aspect words to form prompts;
[0145] A representation extraction module is used to generate embedded representations for the input text and prompts using a pre-trained language model (such as BERT); and to extract visual aspect representations for the input visual information and prompts using an existing multimodal aspect sentiment analysis model (DQPSA);
[0146] Shallow and deep aspect representation building blocks for computing shallow aspect representations of text and prompts using a shallow attention mechanism; using a multi-layer Transformer to further extract deep aspect representations;
[0147] The gated representation filtering module is used to filter out the visual aspect representation, shallow aspect representation, deep aspect representation and several tokens with the highest correlation with the CLS mark respectively, and integrate them to filter out noise, and respectively filter out the visual gated representation, shallow gated representation and deep gated representation;
[0148] A distillation loss calculation module is used to calculate the shallow distillation loss for the shallow gated representation and the visual gated representation; and to calculate the deep distillation loss for the deep gated representation and the visual gated representation;
[0149] The sentiment classification module is used to train the model's sentiment classification ability in multimodal scenarios by jointly optimizing classification loss, shallow distillation loss, and deep distillation loss.
[0150] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the multimodal aspect sentiment analysis method based on prompt-based two-layer cross-modal distillation learning is implemented.
[0151] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multimodal aspect sentiment analysis method based on the above-mentioned prompt-based two-layer cross-modal distillation learning.
[0152] In order to verify the effectiveness of the model proposed in this invention, the model of the method of this invention is compared with other latest sentiment analysis models, covering three data sets, including plain text, image-text and other scenarios, as follows:
[0153] In the plain text scenario:
[0154] ATAE-LSTM uses aspect-oriented attention mechanism to extract important information from sentences.
[0155] IAN combines the long short-term memory (LSTM) network and the attention mechanism to model the interaction between the target and the context, accurately capture key information and predict sentiment tendency.
[0156] TNet uses convolutional neural networks (CNNs) to extract features and generate aspect-specific sentence representations while preserving the original contextual information.
[0157] MGAN utilizes multi-granularity attention networks to understand aspects.
[0158] BERT captures the interaction between aspects and text using the pre-trained BERT model.
[0159] DualGCN is a dual graph convolutional network model that considers the complementarity between syntactic structure and semantic association.
[0160] In the text-image multimodal scenario:
[0161] MIMN explores attention-based interactions between aspects, sentences, and related images via a multi-hop memory network.
[0162] TomBERT learns aspect-sensitive text features, matches aspect-image pairs to obtain visual features, and uses a self-attention strategy to capture multimodal interactions.
[0163] MMAP reveals hidden relations among sentences, aspects, and images through multimodal interaction layers and adversarial training.
[0164] AMIFN improves the accuracy of fine-grained sentiment analysis by integrating attention mechanism and graph convolutional network to achieve multi-view interaction and fusion for specific aspects.
[0165] ESAFN divides a sentence into left and right contexts, exploits the attention mechanism to explore aspect-text and aspect-image interactions, and finally fuses these features through a bilinear layer for sentiment prediction.
[0166] Res-BERT combines the visual features extracted by ResNet and the hidden representations of BERT.
[0167] HIMT conducts aspect-text and aspect-image interactions and builds an auxiliary reconstruction module to eliminate the semantic differences between different modalities.
[0168] VLP-MABSA is a task-specific vision-language pre-training model that models aspect, viewpoint, and alignment by performing five specialized pre-training tasks.
[0169] KEF-TomBERT is a knowledge-enhanced framework that improves task performance by associating images with adjective-nouns.
[0170] Evaluation metrics: Accuracy (Acc) and macro-average F1 score (F1) are used as evaluation metrics. Higher Acc and F1 values indicate better performance.
[0171] The method of the present invention firstly conducts comparative experiments on two benchmark datasets, Twitter2015 and Twitter2017, in the scenarios without images and with images and texts combined, and the results are shown in Table 2. The present invention draws the following conclusions: 1) In the conditions without images and with images and texts combined, the method proposed in the present invention is superior to other SOTA methods; 2) In the pure text scenario, the method of the present invention has a significant improvement over all methods, which verifies the effectiveness of the method of the present invention in fine-grained emotion recognition; 3) In the two evaluation indicators of the two tasks, the method proposed in the present invention is also significantly superior to other methods, further proving the effectiveness of the method in sentiment analysis; 4) In the two datasets, the method proposed in the present invention is also superior to other MABSA methods based on pre-training, which shows the effectiveness of the distillation mechanism in sentiment analysis.
[0172] Experimental results on public datasets such as twitter2015 and twitter2017 show that the cue-based two-layer cross-modal distillation learning method exhibits extremely high accuracy and robustness in sentiment analysis tasks involving noisy vision.
[0173] Table 2 compares the experimental results on the twitter2015 and twitter2017 datasets.
[0174]
[0175]
[0176] Then, on the MASAD dataset, the present invention also conducted comparative tests in two scenarios. Since other models obtained experimental results in seven fields of the dataset, the present invention took the average results of each model for comparison, and the results are shown in Table 3. The present invention draws the following conclusions: 1) The method of the present invention shows strong robustness and adaptability in processing tasks with diverse data in different fields, which proves that the method of the present invention is capable of handling a variety of sentiment analysis scenarios; 2) Whether in pure text or graphic scenarios, the method of the present invention also shows significant advantages over most models in the dataset, which once again proves the effectiveness and robustness of the method of the present invention.
[0177] Table 3 shows the comparison of experimental results on the MASAD dataset
[0178]
[0179] In addition, in order to evaluate the robustness of the proposed method, the present invention conducted experiments in multiple scenes with visual noise, including visual occlusion noise (VON), visual interference noise (VIN), image-text mismatch noise (ITMN) and visual missing noise (VMN). The experimental results are shown in Table 4.
[0180] 1) In all three datasets, adding VIN and VMN leads to a significant drop in sentiment analysis performance, because VIN severely destroys visual information, which in turn has a negative impact on performance; 2) The sentiment analysis performance in the VMN scenario is the worst, probably due to the complete lack of visual information, which causes the model to lose the most information; 3) The introduction of VON has little impact on performance, demonstrating the robustness of the proposed method. This also shows that Gaussian white noise has less loss of visual information, because compared with other more destructive noises, the impact of additive visual noise is less severe.
[0181] Table 4 shows the comparison of experimental results under different visual noises
[0182]
[0183]
[0184] In order to further verify the effectiveness of the module proposed in the present invention, the present invention conducts ablation experiments on three datasets. The experimental results are shown in Table 5.
[0185] Table 5 shows the Ablation Study on three datasets
[0186]
[0187] w / o Prompt-Guided and w / o Gating represent the removal of the prompt-guided two-layer aspect-guided representation and gating mechanism from the proposed model. w / o L HV and w / o L LV Represents the removal of shallow distillation loss and deep distillation loss.
[0188] The present invention draws the following conclusions: 1) Removing the two distillation learning losses will lead to a significant drop in performance on the three datasets. This proves the effectiveness of the distillation mechanism in sentiment analysis, especially in improving the model's adaptability to complex scenarios; 2) After removing Prompt-Guided, the performance of the model in each dataset has declined, indicating that the aspect perception module is crucial to the model's information focusing ability in fine-grained sentiment analysis; 3) After removing Gating, the model performance also drops significantly, which proves the role of the gating mechanism in reducing the semantic gap between modalities and the impact of noise; 4) Removing L HV After that, the distillation learning ability of the model is incomplete and the model performance has declined, which proves that double-layer distillation learning can more effectively enhance the sentiment analysis ability of the model. In summary, the ablation experiment results fully verify the importance of each module and loss function in the fine-grained sentiment analysis task.
[0189] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.
Claims
1. A multimodal aspect sentiment analysis method based on two-layer cross-modal distillation learning based on prompts, characterized by: The method comprises: Step 1, prompt generation: construct prompts through text templates with aspect words as the core, and regard the text templates as the extension of aspect words to form prompts; Step 2: Embedding text and visual representations: Use the pre-trained language model to generate embedded representations for the input text and prompts; use the existing multimodal aspect sentiment analysis model to extract visual aspect representations for the input visual information and prompts; Step 3: Construction of shallow and deep aspect representations: Calculate shallow aspect representations of text and prompts using shallow attention mechanism; Use multi-layer Transformer to further extract deep aspect representations; Step 4, gating mechanism to filter noise: by filtering out several tokens with the highest correlation between visual aspect representation, shallow aspect representation, deep aspect representation and CLS tag, and integrating them, the noise is filtered out, and the visual gating representation, shallow gating representation and deep gating representation are respectively selected; Step 5, two-layer cross-modal distillation: calculate the shallow distillation loss for the shallow gated representation and the visual gated representation; calculate the deep distillation loss for the deep gated representation and the visual gated representation; Step 6, model training and optimization: By jointly optimizing classification loss, shallow distillation loss and deep distillation loss, the model's sentiment classification ability in multimodal scenarios is trained.
2. The multimodal aspect sentiment analysis method based on prompt-based two-layer cross-modal distillation learning according to claim 1, characterized in that: In the Step 1, the construction of the prompt includes: using a predefined text template containing a fixed context structure to construct a prompt Q; used to enhance the model's attention to the semantics related to aspect words; the input sentence S and the generated prompt Q will serve as the basis for subsequent representation.
3. The multimodal aspect sentiment analysis method based on prompt-based two-layer cross-modal distillation learning according to claim 1, characterized in that: The Step 2 includes: Step 2.
1. For sentence S and prompt Q, get sentence embedding T through the pre-trained language model BERT S and prompts embedded in T Q ; Sentence Embedding T S and prompts embedded in T Q It is expressed as: T S =BERT emb (S) (1) T Q =BERT emb (Q) (2) where t S ∈R n×d , T Q ∈R m×d , n represents the sentence sequence length, m represents the prompt sequence length, and d is the dimension of the BERT embedding layer; Step 2.2: Use the pre-trained PDQ module in the DQPSA model to extract the visual representation H by inputting the noisy image V and the prompt Q V ; Visual Representation H V It is expressed as: H V =PDW_module(V,Q) (3) Among them, H V ∈R m×d Representation of visual aspects extracted by the representation model.
4. The multimodal aspect sentiment analysis method based on prompt-based two-layer cross-modal distillation learning according to claim 1, characterized in that: The Step 3 includes: Step 3.1, Obtain shallow aspect representation: For a given sentence embedding T S and prompts embedded in T Q , first calculate the attention score α through the interactive attention mechanism; the calculation formula of the attention score α is: α=softmax(T Q W1T S T ) (4) where W1∈R d×d is the shallow learning weight matrix, α∈R m×k ; Based on this attention score, we obtain the cue-guided shallow aspect representation H L , shallow aspect representation H L It is expressed as: H L =relu(α)T S (5) Among them, H L ∈R m×d It is a superficial aspect representation; Step 3.2, obtain deep aspect representation: shallow aspect representation H through multi-layer Transformer structure L Further extraction is performed and the deep aspect representation is calculated; the deep aspect representation is: H H =TransfoemerEncoder(H L ) (6) Among them, H H ∈R m×d Represents the deep aspects.
5. The multimodal aspect sentiment analysis method based on prompt-based two-layer cross-modal distillation learning according to claim 1, characterized in that: The Step 4 includes: Step 4.
1. Represent H in terms of vision V , by calculating the cosine similarity between each token and the CLS tag, the k tokens with the highest correlation are selected and integrated to filter out the noise and screen out the key representations. Among them, the CLS tag is the first feature in the visual representation, that is, the global representation; The similarity score calculation formula is as follows: Among them, γ i Indicates the i-th token With CLS tag The similarity score between Select the top k tokens in the similarity score ranking as the new visual gating representation in, k is set as the gating coefficient and participates in training as a hyperparameter; Step 4.2: Characterize H in terms of shallow aspects L and deep aspect representation H H , and the gating mechanism is also applied to execute Step 4.1 to generate the shallow gating representation after the gating mechanism is processed and deep gating representation 6. The multimodal aspect sentiment analysis method based on prompt-based two-layer cross-modal distillation learning according to claim 1, characterized in that: The Step 5 includes: Step 5.1: Deep gated representation after gating mechanism processing Shallow gating representation and visual gating representation Converted into probability distributions: shallow gated representation Deep gated characterization Visual gating representation The probability distribution P L , P H , P V It is expressed as: Where τ is a temperature hyperparameter that controls the smoothness of softmax; Step 5.2, calculate the shallow distillation loss and deep distillation loss; Shallow distillation loss L LV It is used to measure the KL divergence between the shallow gated representation and the visual representation after the gated mechanism, and the shallow distillation loss L LV for: Deep distillation loss L HV It is used to measure the KL divergence between the deep gated representation and the visual aspect representation after the gating mechanism, and the deep distillation loss L HV for: Where P V j , P L j , P H j Respectively represent P L , P H , P V The jth element in .
7. The multimodal aspect sentiment analysis method based on prompt-based two-layer cross-modal distillation learning according to claim 1, characterized in that: The Step 6 includes: In the training phase, by minimizing the shallow distillation loss and the deep distillation loss L LV , L HV , so that the model can fully learn the visual representation; then the deep and shallow representations of the model are combined and sent to the classifier for sentiment prediction; the classifier processing process is as follows: First, the deep and shallow aspect representations are added and pooled to obtain a fused representation: The fused representation is then processed through the relu activation function, and the logarithmic probability distribution is calculated through the log_softmax function: Y=log_softmax(w1relu(w0η)) (13) in Represents the addition of two matrices, η∈R 1×d represents the fusion representation finally fed into the classifier, Y∈R num Represents the prediction result, num represents the number of classifications, w0∈R d×0.5d , w1∈R 0.5d×num is the learned weight matrix; The cross entropy loss between the predicted result Y and the true label M is used as the prediction loss: M i represents the i-th true label, Y i Represents the i-th prediction result; Finally, the total loss is as follows: L=λ pre L pre +λ LV L LV +λ HV L HV (15) The three hyperparameters λ are pre , LV , HV They are used to control the contribution rates of the three losses respectively; During the training phase, the model accepts text, visual information, and prompts as input; In the inference phase, the model only needs text and prompts to complete the sentiment classification task.
8. A multimodal aspect sentiment analysis system based on two-layer cross-modal distillation learning based on prompts, characterized by: The system includes: a module for executing the multimodal aspect sentiment analysis method based on prompt-based two-layer cross-modal distillation learning as described in any one of claims 1 to 7.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the multimodal aspect sentiment analysis method based on prompt-based two-layer cross-modal distillation learning as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the multimodal aspect sentiment analysis method based on prompt-based two-layer cross-modal distillation learning as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Aspect-level sentiment analysis method based on knowledge distillation framework
CN116384373A
Multi-modal aspect level sentiment analysis method for multi-level fusion image and text
CN117708642A
Multi-modal emotion recognition method and device
CN119167303A
Cited By
Text-guided knowledge distillation method based on attribute-driven fusion
CN120494040A
Multimodal sentiment analysis method and device based on selective alignment of teacher-guided distillation and text query
CN122065290A