Multi-modal aspect-emotion pair extraction method and system for emotion reason knowledge enhancement
By generating emotional cause knowledge in the multimodal aspect-emotional pair extraction task and assisting the extraction-emotional pair, the problem of insufficient context information in the multimodal data is solved, and more accurate extraction effects and richer context understanding are achieved.
Patent Information
- Application Number
- CN202411968911.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-30
AI Technical Summary
The existing multimodal aspect-emotional pair extraction method is difficult to learn the representation of modal alignment when facing multimodal data, and the text-text matching pattern has problems of insufficient context information and external knowledge redundancy in the field of sentiment analysis.
A multimodal aspect-emotional pair extraction method for enhancing emotional cause knowledge is proposed. Through the collaborative work of manual annotation and large language model, emotional reason knowledge is generated for each graphic and text data pair, and the aspects-emotional pair are obtained with its assistance. This method converts the picture-text matching pattern into the text-text matching pattern, uses a large language model to extract language knowledge that contains the description of the picture content and the explanation of emotional reasons, and enriches the context of the text content.
Modal fusion without semantic differences improves the accuracy of multimodal aspect-emotional extraction tasks, provides a richer context, and enhances the model's ability to understand emotional causes.
Smart Images

Figure CN120067743A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly relates to a method and system for extracting multimodal aspect-sentiment pairs enhanced by sentiment reason knowledge. Background Art
[0002] In recent years, the wide popularity of social media has enabled people to no longer be restricted to expressing their viewpoints and attitudes only in words, and often supplemented with other forms of content such as pictures, audio, and video. With the large amount of unstructured content including pictures and text generated by network users on social media, it has attracted extensive attention from researchers in the field of fine-grained sentiment analysis to the task of multimodal aspect-sentiment pair extraction (MASPE). In the content with both pictures and texts, the text information has inherent characteristics such as concise text content and informal writing style. These unique characteristics pose challenges to the traditional joint extraction of aspect-sentiment pairs. In order to make full use of the features of multiple modalities and improve the performance of the joint extraction task of aspect-sentiment pairs, a large number of research works have tried to use strategies such as cross-attention mechanisms and their various variants or maximizing the mutual information between modalities to implicitly align text and picture features. However, this series of methods focusing on the picture-text matching mode (I+T) have two limitations: First, there are biases in the feature distributions of different modalities, making it difficult for the model to learn the aligned representations between modalities when facing multiple modalities; Second, the picture feature extractors used in existing methods are all trained on large-scale datasets such as ImageNet and COCO. The labels in these datasets are mainly composed of specific nouns related to the task, rather than aspect targets related to the sentiment analysis task. There is an obvious deviation between the two types of datasets. Due to the limitations caused by the above two biases, some multimodal fusion methods may not be as effective as the current best language models that only focus on text.
[0003] Essentially, the MASPE task is a multi-modal task, and the contribution degrees of text and images to the task are not equivalent. When images cannot provide more interpretable information for understanding text semantics, image information can even be discarded or ignored. In addition, research work has proven that, on the basis of the original text, the performance of the named entity recognition task can be significantly improved by introducing document-level context. Therefore, some researchers have attempted to solve the multi-modal named entity recognition task by using the text-text matching mode (T+T). In this type of method, images will be converted into text representations, which can be achieved by techniques such as "Image Caption" and handwritten digit recognition. Obviously, due to the absence of the problem of feature distribution differences, modeling the attention mechanism between texts is superior to modeling the cross-modal attention mechanism. Considering that aspect extraction and opinion extraction in fine-grained sentiment analysis are similar to the named entity recognition task, and aspect-opinion relationship classification is similar to the entity relationship classification task, inspired by the above work, the present invention attempts to explore the possibility of converting the I+T mode of the existing multi-modal aspect-sentiment pair extraction method into the T+T mode to avoid the problem of modal semantic inconsistency. However, the existing text-text matching mode still has two potential defects in the field of sentiment analysis: (1) For methods that only rely on in-sample information, when the data content itself is relatively simple, there is a lack of sufficient sentiment-related context information, and additional external knowledge is often required to enhance the understanding of the text; (2) For methods that introduce external knowledge, the relevant knowledge retrieved from external knowledge bases (such as Wikipedia) is usually very redundant, and in some cases, the extended knowledge with low relevance may even mislead the model to understand the text information.
[0004] Recently, the field of large language models (LLMs) has been developing rapidly, with some interesting new discoveries and progress. On the one hand, research work on LLMs has shown that generative models have obvious deficiencies in sequence labeling tasks. On the other hand, LLMs have achieved surprising results in a variety of natural language processing tasks and multi-modal tasks. These LLMs with in-context learning capabilities can be regarded as encyclopedias that understand comprehensive world knowledge and can usually provide high-quality knowledge to assist in text content understanding. From this, a conjecture arises: Can the potential of LLMs in the MASPE task be activated by endowing them with reasonable heuristic methods? Summary of the Invention
[0005] To address the above technical problems, the present invention proposes a multi-modal aspect-sentiment pair extraction method and system enhanced with sentiment reason knowledge to solve the problem that the previous methods ignored the lack of sufficient context information in the text content by the multi-modal graphic and text data.
[0006] To solve the above technical problems, the technical solution of the present invention is as follows:
[0007] A method for extracting multi-modal aspect-sentiment pairs enhanced by sentiment reason knowledge, the method comprising:
[0008] Generating sentiment reason knowledge for each text-image data pair; wherein, the text-image data pair includes: an original text and an original image;
[0009] With the assistance of the sentiment reason knowledge, obtaining the aspect-sentiment pair of the text-image data pair.
[0010] Further, generating sentiment reason knowledge for each text-image data pair includes:
[0011] Randomly sampling from the training set, and manually annotating the sentiment reason knowledge of the selected text-image data pair samples to construct a set of manually annotated examples; wherein, the content of each manually annotated example includes: the original text of the text-image data pair sample, the picture text description of the text-image data pair sample, a question, and sentiment reason knowledge;
[0012] Based on the similarity between the text-image data pair and the text-image data pair samples, selecting K most similar manually annotated examples for the text-image data pair from the set of manually annotated examples;
[0013] Obtaining the original text and picture text description of the text-image data pair;
[0014] Constructing a test example, the content of the test example including: the original text of the text-image data pair, the picture text description of the text-image data pair, and a question;
[0015] Embedding the test example and the corresponding K most similar manually annotated examples into a sentiment reason knowledge generation prompt template, and generating sentiment reason knowledge for the text-image data pair based on a large language model.
[0016] Further, based on the similarity between the text-image data pair and the text-image data pair samples, selecting K most similar manually annotated examples for the text-image data pair from the set of manually annotated examples includes:
[0017] Encoding the text-image data pair and each text-image data pair sample in the set of manually annotated examples respectively;
[0018] Based on the encoding results, calculating the cosine similarity between the text-image data pair and each text-image data pair sample;
[0019] According to the cosine similarity, obtaining the K most similar manually annotated examples of the text-image data pair.
[0020] Further, the picture text description of the picture-text data pair is generated based on the vision-language pre-training model BLIP-2.
[0021] Further, with the assistance of the emotional cause knowledge, the aspect-sentiment pair of the picture-text data pair is obtained, including:
[0022] Concatenate the original text and the emotional cause knowledge to obtain a new text;
[0023] Input the new text into the Transformer-based encoder for encoding to obtain an encoded vector;
[0024] Input the encoded vector into a linear-chain conditional random field for predicting sequence labels to obtain label sequence probabilities;
[0025] According to the label sequence probabilities, obtain the aspect-sentiment pair of the picture-text data pair.
[0026] Further, the process of training the Transformer-based encoder and the conditional random field includes:
[0027] Calculate the label sequence probabilities of the picture-text data pair samples;
[0028] Use negative log-likelihood to define the loss function, and calculate the loss value based on the label sequence probabilities of the picture-text data pair samples and the true label sequences of the picture-text data pair samples;
[0029] Perform backpropagation based on the loss value to update the parameters of the Transformer-based encoder and the conditional random field.
[0030] An emotional cause knowledge-enhanced multi-modal aspect-sentiment pair extraction system, the system includes:
[0031] An emotional cause knowledge generation module, configured to generate emotional cause knowledge for each picture-text data pair; wherein, the picture-text data pair includes: an original text and an original picture;
[0032] An aspect-sentiment pair generation module, configured to obtain the aspect-sentiment pair of the picture-text data pair with the assistance of the emotional cause knowledge.
[0033] An electronic device, the electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the emotional cause knowledge-enhanced multi-modal aspect-sentiment pair extraction method described in any one of the above is implemented.
[0034] A computer-readable storage medium stores computer program instructions thereon, and when the computer program instructions are executed by a processor, the multi-modal aspect-sentiment pair extraction method with enhanced sentiment cause knowledge described in any one of the above is implemented.
[0035] A computer program product, when running on a computer device, causes the computer device to execute the multi-modal aspect-sentiment pair extraction method with enhanced sentiment cause knowledge described in any one of the above.
[0036] Compared with the prior art, the present invention has at least the following beneficial effects.
[0037] 1) The present invention converts the picture-text matching mode into a text-text matching mode, and can perform modal fusion without semantic differences.
[0038] 2) Empowered by a large language model, the present invention extracts language knowledge containing picture content descriptions and sentiment cause explanations from picture-text data pairs, provides a richer context for simple text content to assist in understanding the semantics of the original text, thereby improving the accuracy of the extraction task and having good practicality. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a flowchart of the multi-modal aspect-sentiment pair extraction method with enhanced sentiment cause knowledge provided by an embodiment of the present invention.
[0040] Figure 2 It is a prompt template for ChatGPT to generate sentiment cause knowledge in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0041] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below through specific implementation cases and in conjunction with the accompanying drawings.
[0042] Figure 1 It is a flowchart of the multi-modal aspect-sentiment pair extraction method with enhanced sentiment cause knowledge. As Figure 1 shown, the method mainly includes two stages, namely heuristic generation of sentiment cause knowledge and extraction of aspect-sentiment pairs with enhanced sentiment cause knowledge. The whole process is to first randomly sample some samples from the training data for manual annotation, and then let the large language model annotate all the data under the guidance of the manual annotation prompts. After obtaining the sentiment cause knowledge, it is spliced as auxiliary content after the original input text, and then the text encoder and sequence tagger are trained on the training data, and then applied to the actual extraction.
[0043] (1) Heuristic generation stage of sentiment cause knowledge.
[0044] In the heuristic generation stage of emotional reason knowledge, the present invention uses manually annotated examples as a guide to stimulate the generation ability of the large language model, and generates emotional reason knowledge containing multi-modal semantics for each text-image data pair. Among them, the text-image data pair described in this embodiment includes a piece of text and a picture.
[0045] Step 1, predefined artificial examples: The present invention randomly samples a certain proportion (such as 1 / 21) of text-image data pair samples from the training data for manual annotation of emotional reason knowledge, obtaining a set of artificial examples. The annotation process is divided into two steps:
[0046] (1-1) Identify all the target aspect item contents mentioned in the text, including the aspect items mentioned in the text and their corresponding emotional polarities;
[0047] (1-2) Considering the text content, picture information and corresponding true labels at the same time, provide a comprehensive argument for judging each aspect item and its corresponding emotional polarity.
[0048] In this way, a set of predefined manually annotated example sets can be obtained, and each element in the set is composed of the text content of the text-image data pair sample, the picture information, and the manually annotated emotional reason knowledge.
[0049] Step 2, similarity-aware example selection: The present invention first obtains the multi-modal representations of each text-image data pair. Based on this, by measuring the similarity between the multi-modal representations of the text-image data pairs, the most similar K examples are selected from the set of artificial examples for each text-image data pair to construct an example demonstration for the large language model to reason.
[0050] (2-1) Define the pre-trained multi-modal aspect-emotion pair extraction model by predecessors as M, and this model is composed of the base encoder M b and the sequence tagger M c . Then, for any input of a multi-modal text-image data pair containing a picture I and text T, use the encoder M b to encode:
[0051] H = M b (T, I)
[0052] (2-2) Given the data set D and the predefined set of manually annotated examples G, calculate the cosine similarity between each text-image data pair in the data set D and each text-image data pair in the set of manually annotated examples G, and select the K examples with the highest similarity from the set of manually annotated examples G. For the i-th text-image data pair in the data set D, calculate the set S i of its K most similar examples:
[0053]
[0054] Step 3, Heuristic-Enhanced Prompt Generation: For each text-image data pair in the data set D, construct a prompt for the large language model ChatGPT to reason. The prompt content includes a prompt header, a set of context example demonstrations, and a test input. The prompt header is the natural language of the task description, and each example in the context example demonstrations is constructed according to the specified template:
[0055] Text: T k , Image: V k , Question: Q, Answer: A k
[0056] Among them, T k is the text content in the text-image data pair input, V k is the text description of the image in the text-image data pair input, Q is the natural language for asking ChatGPT questions, and A k is the sentiment reason knowledge manually annotated for the text-image data pair (T k , V k ). Accordingly, the context example demonstrations are obtained by combining the content constructed from K examples in the example set S according to the above specified template. The example template for the test input is:
[0057] Text: T, Image: V, Question: Q, Answer:
[0058] ChatGPT will generate a new answer for this test input example under the guidance of the examples in the context example demonstrations, as the sentiment reason knowledge Z for this example.
[0059] (2) Aspect-Sentiment Pair Extraction Phase with Sentiment Reason Knowledge Enhancement.
[0060] In the aspect-sentiment pair extraction phase with sentiment reason knowledge enhancement, the present invention uses the sentiment reason knowledge from the previous phase to enrich the context of the original input text content to assist in extracting aspect-sentiment pairs.
[0061] Step 1, Construct Input: Concatenate the text content T in the text-image data pair and its corresponding sentiment reason knowledge Z to construct a new text content [T; Z], as the model input for the aspect-sentiment pair extraction phase, where the text length of T is n and the text length of Z is m;
[0062] Step 2, Text Encoding: Use the Transformer-based text encoder XLM-RoBERTa to encode the text input [T; Z] to obtain the representation vectors (h 1 ,..., h n ,..., h n+m );
[0063] Step 3, Label Prediction: Use a linear-chain conditional random field (CRF) to perform sequence label prediction. Given the text content T and the corresponding sentiment reason knowledge Z, define the calculation process of the label sequence probability for this text-image data pair:
[0064]
[0065] where φ(y i-1 , y i , h i ) and φ(y′ i-1 , y′ i , h i ) are both potential functions, and Y represents the set of all possible label sequences;
[0066] Step 4, Model Training: Define the loss function using negative log-likelihood:
[0067]
[0068] where y * represents the true label sequence, and θ represents the trainable parameters in the model;
[0069] Step 5, Adopt the backpropagation algorithm and train XLM-RoBERTa and CRF according to .
[0070] In summary, to address the problem of insufficient context in the text content of the multi-modal aspect-sentiment pair task, an aspect-sentiment pair extraction method enhanced with sentiment reason knowledge is proposed to extract the (aspect, sentiment) binary pairs mentioned in the text-image data pair. Specifically: By means of manual annotation, label the sentiment reason knowledge of a small amount of text-image data; then, through the similarity-aware example selection module, select relevant examples for each sample in the dataset from the set of manually annotated examples and integrate them into the prompt template to construct example samples for context learning; afterwards, relying on the few-shot context learning ability of the large language model, use a heuristic method to refine auxiliary knowledge similar to the manually annotated content from the original text; finally, with the assistance of sentiment reason knowledge, extract all the mentioned aspect-sentiment pairs from the original text. In this way, first, the present invention converts the picture-text matching mode into a text-text matching mode, enabling modal fusion without semantic differences; second, empowered by the large language model, it extracts language knowledge containing picture content descriptions and sentiment reason explanations from the text-image data pair, providing a richer context for simple text content to assist in understanding the semantics of the original text, thereby improving the accuracy of the extraction task and having good practicality.
[0071] Another embodiment of the present invention provides a multi-modal aspect-sentiment pair extraction system enhanced with sentiment reason knowledge, including:
[0072] An emotional reason knowledge generation module for generating emotional reason knowledge for each text-image data pair; wherein, the text-image data pair includes: an original text and an original image;
[0073] An aspect-emotion pair generation module for obtaining the aspect-emotion pair of the text-image data pair with the assistance of the emotional reason knowledge.
[0074] For the specific implementation process of each module, refer to the previous description of the method of the present invention. For example, the emotional reason knowledge generation module first annotates a small number of text-image data pairs in an artificial annotation manner, and the annotation content includes the description of the image content and a brief reason explanation related to the task; then, through the similarity-aware example selection module, relevant examples are selected for each sample in the dataset from the set of artificially annotated examples, and they are integrated into the prompt template to construct example samples for context learning; afterwards, with the help of the few-shot context learning ability of the large language model, auxiliary knowledge similar to the artificially annotated content is refined for the original text through a heuristic method, so as to provide rich document-level information for the original text. Then, the aspect-emotion pair generation module splices the generated emotional reason knowledge after the input text content as auxiliary information, constructs a new text input, encodes it using the XLM-RoBERTa encoder, and then uses the CRF sequence tagger for label prediction, and finally infers the (aspect, emotion) binary group.
[0075] Another embodiment of the present invention provides a computer device (such as a computer, a server, a smart phone, etc.), which includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in the method of the present invention.
[0076] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, a disk, an optical disc), the computer-readable storage medium stores a computer program, and when the computer program is executed by the computer, each step of the method of the present invention is realized.
[0077] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the concept of the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as within the protection scope of the present invention.
Claims
1. A multimodal aspect-sentiment pair extraction method enhanced with sentiment cause knowledge, characterized in that: The method comprises: Generate emotional reason knowledge for each image-text data pair; wherein the image-text data pair includes: original text and original picture; With the assistance of the emotion cause knowledge, the aspect-emotion pair of the image-text data pair is obtained.
2. The method according to claim 1, characterized in that The generation of emotional reason knowledge for each image-text data pair includes: Random sampling is performed from the training set, and the selected image-text data pairs are manually annotated with emotional cause knowledge to construct a manually annotated example set; wherein the content of each manually annotated example includes: the original text of the image-text data pair sample, the image text description of the image-text data pair sample, the question and emotional cause knowledge; Based on the similarity between the image-text data pair and the image-text data pair sample, selecting K most similar manually labeled examples for the image-text data pair in the manually labeled example set; Obtain the original text and picture text description of the picture-text data pair; Constructing a test example, wherein the content of the test example includes: the original text of the image-text data pair, the image text description of the image-text data pair, and the question; The test example and the K most similar manually annotated examples corresponding to the test example are embedded into the sentiment cause knowledge generation prompt template, and the sentiment cause knowledge is generated for the image-text data pair based on the large language model.
3. The method according to claim 2, characterized in that Based on the similarity between the image-text data pair and the image-text data pair sample, selecting K most similar manual annotation examples for the image-text data pair from the manual annotation example set includes: Encoding the image-text data pairs and each image-text data pair sample in the manually annotated example set respectively; Based on the encoding result, the cosine similarity between the image-text data pair and each image-text data pair sample is calculated; According to the cosine similarity, the K most similar manually labeled examples of the image-text data pair are obtained.
4. The method according to claim 2, characterized in that: The image text description of the image-text data pair is generated based on the vision-language pre-training model BLIP-2.
5. The method according to claim 1, characterized in that: With the help of the emotion cause knowledge, the aspect-emotion pair of the image-text data pair is obtained, including: splicing the original text and the emotional cause knowledge to obtain a new text; Input the new text into the Transformer-based encoder for encoding to obtain the encoding vector; Input the encoding vector into the conditional random field of the linear chain to predict the sequence label and obtain the label sequence probability; According to the label sequence probability, the aspect-sentiment pair of the image-text data pair is obtained.
6. The method according to claim 5, characterized in that The process of training the Transformer-based encoder and the conditional random field comprises: Calculate the label sequence probability of the image and text data for the sample; The loss function is defined using negative log-likelihood, and the loss value is calculated based on the label sequence probability of the image-text data pair sample and the true label sequence of the image-text data pair sample; Back propagation is performed based on the loss value to update the parameters of the Transformer-based encoder and the conditional random field.
7. A multimodal aspect-sentiment pair extraction system enhanced with sentiment cause knowledge, characterized in that: The system comprises: The emotion cause knowledge generation module is used to generate emotion cause knowledge for each image-text data pair; wherein the image-text data pair includes: original text and original picture; The aspect-emotion pair generation module is used to obtain the aspect-emotion pair of the image-text data pair with the assistance of the emotion cause knowledge.
8. An electronic device, characterized in that: The electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the multimodal aspect-emotion pair extraction method with enhanced emotion cause knowledge as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the multimodal aspect-emotion pair extraction method with enhanced emotion cause knowledge as described in any one of claims 1-6.
10. A computer program product, characterized in that When the computer program product runs on a computer device, the computer device executes the multimodal aspect-emotion pair extraction method enhanced with emotion cause knowledge as described in any one of claims 1 to 6.