Emotion analysis method and system based on structure perception language model

By constructing a structure-aware language model, combining global grammatical information and subgraphs, the problem of ignoring structured information in traditional methods is solved, and a more refined and fine-grained sentiment analysis is achieved.

CN120297262APending Publication Date: 2025-07-11GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510140392.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Traditional emotion analysis methods are difficult to accurately capture detailed emotional information in the text, especially ignoring structured information, resulting in poor analysis results.

Method used

By building a structure-aware language model, obtaining aspects-opinions pairs, forming a natural language description of the global structure, combining sub-graphs for sentiment analysis, introducing global grammar information and element link prediction, model training and fine-tuning, and improving the fineness of sentiment analysis.

Benefits of technology

It improves the depth and accuracy of understanding of emotional information, enhances the understanding of context and emotional transmission paths, and improves the precision and accuracy of analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297262A_ABST
    Figure CN120297262A_ABST
Patent Text Reader

Abstract

The invention relates to the field of text analysis, in particular to an emotion analysis method and system based on a structure perception language model.The method comprises the steps that a to-be-analyzed text is obtained, and aspect-opinion pairs are obtained according to the to-be-analyzed text; connecting aspect-opinion pairs to obtain natural language description of a global structure; constructing a subgraph according to the aspect-opinion pair; constructing a language model and training to obtain an emotion analysis model; combining the aspect-opinion pair, the natural language description of the global structure and the subgraph with an emotion analysis model to obtain a structure perception language model; and performing sentiment analysis on the to-be-analyzed text to obtain an analysis result. According to the method, the aspect-opinion pair, the natural language description of the global structure and the structure information of the subgraph are combined, so that the understanding depth and accuracy of the emotion information are improved. Meanwhile, the sub-graph reveals the relation between words, the model is helped to better understand the context and the emotion transmission path, and therefore the analysis fineness is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of text analysis, and more specifically, to a sentiment analysis method and system based on a structure-aware language model. Background Art

[0002] With the rapid development of the mobile Internet, e-commerce platforms and social networks have deeply integrated into people's daily lives and become an indispensable part. These platforms not only provide users with a convenient online shopping experience but also build a huge interactive community where users can freely share their true feelings about products and services. Through the online review system, consumers can elaborate on their views on various aspects such as product usage experience, service attitude, and product quality, which provides rich data resources for intelligent business analysis. Enterprises can formulate more accurate product improvement plans and service optimization strategies by intelligently analyzing these massive user evaluations. At the same time, a smooth consumer feedback channel helps enterprises quickly respond to market demands, improve customer satisfaction, and enhance brand loyalty. In today's digital age, online media has replaced traditional paper media as the main means of information dissemination. Social platforms such as Weibo, WeChat, and forums not only allow users to obtain the latest hot events but also provide a space for public discussion and personal expression, enabling the rapid spread and exchange of public opinions. In this context, online public opinion analysis has become particularly important. It is not only a key tool for enterprises to maintain their brand images and respond to emergencies but also an important means for the government to understand social dynamics, grasp public opinion trends, and create a healthy online environment. The rich information hidden in online texts provides valuable decision-making support for all walks of life, but at the same time, it also brings huge challenges: Facing such a large and growing amount of data, traditional manual processing methods are obviously unable to meet the requirements. Therefore, how to automatically extract and deeply analyze the opinions and emotions of users towards specific objects from massive online texts has become an important research direction in the field of natural language processing (NLP). The development of fine-grained sentiment analysis technology aims to solve this problem by identifying and understanding the subtle emotional differences in texts, helping enterprises and researchers more accurately capture the true intentions and emotional tendencies of consumers. This can not only improve the efficiency of data analysis but also provide more specific and valuable insights for decision-makers, helping enterprises stand out in the highly competitive market and promoting the healthy interaction and communication between society and the government.

[0003] Traditional sentiment analysis methods usually conduct an overall evaluation of a piece of text to determine whether the sentiment expressed in it is positive, negative, or neutral. However, this type of text-level or sentence-level sentiment analysis often fails to accurately capture the user's detailed sentiment towards specific things. For example, in a comment like "Although the screen display effect is good, the battery life is very short", it contains both positive and negative sentiment information. In this case, if only the overall sentiment is focused on, these important detailed differences may be overlooked, leading to a misunderstanding of the user's true feelings. To obtain more accurate analysis results, a fine-grained sentiment analysis method is needed.

[0004] Most current sentiment analysis methods focus on sentiment vocabulary and sentence-level classification, and involve less in the processing of more detailed sentiment information. Introducing structure-enhanced language models provides new possibilities for addressing this challenge. Structural information plays a crucial role in natural language, including syntactic structures (such as the modification relationships between words), and these structures provide additional information for understanding and generating natural language, making language expressions more accurate and coherent. When traditional graph neural network (GNN)-based structure-aware models process natural language, they usually rely on constructing explicit relationship graphs between words. However, with the rise of large language models (LLMs), it is difficult for structure-aware models to be extended to the current mainstream autoregressive large language models (GNNs require explicit graph inputs, while the input of autoregressive models is a simple text sequence, and the conversion between the two is not intuitive). Therefore, current sentiment analysis methods often ignore structural information, resulting in poor analysis effects. Summary of the Invention

[0005] The purpose of the present invention is to disclose a sentiment analysis method and system based on a structure-aware language model with better analysis effects.

[0006] To achieve the above purpose, the present invention provides a sentiment analysis method based on a structure-aware language model, including:

[0007] S1: Obtain the text to be analyzed, and obtain aspect-opinion pairs according to the text to be analyzed;

[0008] S2: Connect the aspect-opinion pairs according to the order in which they appear in the text to be analyzed to obtain a natural language description of the global structure; regard the aspect-opinion pairs as the central nodes in the dependency tree, and extract adjacent words from them to construct a subgraph; S3: Construct a language model and train it to obtain a sentiment analysis model;

[0009] S4: Combine the aspect-opinion pairs, the natural language description of the global structure, and the subgraph with the sentiment analysis model to obtain a structure-aware language model;

[0010] S5: Perform sentiment analysis on the text to be analyzed according to the structure-aware language model to obtain the analysis result.

[0011] Further, in step S1, it includes: introducing global syntactic information and extracting it by considering the modification relationships between various elements in the sentence.

[0012] Further, in step S2, connecting the aspect-opinion pairs to obtain the natural language description of the global structure includes:

[0013] Connect the aspect-opinion pairs according to the order in which they appear in the sentence to obtain the natural language description of the global structure; the natural language description of the global structure takes into account the order of multiple elements in the sentence and the mutual relationships between the aspect-opinion pairs, forming a more comprehensive context information. The natural language of the global structure is then appended to the original sentence, thereby enhancing the syntactic context of the sentence. At the same time, use the instruction fine-tuning method to fine-tune the model with instruction I1 and output O1; the goal of supervised fine-tuning is to minimize the loss function, and the loss function quantifies the gap between the model prediction and the actual label through a specific mathematical form.

[0014] Further, in step S2, constructing a subgraph according to the aspect-opinion pairs includes:

[0015] Regard the aspect-opinion pairs as the central nodes in the dependency tree, and extract the adjacent words around the central nodes to construct a subgraph: the subgraph around a certain central node includes not only the direct adjacent nodes of the node, but also considers the adjacent node information with a maximum jump distance of one.

[0016] Further, in step S3, it includes: the model is a large language model.

[0017] Further, in step S3, the training dataset used is the public dataset ACOS, including a restaurant dataset and a computer dataset. There are 1,500 items in the restaurant dataset and 2,800 items in the computer dataset; in the restaurant dataset, it contains evaluations of food and different restaurants, and for each piece of text, the aspect words, opinion words, categories, and sentiment polarities are labeled. The computer dataset contains evaluations of the batteries and graphics cards of laptops; the computer dataset also has labels for the aspect words, opinion words, categories, and sentiment polarities of each piece of text.

[0018] Further, in step S3, it includes: the performance metric for training is the F1 value. F1 is used to measure the performance of the model in classification or information retrieval tasks; F1 comprehensively considers the precision and recall of the model. In classification tasks, precision represents the proportion of samples predicted as positive examples that are actually positive examples, and recall represents the proportion of samples that are actually positive examples that are correctly predicted as positive examples.

[0019] Further, in step S4, it further includes: introducing element link prediction in the structure-aware language model, and the element link prediction includes:

[0020] Guiding the model to associate elements with their corresponding entities; in the element link prediction, directly exposing opinion words to the model, so as to find corresponding aspect words in the dependency tree to enhance the model's cognition of structural information and the matching ability of sentiment elements.

[0021] Further, in step S4, it further includes: introducing node classification in the structure-aware language model, and the node classification includes: the node classification prompts the model to assign accurate labels to each node in the structural framework.

[0022] In addition, the present invention also provides a fine-grained sentiment analysis system based on a structure-aware language model, including:

[0023] An acquisition module: acquiring the text to be analyzed, and obtaining aspect-opinion pairs according to the text to be analyzed;

[0024] A connection construction module: connecting the aspect-opinion pairs according to the order in which they appear in the text to be analyzed to obtain a natural language description of the global structure; regarding the aspect-opinion pairs as the central nodes in the dependency tree, and extracting adjacent words therefrom to construct a subgraph; A training module: constructing a language model and training it to obtain a sentiment analysis model;

[0025] A combination module: combining the aspect-opinion pairs, the natural language description of the global structure, and the subgraph with the sentiment analysis model to obtain a structure-aware language model;

[0026] An analysis module: performing sentiment analysis on the text to be analyzed according to the structure-aware language model to obtain an analysis result.

[0027] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0028] By combining structural information such as aspect-opinion pairs, natural language descriptions of global structures, and subgraphs, the present invention improves the depth and accuracy of understanding of sentiment information. At the same time, the subgraph reveals the relationships between words, helps the model better understand the context and sentiment transmission path, thereby improving the fineness of analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a flowchart of a sentiment analysis method based on a structure-aware language model described in Embodiment 1;

[0030] Figure 2 It is a graph of the training result of the structure-aware language model described in Embodiment 2;

[0031] Figure 3Block diagram of the fine-grained sentiment analysis system based on the structure-aware language model described in Embodiment 3; Detailed implementation manners

[0032] The accompanying drawings are only for illustrative purposes and should not be construed as limitations on this patent;

[0033] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0034] Embodiment 1:

[0035] This embodiment provides a sentiment analysis method based on the structure-aware language model as shown in Figure 1 and includes:

[0036] S1: Obtain the text to be analyzed, and obtain the aspect-opinion pairs according to the text to be analyzed;

[0037] S2: Connect the aspect-opinion pairs according to the order in which they appear in the text to be analyzed to obtain a natural language description of the global structure; regard the aspect-opinion pairs as the central nodes in the dependency tree, and extract adjacent words therefrom to construct a subgraph; S3: Construct a language model and train it to obtain a sentiment analysis model;

[0038] S4: Combine the aspect-opinion pairs, the natural language description of the global structure, and the subgraph with the sentiment analysis model to obtain a structure-aware language model;

[0039] S5: Perform sentiment analysis on the text to be analyzed according to the structure-aware language model to obtain the analysis result.

[0040] In this embodiment, by combining the structural information such as aspect-opinion pairs, the natural language description of the global structure, and the subgraph, the depth and accuracy of the understanding of sentiment information are improved. At the same time, the subgraph reveals the relationships between words, helps the model better understand the context and the sentiment transmission path, thereby improving the fineness of the analysis.

[0041] Embodiment 2:

[0042] This embodiment further discloses on the basis of Embodiment 1:

[0043] In step S1, it includes: introducing global syntactic information and extracting it by considering the modification relationships between various elements in the sentence.

[0044] , in step S2, connecting the aspect-opinion pairs to obtain the natural language description of the global structure includes:

[0045] Connect according to the order of appearance in the sentence for the aspect-opinion pairs to obtain the natural language description of the global structure; the natural language description of the global structure takes into account the order of multiple elements in the sentence and the mutual relationship between the aspect-opinion pairs, forming a more comprehensive context information. The natural language of the global structure is then appended to the original sentence, thereby enhancing the syntactic context of the sentence. At the same time, using the instruction fine-tuning method, the model is fine-tuned with instruction I1 and output O1; the goal of supervised fine-tuning is to minimize the loss function, and the loss function quantifies the gap between the model prediction and the actual label through a specific mathematical form.

[0046] In step S2, construct a subgraph according to the aspect-opinion pairs, including:

[0047] Regard the aspect-opinion pair as the central node in the dependency tree, and extract adjacent words around the central node to construct a subgraph: the subgraph around a certain central node includes not only the direct adjacent nodes of the node, but also considers the adjacent node information with a maximum jump distance of one.

[0048] In step S3, include: the model is a large language model.

[0049] In step S3, the training dataset used is the public dataset ACOS, including a restaurant dataset and a computer dataset. There are 1500 entries in the restaurant dataset and 2800 entries in the computer dataset; in the restaurant dataset, it contains evaluations of food and different restaurants, and for each piece of text, the aspect words, opinion words, categories, and sentiment polarities are labeled. While the computer dataset contains evaluations of the batteries and graphics cards of laptop computers; the computer dataset also has the labels of aspect words, opinion words, categories, and sentiment polarities for each piece of text.

[0050] In step S3, include: the performance metric for training is the F1 value. F1 is used to measure the performance of the model in classification or information retrieval tasks; F1 comprehensively considers the precision and recall of the model. In the classification task, the precision represents the proportion of samples predicted as positive examples that are actually positive examples, while the recall represents the proportion of samples that are actually positive examples and are correctly predicted as positive examples.

[0051] In step S4, also include: introduce element link prediction in the structure-aware language model, and the element link prediction includes:

[0052] Guide the model to associate elements with their corresponding entities; in element link prediction, expose the opinion words directly to the model, so as to find the corresponding aspect words in the dependency tree to enhance the model's awareness of structural information and the matching ability of sentiment elements.

[0053] In step S4, it further includes: introducing node classification into the structure-aware language model, where node classification includes: node classification prompts the model to assign accurate labels to each node in the structure framework.

[0054] In this embodiment, by combining aspect-opinion pairs, natural language descriptions of the global structure, and subgraphs of these structural information, the depth and accuracy of the understanding of sentiment information are improved. At the same time, the subgraph reveals the relationships between words, helping the model better understand the context and sentiment transmission path, thereby improving the fineness of the analysis.

[0055] In this embodiment, the sentiment tuple extraction task is completed through a two-stage fine-tuning method. This method can effectively improve the performance of the model in sentiment analysis tasks. Specifically, in the first stage, the goal is to extract potential aspect-opinion pairs from sentences. Traditional methods often directly focus on single words or phrases. The present invention introduces global syntactic information, which can guide the extraction by considering the modification relationships between various elements in the sentence. By injecting these relationships into large language models (LLMs) and using the natural language description of triple relationships (for example, assuming x i as the head node of x h ), and recording the details of these syntactic structures. This process can better capture the syntactic structure in the sentence, thereby improving the extraction accuracy of aspect-opinion pairs. In this way, the model can effectively understand complex grammar and modification relationships, and then extract accurate sentiment tuples.

[0056] After extracting the potential aspect-opinion pairs in this embodiment, the next step is to connect them according to the order in which they appear in the sentence, thereby forming a natural language description of the global structure. This description not only focuses on individual aspect-opinion pairs, but also considers the order of multiple elements in the sentence and their mutual relationships, forming a more comprehensive context information. This natural language description is then appended to the original sentence, thereby enhancing the syntactic context of the sentence and helping the model better understand the overall structure of the sentence. This process effectively captures the syntactic information of the sentence and avoids the problem of traditional methods ignoring the overall structure of the sentence. To optimize this process, the instruction fine-tuning method is used, and the model is fine-tuned using instruction I1 and output O1. The goal of supervised fine-tuning is to minimize the loss function, which quantifies the gap between the model prediction and the actual label through a specific mathematical form, thereby guiding the model to gradually improve its performance in each round of fine-tuning.

[0057] After obtaining the aspect-opinion pairs in this embodiment, these pairs are regarded as the central nodes in the dependency tree, and adjacent words are extracted around these central nodes to construct a subgraph. Specifically, the subgraph around a certain central node includes not only the direct adjacent nodes of this node, but also considers the adjacent node information with a maximum jump distance of one. Such a design enables the subgraph to not only capture directly relevant words, but also incorporate certain context information, thereby providing more background knowledge to help the model better understand the relationships between various words. The establishment of this structure not only increases the accuracy of sentiment tuple extraction, but also optimizes the sentiment analysis results through more comprehensive context information. For this reason, this embodiment designs a loss function to perform supervised fine-tuning on the model, further improving the model's sensitivity and understanding ability to context information.

[0058] This embodiment also further improves the model's adaptability to specific structured knowledge tasks by introducing two auxiliary tasks: element link prediction and node classification. These auxiliary tasks can help the model better understand complex syntactic structures and improve its performance in the sentiment tuple extraction task. Specifically, the element link prediction task aims to help the model more accurately identify the relationships between different elements in a sentence, while the node classification task is to better assign accurate labels (such as sentiment polarity and aspect category) to each node in the structural framework. By introducing these two auxiliary tasks, the model can better capture the multi-dimensional information of sentiment tuples, thereby improving the performance of the overall task.

[0059] Embodiment Three:

[0060] This embodiment provides a Figure 3 fine-grained sentiment analysis system based on a structure-aware language model as shown in

[0061] An acquisition module: acquires the text to be analyzed and obtains aspect-opinion pairs according to the text to be analyzed;

[0062] A connection construction module: connects the aspect-opinion pairs according to the order in which they appear in the text to be analyzed to obtain a natural language description of the global structure; regards the aspect-opinion pairs as the central nodes in the dependency tree and extracts adjacent words from them to construct a subgraph; A training module: constructs a language model and trains it to obtain a sentiment analysis model;

[0063] A combination module: combines the aspect-opinion pairs, the natural language description of the global structure, and the subgraph with the sentiment analysis model to obtain a structure-aware language model;

[0064] An analysis module: performs sentiment analysis on the text to be analyzed according to the structure-aware language model to obtain an analysis result.

[0065] In this embodiment, by combining aspect-opinion pairs, natural language descriptions of the global structure, and subgraphs of these structural information, the depth and accuracy of the understanding of sentiment information are improved. At the same time, the subgraph reveals the relationships between words, helping the model better understand the context and sentiment transmission path, thereby improving the fineness of the analysis.

[0066] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limiting the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A sentiment analysis method based on a structure-aware language model, characterized in that, Including: S1: Obtain the text to be analyzed, and get the aspect-opinion pairs according to the text to be analyzed; S2: Connect the aspect-opinion pairs according to the order in which they appear in the text to be analyzed to obtain a natural language description of the global structure; regard the aspect-opinion pairs as the central nodes in the dependency tree, and extract adjacent words from them to construct a subgraph; S3: Construct a language model and train it to obtain a sentiment analysis model; S4: Combine the aspect-opinion pairs, the natural language description of the global structure, and the subgraph with the sentiment analysis model to obtain a structure-aware language model; S5: Perform sentiment analysis on the text to be analyzed according to the structure-aware language model to obtain the analysis result.

2. The sentiment analysis method based on a structure-aware language model according to claim 1, characterized in that In step S1, it includes: introducing global syntactic information, and extracting aspect-opinion pairs by considering the modification relationships between various elements in the text to be analyzed.

3. The sentiment analysis method based on a structure-aware language model according to claim 1, characterized in that In step S2, connecting the aspect-opinion pairs to obtain the natural language description of the global structure includes: Connecting according to the order in which the aspect-opinion pairs appear in the sentence to obtain a natural language description of the global structure; the natural language description of the global structure considers the order of multiple elements in the sentence and the mutual relationships between the aspect-opinion pairs, forming a more comprehensive context information; the natural language of the global structure is then appended to the original sentence, thereby enhancing the syntactic context of the sentence. At the same time, the instruction fine-tuning method is used to fine-tune the model with instruction I1 and output O1; the goal of supervised fine-tuning is to minimize the loss function, and the loss function quantifies the gap between the model prediction and the actual label through a specific mathematical form.

4. The sentiment analysis method based on a structure-aware language model according to claim 1, characterized in that In step S2, constructing a subgraph according to the aspect-opinion pairs includes: Regarding the aspect-opinion pairs as the central nodes in the dependency tree, and extracting adjacent words around the central nodes to construct a subgraph: the subgraph around a certain central node includes not only the direct adjacent nodes of the node, but also considers the adjacent node information with a maximum jump distance of one.

5. The sentiment analysis method based on a structure-aware language model according to claim 1, wherein In step S3, it includes: the language model is a large language model.

6. The sentiment analysis method based on a structure-aware language model according to claim 1, wherein In step S3, the dataset used for training is the publicly available dataset ACOS, including a restaurant dataset and a computer dataset. There are 1500 items in the restaurant dataset and 2800 items in the computer dataset; in the restaurant dataset, it contains evaluations of food and different restaurants, and for each piece of text, the aspect words, opinion words, categories, and sentiment polarities are annotated; while the computer dataset contains evaluations of the batteries and graphics cards of laptop computers; the computer dataset also has annotations of aspect words, opinion words, categories, and sentiment polarities for each piece of text.

7. The sentiment analysis method based on a structure-aware language model according to claim 1, characterized in that In step S3, it includes: the performance metric for training is the F1 value, and F1 is used to measure the performance of the model in classification or information retrieval tasks; F1 comprehensively considers the precision and recall of the model. In classification tasks, precision represents the proportion of samples predicted as positive examples that are actually positive examples, and recall represents the proportion of samples that are actually positive examples that are correctly predicted as positive examples.

8. The sentiment analysis method based on a structure-aware language model according to claim 1, characterized in that In step S4, it also includes: introducing element link prediction in the structure-aware language model, and element link prediction includes: The guiding model associates elements with their corresponding entities; in element link prediction, the opinion words are directly exposed to the model, so as to find the corresponding aspect words in the dependency tree to enhance the model's understanding of structural information and its ability to match sentiment elements.

9. The sentiment analysis method based on a structure-aware language model according to claim 1, characterized in that In step S4, it also includes: introducing node classification in the structure-aware language model, and the node classification includes: the node classification prompts the model to assign accurate labels to each node in the structural framework.

10. A fine-grained sentiment analysis system based on a structure-aware language model, characterized in that, It includes: Acquisition module: Acquire the text to be analyzed and obtain the aspect-opinion pairs according to the text to be analyzed; Connection construction module: Connect the aspect-opinion pairs according to the order in which they appear in the text to be analyzed to obtain a natural language description of the global structure; regard the aspect-opinion pairs as the central nodes in the dependency tree, and extract adjacent words from them to construct a subgraph; Training module: Construct a language model and train it to obtain a sentiment analysis model; Combination module: Combine the aspect-opinion pairs, the natural language description of the global structure, and the subgraph with the sentiment analysis model to obtain a structure-aware language model; Analysis module: Perform sentiment analysis on the text to be analyzed according to the structure-aware language model to obtain an analysis result.