A Sentiment Analysis Classification Method under a Multi-Class Knowledge System

The integration of BERT and LLM models with prompt learning and data enhancement techniques addresses accuracy issues in multi-class sentiment classification, enhancing model performance and reducing errors in large language models.

CN117235256BActive Publication Date: 2025-07-15ZHONGKE ZHIHE DIGITAL TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311007745.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-10
Publication Date
2025-07-15
Estimated Expiration
2043-08-10

AI Technical Summary

Technical Problem

In the sentiment analysis classification under multiple knowledge systems, traditional methods have low accuracy and large language models are prone to randomly generate answers.

Method used

Using the combination of BERT model and LLM model, through the method of data augmentation and prompt learning, an emotion analysis classification model under a multi-class knowledge system is constructed. The specific steps include data preprocessing, LLM model feature extraction, BERT model feature extraction and prompt learning application, and the binary classification results of BERT are used to supervise the LLM generation results.

Benefits of technology

It significantly improves the accuracy of text sentiment analysis, reduces the error generation of large language models, and improves the accuracy and stability of the model generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117235256B_ABST
    Figure CN117235256B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of natural language processing, and proposes a sentiment analysis and classification method under a multi-class knowledge system, aiming to solve the problem of low accuracy of existing multi-classification methods. The main solutions include performing data preprocessing and data augmentation on the tweets to obtain a text data set G. Then, the LLM strategy is adopted. After encoding the text and labels, they are put into a large language model for fine-tuning to obtain a five-classification vector. The BERT model is used for encoding to obtain the corresponding hidden vector, and then the feature extraction is performed on the hidden vector to obtain a feature vector. The training parameters are continuously updated. The feature vector passes through Softmax to obtain a probability value, and the probability is mapped to the solution space to obtain a binary classification result. The binary classification result is put into the constructed prompt template to perform prompt intervention on the output of the LLM result layer to obtain the final answer of five-classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and provides a sentiment analysis and classification method under a multi-class knowledge system. Background Art

[0002] Existing text sentiment classification techniques are divided into three types: rule-based, traditional machine learning-based, and deep learning-based. Rule-based classification algorithms mainly rely on manually constructing dictionaries for classification, with relatively high costs; traditional machine learning algorithms use manual annotation of part of the dataset. Common clustering algorithms and SVM classifiers perform well on small datasets, but have poor feature extraction capabilities on massive datasets; with the development of neural networks and deep learning, deep learning has become the mainstream direction for natural language processing problems. Deep learning algorithms represented by BERT and GPT series have strong learning capabilities for massive data features.

[0003] The application of text sentiment classification tasks in scenarios such as binary classification is numerous, and the BERT model has strong performance and is sufficient for most scenarios. However, when exploring its application in complex scenarios under the multi-classification task system, traditional text classification strategies have shortcomings and need to be optimized and iterated. Traditional multi-classification tasks, similar to binary classification tasks, need to perform a classifier judgment at the output layer stage. This output is a probability-based form, judging which class has the highest possibility. However, as the number of classes increases, the model needs to make judgments on more classes, and the training may be affected by more data noise, resulting in a relatively low prediction accuracy. LLM (Large Language Model) is a mainstream generative language model that has emerged in recent years. It can perform semantic modeling on natural language and is trained using a large-scale training set, so it can identify and simulate some complex language structures and semantic concepts. The sentiment multi-classification problem can achieve good results on the LLM model.

[0004] The present invention makes full use of the ideas of the BERT model, LLM model, and prompt learning, and proposes and implements a sentiment analysis and classification model under a multi-class knowledge system in the field of deep learning. The LLM model can effectively improve the accuracy of text sentiment, and the binary classification accuracy of BERT can prompt and supervise the results generated by the LLM, solving the problems of relatively low accuracy in traditional text multi-classification tasks and the random generation of answers by large language models. Summary of the Invention

[0005] The object of the present invention is to solve the problems of low accuracy in sentiment analysis classification under multiple knowledge systems and the random generation of answers by large language models. The data sources are divided into two categories: the publicly available SST-5 dataset and recently crawled Twitter data. First, the EDA algorithm is used to expand the dataset to complete data augmentation. The text statements are tokenized to obtain the context word vectors of each word in each statement. The training process is divided into three stages: first, a 5-classifier is constructed using the LLM model; second, BERT is used to distinguish positive and negative; finally, the BERT classification results are used to prompt the LLM generation results as the final model output, and the merged results are obtained as 5-class results. Experiments show that the accuracy of this model has been significantly improved in text sentiment analysis.

[0006] The present invention provides a sentiment analysis classification method under multiple knowledge systems, including the following steps:

[0007] Step 1: Both the training data and the validation data adopt the publicly available sst-5 dataset, and the test data comes from Twitter and is obtained by using a selenium crawler. It is stored in the format of [1abel, sentence], where label = {1, 2, 3, 4, 5} represents five kinds of sentiments. label = 1 represents extremely negative sentiment, label = 2 represents slightly negative sentiment, label = 3 represents neutral emotion, label = 4 represents slightly positive sentiment, and label = 5 represents extremely positive sentiment;

[0008] Step 2: The amount of data with neutral sentiment is small and the data distribution is uneven. Therefore, data augmentation needs to be performed on the training data. To ensure the emotional accuracy of the semantics, only entity extraction is performed on nouns, and the synonym replacement is implemented using the EDA-SR algorithm to expand the original dataset. In addition, to solve the data noise problem, the data needs to be preprocessed, and after the preprocessing is completed, the text dataset G is obtained;

[0009] Step 3: Based on the samples in the text dataset G obtained after the preprocessing in Step 2, for the LLM model, a 5-class dataset is constructed and encoded to obtain the encoded set E. For the BERT model, a positive and negative 2-class dataset is constructed and encoded to obtain the encoded set E';

[0010] Step 4: Use the LLM model to extract features from the preprocessed encoded dataset E using the generative large language model to obtain feature vectors, and through operations, obtain the sentiment 5-class vector:

[0011] llm_ans = LLM(E);

[0012] Step 5: Use the BERT model to extract features from the preprocessed and classified encoded dataset E′ to obtain feature vectors, and then perform a softmax operation to output the 2-classification result of positive and negative sentiment:

[0013] pos_neg_ans = softmax(BERT(E′));

[0014] Step 6: Package the 2-classification result obtained by BERT into the prompt template:

[0015] pos_neg_ans = 0 indicates negative, and the prompt template is as follows:

[0016]

[0017] Step 7: The LLM model generates a dialogue according to the prompt information prompt template;

[0018] Step 8: Under the guidance of the prompt template, the llm_ans vector generated in Step 4 generates a more accurate 5-classification result, and answer is the final classification result:

[0019] answer = LLM(llm_ans + prompt)

[0020] Among them, prompt can prompt the positive and negative information of semantics, thereby controlling the prediction direction of the model, and further reducing the wrong generation of the LLM model and improving the generation accuracy of the model;

[0021] In the above technical solution, the text dataset G is expressed as:

[0022] G = [(label1, sentence1), (label2, sentence2),..., (label N , sentence N )]

[0023] Among them, (label i , sentence i ) represents the i-th sample, i ∈ N, sentence i represents the i-th user comment, and label i is the sentiment category label corresponding to the text, and label i = {1, 2, 3, 4, 5}, and N represents the number of samples in the dataset G.

[0024] Because the present invention adopts the above technical solution, it has the following beneficial effects:

[0025] 1. By performing word segmentation and encoding on each sentence in all text information and combining the attention mechanism, it is possible to obtain the feature of the context word vector of each word in each sentence, which is beneficial for semantic understanding.

[0026] 2. Adopting data augmentation and prompt learning is beneficial for identifying the sentiment type in the text. After obtaining the LLM model through such training, it can effectively identify the multi-classification task type of sentiment text data, with better generalization and stability, thereby improving the accuracy of sentiment recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a model architecture diagram. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] The following will give a detailed description of the embodiments of the present invention. Although the present invention will be described and explained in conjunction with some specific embodiments, it should be noted that the present invention is not limited to these embodiments only. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments. Any modification or equivalent replacement made to the present invention shall be covered by the scope of the claims of the present invention.

[0029] In addition, in order to better illustrate the present invention, numerous specific details are given in the following detailed description. Those skilled in the art will understand that the present invention can be implemented without these specific details.

[0030] The present invention provides a sentiment analysis and classification method under a multi-class knowledge system, including the following steps:

[0031] Step 1: Selenium crawler is an automated testing tool that can be used to simulate user operations in a browser, such as opening web pages, entering text, clicking buttons, etc. By using the Selenium library, we can write Python code to control a real browser and simulate user operations such as accessing web pages, which is very useful for obtaining Twitter public opinion text. The test set data comes from the crawled Twitter user comments. The training set and validation set use the open-source dataset sst-5. The training, validation, and test datasets can be expressed as:

[0032] Set = [(label1, sentence1), (label2, sentence2),..., (label N , sentence N )

[0033] where, (label i , sentencei ) represents the i-th sample, sentence i represents the i-th user comment, label i is the sentiment category label for the corresponding text, label i = {1, 2, 3, 4, 5}, N represents the number of samples in the dataset Set;

[0034] Step 2: The amount of neutral sentiment data is small and the data distribution is uneven. Therefore, data augmentation needs to be performed on the training data. To ensure the emotional accuracy of the semantics, only entity extraction is performed on nouns, and the algorithm EDA-SR is used to implement synonym replacement to expand the original dataset. In addition, to solve the data noise problem, the data needs to be preprocessed, and after the preprocessing is completed, the text dataset G is obtained;

[0035] Step 3: For the LLM model, use the text dataset G to construct a 5-classification dataset and encode it to obtain the encoded set E. For the BERT model, use the text dataset G to construct a positive and negative 2-classification dataset and encode it to obtain the encoded set E';

[0036] Step 4: Use the LLM model to extract features from the preprocessed encoded dataset E using the generative large language model to obtain feature vectors, and through operations, obtain the sentiment 5-classification vector:

[0037] llm_ans = LLM(E)

[0038] Step 5: Use the BERT algorithm to extract features from the preprocessed encoded dataset E' respectively to obtain feature vectors, and then perform a softmax operation to output the positive and negative sentiment 2-classification results:

[0039] pos_neg_ans = softmax(BERT(E'))

[0040] Step 6: Enclose the 2-classification results obtained by BERT into the prompt template;

[0041] Step 7: The LLM model generates a dialogue according to the prompt information prompt;

[0042] Step 8: Under the guidance of the prompt template, the llm_ans vector generated in Step 4 generates a more accurate 5-classification result, and answer is the final classification result:

[0043] answer = LLM(llm_ans + prompt)

[0044] Among them, prompt can prompt the positive and negative information of the semantics, thereby controlling the prediction direction of the model, and further reducing the incorrect generation of the LLM model and improving the generation accuracy of the model;

[0045] Experimental results and explanations:

[0046]

[0047] The key point of this proposal lies in using prompt learning, constructing a binary classifier with BERT, and controlling the output of the LLM model after obtaining the results. In addition, the real-time EDA-SR data augmentation strategy in this scheme is also one of the reasons for improving the accuracy.

[0048] Prompt Learning is a natural language processing technique that guides the model to generate specific text or perform tasks by providing short prompt information. In NLP tasks, people usually define a specific task and train it with a specific dataset. This approach can not only improve the accuracy of the model but also reduce the data required for training. Prompt learning has achieved good results in many NLP tasks, such as question answering, language models, and translation tasks.

[0049] EDA (Easy Data Augmentation) is a very effective data augmentation technique. Synonym replacement can not only increase the size of the dataset but also improve the diversity of the data, enabling the model to better adapt to different scenarios. At the same time, synonym replacement can also effectively alleviate the problems of scarce vocabulary and low-frequency words in the dataset. This requires selecting an appropriate synonym generation method according to different tasks.

[0050] In natural language processing tasks, the BERT model has become a very mature and effective processing method. For multi-classification tasks, the LLM model is more efficient. However, large language models have the defect of generating randomly. To optimize this shortcoming, this invention adopts a prompt learning method based on BERT to control the output of the LLM. Compared with traditional multi-classification methods, it has the advantages of high availability, high scalability, and high accuracy.

[0051] 1. High availability: Well-designed prompt information can help the model converge faster, reduce the dependence on a large amount of data. In addition, the prompt information can also help the model generate text that is more in line with human language habits, improve the quality of the generated text, and thus improve the usability of the model.

[0052] 2. Scalability: The LLM has achieved remarkable results in many natural language processing tasks, can establish a complete language model, and has high language understanding and generation capabilities. In the future, the application fields of large language models will become more and more extensive. It can be foreseen that large language models will play an increasingly important role in fields such as human-computer interaction and natural language understanding.

[0053] 3. High accuracy: Compared with traditional multi-classification methods, the classification method of the present invention can greatly improve the accuracy rate compared with traditional BERT multi-classification and multiple binary classifier methods.

Claims

1. A sentiment analysis and classification method under a multi-class knowledge system, characterized in that, It includes the following steps: In step 1, both the training data and the validation data adopt the sst-5 public dataset, and the test data comes from Twitter user tweets and comments, which are obtained using Selenium crawlers. The training, validation, and test datasets are all stored in the format of [label, sentence], where label = {1, 2, 3, 4, 5} represents 5 kinds of emotions. label = 1 represents extremely negative emotion, label = 2 represents slightly negative emotion, label = 3 represents neutral emotion, label = 4 represents slightly positive emotion, and label = 5 represents extremely positive emotion. In step 2, data augmentation is performed on the training data, entity extraction is performed on nouns, and synonym replacement is implemented using the EDA-SR algorithm to expand the original dataset. Then, after preprocessing the data, the text dataset G is obtained. In step 3, for the LLM model, the text dataset G is used to construct a 5-classification dataset and encoded to obtain the encoded set E. For the BERT model, the text dataset G is used to construct a positive and negative 2-classification dataset and encoded to obtain the encoded set E'. In step 4, using the LLM model, the generative large language model is used to extract features from the encoded dataset E to obtain feature vectors, and through operations, 5 types of vectors for preliminary classification are obtained: llm_ans = LLM(E); In step 5, the BERT model is used to extract features from the encoded dataset E' to obtain feature vectors, and then the softmax operation is used to output the 2-classification result of positive and negative sentiment: pos_neg_ans = softmax(BERT(E')); In step 6, the 2-classification result obtained by the BERT model is encapsulated into the prompt template. In step 7, the LLM model is a generative dialogue model. The prompt template is embedded into the input of the LLM model to control the generation result and direction of the LLM model. The LLM model will generate dialogues according to the prompt information. In step 8, under the guidance of the prompt module, the llm_ans vector of the LLM model generates a 5-classification answer that meets the user's expectations. answer is the final classification result: answer = LLM(llm_ans + prompt) Among them, prompt can prompt the positive and negative information of semantics, thereby controlling the prediction direction of the model, and further reducing the incorrect generation of the LLM model and improving the generation accuracy of the model.

2. The emotional analysis and classification method under a multi-class knowledge system according to claim 1, characterized in that, The crawler in step 1 includes the following steps: 1) Start a real browser; 2) Open the Twitter website and log in to the website; 3) Control the browser to perform a series of operations through Python code to obtain the required information; 4) Process and store the captured data.

3. A sentiment analysis and classification method under a multi-class knowledge system according to claim 1, characterized in that, The crawler includes the following steps: The preprocessing process in step 2 includes the following steps: 1) Word segmentation, stop word removal, and conversion of English uppercase to lowercase; 2) For each sentence sample in the dataset, if the length is greater than d, truncate it; if the length is less than d, use the symbol <pad>Padding;< / pad> After the preprocessing is completed, the text dataset G is obtained, which is expressed as: G = [(label1, sentence1), (label2, sentence2),..., (label N , sentence N )] Among them, (label i , sentence i ) represents the i-th sample, where i ∈ N, and sentence i represents the i-th user comment, and label i is the sentiment category label of the corresponding text, and label i = {1, 2, 3, 4, 5}, and N represents the number of samples in the text dataset G.

4. A sentiment analysis and classification method under a multi-class knowledge system according to claim 1, characterized in that, The crawler includes the following steps: The template construction format in step 6 is as follows: pos_neg_ans = 0 represents negative, and the prompt template is as follows: pos_neg_ans = 1 indicates positive, and the prompt template is as follows:

Citation Information

Patent Citations

  • Text sentiment analysis method based on BERT model and double-channel attention

    CN110717334A

  • Text sentiment analysis method based on deep learning

    CN114595693A