Emotion analysis method based on AUC maximization

Through the sentiment analysis method based on AUC maximization, the BERT model is trained using the AUC loss function, which solves the problem of sentiment analysis performance degradation in the unbalanced training data set, and achieves robust sentiment analysis and improves classification effect.

CN120216696APending Publication Date: 2025-06-27胡伊伊
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510288698.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When the prior art deals with unbalanced training data sets, it is difficult to effectively learn fewer categories of features, affecting the performance of sentiment analysis.

Method used

The sentiment analysis method based on AUC maximization is adopted to train the BERT model through the AUC loss function to ensure the robustness of the model on the unbalanced dataset.

Benefits of technology

Robust sentiment analysis of the unbalanced training data set is realized, avoiding the prior distribution assumptions of the training data and the text categories to be analyzed, and improving the classification effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216696A_ABST
    Figure CN120216696A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of natural language processing, and particularly relates to an emotion analysis method based on AUC maximization. The method comprises the following steps: acquiring an unbalanced emotion corpus data set; word segmentation is carried out on the corpus, and word vectors are constructed; inputting the word vector of the training corpus into a BERT model to be trained; maximizing the training model through the AUC to obtain a trained BERT model; performing word segmentation on a to-be-analyzed text, and constructing word vectors; and inputting the word vector of the to-be-analyzed text into the trained BERT model, and obtaining an emotional tendency analysis result. Compared with the prior art, the method has the advantages that the problem of imbalance of training data is solved through AUC maximization, the robustness of the algorithm is improved, and the method has high practical value and practical significance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of natural language processing, and more particularly, relates to a sentiment analysis method based on maximizing AUC. Background Art

[0002] Sentiment analysis is an important issue in natural language processing and has important application values in fields such as public opinion analysis and e-commerce review analysis. For the text to be analyzed, the purpose of sentiment analysis is to judge the sentiment tendency conveyed by the text to be analyzed.

[0003] Early sentiment analysis relied on manually predefined sentiment lexicons and judged the sentiment tendency of the text by counting the occurrence frequencies of positive and negative sentiment words in the text. Sentiment analysis methods based on statistical learning use machine learning algorithms, such as support vector machines, to train on a pre-collected and labeled training data set, so as to obtain better classification performance. In particular, among the methods based on statistical learning, the methods based on deep learning use deep neural networks as classification models and have better performance.

[0004] The performance of statistical learning-based methods, including deep learning, depends on the training data. In some application scenarios, there will be a serious problem of class imbalance in the training data. For example, in e-commerce review analysis, there may be a situation where most reviews are positive. Imbalanced training data makes it difficult for the model to learn the features of the less represented classes, thus affecting its performance. Summary of the Invention

[0005] The technical problem solved by the present invention is: to overcome the deficiencies of the prior art and provide a sentiment analysis method based on maximizing AUC, which has good robustness for unbalanced training data sets.

[0006] The technical solution of the present invention is as follows: a sentiment analysis method based on maximizing AUC, the steps are as follows:

[0007] Step 1, obtain unbalanced expected training data according to actual needs;

[0008] Step 2, tokenize the training corpus and construct word vectors;

[0009] Step 3, train the model by maximizing AUC to obtain a trained BERT model;

[0010] Step 4, tokenize the text to be analyzed and construct word vectors;

[0011] Step 5, input the word vectors of the text to be analyzed into the trained BERT model and obtain the sentiment tendency analysis result.

[0012] Further, the specific steps of tokenizing the training corpus and constructing word vectors in step 2 are as follows:

[0013] Step 2.1, tokenize each piece of training corpus;

[0014] Step 2.2, use word2vec to obtain the word vector of each word and concatenate the word vectors of each corpus into a text matrix. The text matrix of the i-th corpus is denoted as X i ;

[0015] Further, the specific steps of training the model by maximizing AUC and obtaining the trained BERT model in step 3 are as follows:

[0016] Step 3.1, randomly collect a positive sample set and a negative sample set

[0017] Step 3.2, calculate the loss AUC loss function as follows;

[0018]

[0019] where w is the BERT model parameter vector, and h(X, w) represents the output obtained by inputting the discourse matrix X into the BERT model with parameter w;

[0020] Step 3.3, calculate the gradient through the backpropagation algorithm and update the model parameters through the gradient descent algorithm;

[0021] Step 3.4, if the training round is reached, stop training, otherwise go back to step 3.1;

[0022] Further, the specific steps of tokenizing the text to be analyzed and constructing word vectors in step 4 are as follows:

[0023] Step 4.1, tokenize the text to be analyzed;

[0024] Step 4.2, use word2vec to obtain the word vector of each word;

[0025] The advantages of the present invention compared with the prior art are as follows: The sentiment analysis of the present invention realizes a robust sentiment analysis function. Compared with traditional sentiment analysis methods, it does not require any assumptions about the prior distribution of training data and the categories of texts to be analyzed, and has a good classification effect for unbalanced training data. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] Combined with Figure 1, the sentiment analysis method based on AUC maximization of the present invention specifically comprises the following steps:

[0028] Step 1: Obtain unbalanced expected training data according to actual requirements;

[0029] Step 2: Segment the training corpus and construct word vectors;

[0030] Step 3: Train the model by maximizing AUC to obtain the trained BERT model;

[0031] Step 4: Segment the text to be analyzed and construct word vectors;

[0032] Step 5: Input the word vectors of the text to be analyzed into the trained BERT model and obtain the sentiment analysis result.

[0033] Further, the specific steps of segmenting the training corpus and constructing word vectors in Step 2 are as follows:

[0034] Step 2.1: Segment each piece of training corpus;

[0035] Step 2.2: Use word2vec to obtain the word vectors of each word and splice the word vectors of each corpus into a text matrix. The text matrix of the i-th corpus is denoted as X i ;

[0036] Further, the specific steps of training the model by maximizing AUC to obtain the trained BERT model in Step 3 are as follows:

[0037] Step 3.1: Randomly collect a positive sample set and a negative sample set

[0038] Step 3.2: Calculate the loss AUC loss function as follows;

[0039]

[0040] where w is the BERT model parameter vector, and h(X, w) represents the output obtained by inputting the discourse matrix X into the BERT model with parameter w;

[0041] Step 3.3: Calculate the gradient through the backpropagation algorithm and update the model parameters through the gradient descent algorithm;

[0042] Step 3.4: If the training round is reached, stop training; otherwise, go to Step 3.1;

[0043] Further, the specific steps of segmenting the text to be analyzed and constructing word vectors in Step 4 are as follows:

[0044] Step 4.1, perform word segmentation on the text to be analyzed;

[0045] Step 4.2, use word2vec to obtain the word vectors of each word;

[0046] Embodiment

[0047] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. Sentiment analysis method based on AUC maximization, characterized by Here are the steps: Step 1: Obtain unbalanced expected training data according to actual needs; Step 2: Segment the training corpus and construct word vectors; The step 2 comprises: Step 2.1, segment each training corpus; Step 2.2: Use word2vec to obtain the word vector of each word and concatenate the word vectors of each corpus into a text matrix. The text matrix of the i-th corpus is recorded as X i ; Step 3: Maximize the AUC to train the model and obtain the trained BERT model. Step 4: Segment the text to be analyzed and construct word vectors; The step 4 comprises: Step 4.1, segment the text to be analyzed; Step 4.2, use word2vec to obtain the word vector of each word; Step 5: Input the word vector of the text to be analyzed into the trained BERT model and obtain the sentiment tendency analysis result.

2. The sentiment analysis method based on AUC maximization of the calculation loss AUC loss function according to claim 1 is characterized in that Step 3, in which the model is trained by maximizing the AUC to obtain a trained BERT model, further includes the following steps: Step 3.1: Randomly collect positive sample sets from the training set And negative sample set Step 3.2, calculate the loss AUC loss function as follows; Where w is the BERT model parameter vector, h(X, w) represents the output obtained by inputting the discourse matrix X into the BERT model with parameter w; Step 3.3, calculate the gradient through the back-propagation algorithm, and update the model parameters through the gradient descent algorithm; In step 3.4, if the number of training rounds is reached, stop training; otherwise, go to step 3.1.