Text sentiment analysis method of LDA topic model based on BERT model and combined with K-means thought
The integration of BERT with K-means and LDA topic models addresses the limitations of existing sentiment analysis models by enhancing context capture and emotion classification accuracy.
Patent Information
- Application Number
- CN202510452559.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-15
AI Technical Summary
Existing text sentiment analysis models perform poorly in classification effects, especially when processing long texts, they are prone to losing context information or over-reliance on long-distance context, resulting in emotional classification bias.
Combining the BERT model and the LDA theme model of K-means ideas, the subject words are extracted through multiple iterations and spliced with the original text, the context understanding ability of BERT and the theme model of LDA are used to extract potential emotional topics, and sentiment analysis is performed by combining the SoftMax classifier.
It improves the accuracy and comprehensiveness of text sentiment analysis, and can better capture complex emotional expressions and enhance emotional polarity consistency.
Smart Images

Figure CN120316263A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly relates to a text sentiment analysis method based on a BERT model and an LDA topic model combined with the idea of K-means. Background Art
[0002] Text sentiment analysis is an important task in the field of natural language processing. By extracting the sentiment information in the text to judge the sentiment polarity of the text. In the past few years, with the development of deep learning technology, text sentiment analysis methods have made remarkable progress, but there are still some limitations.
[0003] A convolutional neural network (CNN) can automatically learn the most representative features from the input text, and processes the text through a convolutional layer to capture local sentiment information. However, it has deficiencies in capturing long-distance dependencies and context; recurrent neural networks (RNN) and their variants (LSTM, GRU) perform well in processing sequence data, can better capture the context information in the text, and model the sentiment information in long texts, making up for the deficiency of traditional methods in lacking context information. However, RNN mainly relies on unidirectional information transmission, and may lose some important context information. Especially when processing long texts, the sentiment information may be forgotten as the sequence length increases. Pre-trained language models such as BERT (Bidirectional Encoder Representations from Transformers) enable the model to capture rich syntactic and semantic information through large-scale unsupervised pre-training. The bidirectional encoding ability of BERT can more accurately understand the text context, thereby improving the accuracy of sentiment classification. However, some sentiment information may be concentrated in some short sentences or single words, and BERT may over-rely on long-distance context information, resulting in deviation in sentiment classification. Summary of the Invention
[0004] Object of the Invention: The object of the present invention is to provide a text sentiment analysis method based on a BERT model and an LDA topic model combined with the idea of K-means to solve the problem that the traditional text sentiment analysis model has poor classification effect.
[0005] Technical Solution: The text sentiment analysis method based on a BERT model and an LDA topic model combined with the idea of K-means described in the present invention includes the following steps:
[0006] (1) Preprocess the text data set, including word segmentation, normalization, denoising, and stop word filtering operations;
[0007] (2) Use the LDA topic model to obtain the topic distribution and word distribution of each text, extract the topic words with the highest weights, and splice them with the original text to generate a new dataset;
[0008] (3) Combining the K-means clustering idea, apply the LDA topic model to the new dataset again, iteratively extract topic words and repeat the splicing with the original text until the topic distribution converges;
[0009] (4) Input the text with the finally spliced topic words into the BERT model to generate context-aware word vectors;
[0010] (5) Output the text sentiment category through the SoftMax classifier.
[0011] Further, in step (1), the preprocessing includes: word segmentation, normalization, denoising, and stop word filtering operations; divide the dataset into a training set and a test set.
[0012] Further, in step (2), the construction of the LDA topic model includes: generating the document-topic distribution and topic-word distribution through the Dirichlet distribution; assigning a topic to each word according to the multinomial distribution and generating the corresponding word; extracting the set of topic words with the highest probability in each document and splicing it to the end of the original text.
[0013] Further, the formula for topic word extraction is:
[0014]
[0015] where k d is the index of the topic with the highest weight in the document dm, and θ m (k) represents the probability that the document dm belongs to the topic k.
[0016] Further, in step (3), the K-means idea is as follows: After each LDA topic extraction, perform K-means clustering on the generated new dataset to adjust the topic word distribution to optimize topic consistency; splice the clustered topic words with the original text for the second time and repeat the LDA topic extraction until the topic assignment is stable.
[0017] Further, in step (3), the iteration termination condition is: until the topic with the highest weight obtained by using the LDA topic model for topic extraction of each data in the text dataset no longer changes, the iteration ends.
[0018] Further, in step (4), the processing of the BERT model includes: performing sub-word tokenization on the spliced text to generate a Token sequence; fusing word embeddings, position embeddings, and segment embeddings to generate an input representation; capturing context dependencies through multiple layers of Transformer encoders and outputting the CLS token vector of the sentence as a classification feature.
[0019] Further, the self-attention calculation of the Transformer layer is as follows:
[0020]
[0021] where Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the key.
[0022] An electronic device according to the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements any one of the text sentiment analysis methods based on the BERT model and the LDA topic model combined with the K-means idea.
[0023] A storage medium according to the present invention stores a computer program. When the computer program is executed by a processor, it implements any one of the text sentiment analysis methods based on the BERT model and the LDA topic model combined with the K-means idea.
[0024] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: The present invention uses the pre-trained deep learning model BERT for sentiment analysis. BERT can capture the context dependencies in the text through a bidirectional Transformer architecture, can understand the multiple meanings of words, and effectively process complex sentiment expressions; The present invention uses the LDA topic model to extract potential sentiment topics from the text and combines it with the context understanding of BERT to help aggregate the sentiment information of different parts and obtain a more comprehensive sentiment analysis result; The present invention proposes to combine the K-means idea to use the topic model multiple times on the original LDA topic model for the text, which improves the matching degree between the text and the topic words as much as possible and enhances the sentiment polarity of the text. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is the research framework diagram of the present invention;
[0026] Figure 2 is the model framework diagram of the present invention. DETAILED DESCRIPTION
[0027] The technical solutions of the present invention will be further described below with reference to the accompanying drawings.
[0028] As Figure 1 shown, an embodiment of the present invention provides a text sentiment analysis method based on the BERT model and the LDA topic model combined with the K-means idea. The present invention is based on Figure 2 the model diagram shown, and includes the following steps:
[0029] Step 1, preprocess the text dataset. The preprocessing includes: word segmentation, normalization, noise removal, stop word filtering, etc. Divide the dataset into a training set and a test set;
[0030] Step 2, use the LDA topic model to obtain the topic and its word distribution of each text data, and splice the topic words belonging to the topic with the highest weight of each text data with the original text data to obtain a new text dataset containing topic words. The specific process includes:
[0031] Step 2-1, use the LDA topic model to obtain the topic and its word distribution of each text data;
[0032] θ m ~Dirichlet(α)
[0033] where θ is the topic distribution in document dm (i.e., the probability of each topic in the document), and α is the smoothing parameter of the topic, which controls the sparsity of the topic distribution in each document. For each document dm, draw the topic distribution m from a Dirichlet distribution;
[0034]
[0035] where, is the distribution of words in topic k, and β is the smoothing parameter of the vocabulary distribution, which controls the sparsity of the word distribution in each topic. For each topic k, draw the vocabulary distribution from a Dirichlet distribution
[0036] z m,n ~Multinomial(θ m )
[0037] where z m,n represents the topic of the nth word in document dm. Select a topic Z m from the topic distribution θ m,n .
[0038]
[0039] where w m,n represents the nth word in document dm. Given the selected topic z m,n , select the word w from the vocabulary distribution of this topic m,k .
[0040] Step 2-2, splice the topic words belonging to the topic with the highest weight of each text data with the original text data to obtain a new text dataset containing topic words;
[0041]
[0042] Among them, k d is the index of the topic with the highest weight in the document dm, and θ m (k) represents the probability that the document dm belongs to the topic k.
[0043]
[0044] Among them, w m,n 1 represents the topic word of the topic k with the highest weight in the document dm d .
[0045] dm 1 =(dm, w m,n 1 )
[0046] Among them, dm is the document, and dm 1 is the new text data set containing the topic words.
[0047] Step 3: Combine the K-means idea and use the LDA topic model again on the new text data set containing the topic words to extract topics, obtain the topics and their word distributions of each text data, and splice the topic words of the topic with the highest weight obtained with the original text data to obtain a text data set containing new topic words. The specific process includes:
[0048] Step 3-1: Combine the K-means idea and use the LDA topic model again on the new text data set containing the topic words to extract topics, obtain the topics and their word distributions of each text data;
[0049] θ m 1 ~Dirichlet(α)
[0050] Among them, θ m 1 is the topic distribution in the document dm 1 (i.e., the probability of each topic in the document), and α is the smoothing parameter of the topic, which controls the sparsity of the topic distribution in each document. For each document dm 1 , draw the topic distribution θ m 1 from a Dirichlet distribution;
[0051]
[0052] Among them, is the distribution of words in the topic k, and β is the smoothing parameter of the vocabulary distribution, which controls the sparsity of the word distribution in each topic. For each topic k, draw the vocabulary distribution from a Dirichlet distribution
[0053]
[0054] Among them, z m,n 1 represents the topic of the nth word in document d m 1 Select a topic z from the topic distribution θ of the document m 1 m,n 1 .
[0055]
[0056] Among them, w m,n 1 represents the nth word in document d m 1 Given the selected topic Select the word w from the vocabulary distribution of this topic m,n 1 .
[0057] Step 3-2: Concatenate the topic words of the topic with the highest weight obtained and the original text data to obtain a text data set containing new topic words;
[0058]
[0059] Among them, k d 1 is the index of the topic with the highest weight in document dm 1 , and θ m 1 (k) represents the probability that document dm 1 belongs to topic k.
[0060]
[0061] Among them, w m,n 2 represents the topic word of the topic k with the highest weight in document dm 1 d 1 .
[0062] dm 2 =(dm, w m,n 2 )
[0063] Among them, dm is the document, and dm 2 is the new text data set obtained by concatenating the topic words for the second time and the original text data.
[0064] Step 4: Iterate the operations in (3) multiple times until the topic with the highest weight obtained by using the LDA topic model for topic extraction of each data in the text dataset no longer changes, then the iteration ends, and the finally obtained topic words are concatenated with the original text dataset;
[0065]
[0066] where k d n is the index of the topic with the highest weight in the document dm n and θ m n (k) represents the probability that the document dm n belongs to topic k.
[0067]
[0068] where k d n+1 is the index of the topic with the highest weight in the document dm n+1 and θ m n+1 (k) represents the probability that the document dm n+1 belongs to topic k.
[0069] k d n = k d n+1
[0070] dm n = (dm, w m,n n )
[0071] where dm is the document, and dm n is the new dataset obtained by concatenating the finally determined topic words with the original text data, and w m,n n represents the topic word of the topic with the highest weight in the document dm n-1 in the document dm.
[0072] Step 5: Input the finally obtained text dataset with topic words into the BERT model for word vector training to obtain text word vectors containing topic information. The specific process includes:
[0073] Step 5-1: The input of the BERT model will first be tokenized, splitting the sentence into words or word-characters, and each word or word-character will be mapped to a token;
[0074] Step 5-2, for the input tokens, BERT generates corresponding embedding representations. BERT uses three different embeddings: Token Embeddings, Segment Embeddings, and Position Embeddings;
[0075] Step 5-3, the embedding representations are processed through multiple Transformer layers of BERT. Each layer includes multiple Self-Attention mechanisms to help the model capture context information. The output of each layer updates the token representation. After multiple Transformer layers, the model can obtain context-aware token representations;
[0076]
[0077] where Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the key, used for scaling.
[0078] Step 6, use the SoftMax classifier to obtain the sentiment category of the text;
[0079]
[0080] where P(y = j|x) is the predicted probability of class j, and z j is the logits (the raw scores output by the model) of the input sample j, K is the total number of classes, is the sum of the exponentials of all class logits, ensuring that the output is a valid probability distribution.
Claims
1. A text sentiment analysis method based on the BERT model and the LDA topic model combined with the K-means idea, characterized in that, It includes the following steps: (1) Preprocess the text dataset; (2) Use the LDA topic model to obtain the topic distribution and word distribution of each text, extract the topic words with the highest weights and splice them with the original text to generate a new dataset; (3) Combining the K-means clustering idea, apply the LDA topic model to the new dataset again, iteratively extract topic words and repeat splicing with the original text until the topic distribution converges; (4) Input the text with the finally spliced topic words into the BERT model to generate context-aware word vectors; (5) Output the text sentiment category through the SoftMax classifier.
2. The text sentiment analysis method of an LDA topic model based on the BERT model and combined with the K-means idea according to claim 1, characterized in that, In step (1), the preprocessing includes: word segmentation, normalization, denoising, and stop word filtering operations; divide the dataset into a training set and a test set.
3. The text sentiment analysis method of an LDA topic model based on the BERT model and combined with the K-means idea according to claim 1, characterized in that, In step (2), the construction of the LDA topic model includes: generating the document topic distribution and topic word distribution through the Dirichlet distribution; assigning a topic to each word according to the multinomial distribution and generating the corresponding word; extracting the set of topic words with the highest probability in each document and splicing it to the end of the original text.
4. The text sentiment analysis method of an LDA topic model based on the BERT model and combined with the K-means idea according to claim 3, characterized in that, The formula for topic word extraction is: where k d is the index of the topic with the highest weight in the document dm, and θ m (k) represents the probability that the document dm belongs to the topic k.
5. The text sentiment analysis method of an LDA topic model based on the BERT model and combined with the K-means idea according to claim 1, characterized in that, In step (3), the K-means idea is as follows: After each LDA topic extraction, perform K-means clustering on the generated new dataset, adjust the topic word distribution to optimize the topic consistency; splice the clustered topic words with the original text twice, and repeat the LDA topic extraction until the topic assignment is stable.
6. The text sentiment analysis method of an LDA topic model based on the BERT model and combined with the K-means idea according to claim 1, characterized in that In step (3), the iteration termination condition is: until the topic with the highest weight obtained by performing topic extraction on each data in the text dataset using the LDA topic model no longer changes, the iteration ends.
7. The text sentiment analysis method of an LDA topic model based on the BERT model and combined with the K-means idea according to claim 1, characterized in that In step (4), the processing of the BERT model includes: performing sub-word tokenization on the spliced text to generate a Token sequence; fusing word embeddings, position embeddings, and segment embeddings to generate an input representation; capturing context dependencies through multiple layers of Transformer encoders and outputting the CLS token vector of the sentence as the classification feature.
8. The text sentiment analysis method of an LDA topic model based on the BERT model and combined with the K-means idea according to claim 7, characterized in that, The self-attention calculation of the Transformer layer is: where Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the key.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is loaded into the processor, it implements a text sentiment analysis method based on the BERT model and the LDA topic model combined with the K-means idea according to any one of claims 1-8.
10. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements a text sentiment analysis method based on the BERT model and the LDA topic model combined with the K-means idea according to any one of claims 1-8.