Text Sentiment Recognition Method Based on Maximum Probability Filling and Multi-Head Attention Mechanism
Through the combination method of LDA maximum probability filling and multi-head attention mechanism, the problem of sparse short text data and long text information extraction in text emotion classification is solved, and the accuracy and speed of classification are improved.
Patent Information
- Application Number
- CN202210447939.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-26
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-04-26
AI Technical Summary
The prior art is difficult to effectively handle the sparseness of short text data and the information extraction of long texts in text emotion classification, resulting in sparseness of data and semantic confusion, affecting the classification accuracy.
The text emotion classification method based on LDA maximum probability fill and multi-head attention mechanism is adopted to process the sparseness of short text data through LDA maximum probability fill, and the global features of long text are extracted using the multi-head attention mechanism.
It effectively solves the difficulty of sparse short text data and extracts long text information, improves the accuracy and speed of text emotional classification, and adapts to flexible text length.
Smart Images

Figure CN114742047B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a text sentiment classification method based on LDA maximum probability filling and multi-head attention mechanism, belonging to the technical field of text data recognition. Background Art
[0002] With the development of Internet technology and the advent of the big data era, people are more inclined to express their true opinions on online platforms, and the quantity of sentiment information such as instant chats, current affairs comments, product reviews, etc. has increased rapidly. The sentiment classification of these texts has high application value and can provide support for directions such as decision-making, public opinion guidance, and recommendation systems. At the same time, these online platform texts will present characteristics such as flexible length and simple grammatical structure, which greatly increases the difficulty of discrimination. Therefore, obtaining short text feature information and improving the accuracy of sentiment tendency discrimination is one of the widely concerned issues in the NLP field.
[0003] There are mainly two text sentiment classification methods in related technologies. One is to extract features based on machine learning methods to complete the sentiment classification task. Among them, classifiers such as Naive Bayes, Support Vector Machine (SVM), Latent Dirichlet Allocation (LDA), etc. are used to complete the classification. However, these classifiers often lose the associated information of the text context when extracting features, resulting in the loss of some information. And the classification accuracy often depends on a large-scale high-quality labeled training set, and these data require a high labor cost.
[0004] The second is to automatically extract features based on deep learning methods. Neural networks can better complete the classification task. For example, the Long Short-Term Memory network (LSTM) can fully consider the sequence relationship of the input text, and the Convolutional Neural Network (CNN) can automatically extract deeper features. In recent years, technicians have proposed an attention mechanism with better text feature learning performance, which has also been gradually widely applied to text classification. However, due to the different lengths of the input texts, when performing batch operations, the matrix needs to be filled to make the lengths consistent. Mainstream filling methods such as the zero-padding method and the circular method will cause problems of data sparsity and semantic confusion. Summary of the Invention
[0005] Object of the Invention: Aiming at the problems and deficiencies existing in the prior art, the present invention provides a text sentiment classification method based on LDA maximum probability filling and multi-head attention mechanism, which is a text sentiment recognition method that can adapt to flexible text lengths. This method can not only fully solve the problems of short text data sparsity and filling, but also cope with the difficulty of long-distance information extraction, and ensure the speed and accuracy of recognition.
[0006] Technical solution: A text sentiment classification method based on LDA maximum probability filling and multi-head attention mechanism, which obtains a dataset and performs text preprocessing; obtains a dictionary and a word vector matrix through the Word2Vec model; fills the word vector matrix using the LDA maximum probability method; mines the local features of the text through the convolution operation of the Text-CNN network, performs weight redistribution through the multi-head attention mechanism, and finally obtains the classification probability of the sentiment polarity by the Softmax function. It includes the following steps:
[0007] Step 1: Obtain a text dataset and perform text preprocessing;
[0008] Step 2: Input the dataset into the word2vec model in the specified format to train word vectors, and obtain a word-word vector mapping dictionary and the corresponding word vector matrix A of the dataset;
[0009] Step 3: Input the dataset into the LDA topic model in the specified format, select the optimal number of topic parameter values K, and obtain the probability distribution of text-topic-word;
[0010] Step 4: Based on the longest text length in the dataset, use the LDA filling method under the maximum probability topic to fill the word vector matrix A to obtain a word vector matrix B with the same length;
[0011] Step 5: Input the word vector matrix B into the Text-CNN model to extract local context features; through convolution and pooling operations, obtain the global feature vector T;
[0012] Step 6: Use the Multi-Head Attention model, take the global feature vector T as the input, introduce the multi-head attention mechanism, and obtain the feature vector G composed of the splicing of multiple subspaces;
[0013] Step 7: Use a fully connected network and a SoftMax classifier. The SoftMax classifier is a model for identifying the positive and negative emotions of text. Take the feature vector G as the input, perform sentiment classification, obtain the classification probability of the sentiment polarity and output it.
[0014] The specific steps of performing text preprocessing in the above Step 1 include:
[0015] Step S11: Cleaning, removing uncommon Chinese symbols in the text and converting traditional and simplified Chinese;
[0016] Step S12: Word segmentation, dividing the text into combinations of words and separating them with delimiters;
[0017] Step S13: Stopword removal, removing stopwords and filtering out meaningless and less influential words on the results;
[0018] Step S14: Truncate the text whose length exceeds the set length.
[0019] In step 2, the dataset is input into the word2vec model to train word vectors. The skip-gram algorithm, which is a three-layer small neural network, is used to generate word vectors. Parameters such as the prediction context window size and word vector dimension are set. The words in the dataset are formed into a vocabulary and converted into one-hot encoding. In the hidden layer, the weights of the hidden layer are learned by predicting the probability of the context words appearing in the given window by the central word, and finally, a word-word vector mapping dictionary with a given dimension is obtained. The dataset is mapped into a word vector matrix B through the dictionary.
[0020] In step 3, the specific steps for the dataset to be input into the LDA model to obtain the probability distribution of text-topic-word are as follows:
[0021] Step S31: Set the initial number of topics K and other initial parameters;
[0022] Step S32: Input the dataset into the LDA model to obtain the preliminary topic probability distribution of the document and the vocabulary distribution of the topic
[0023] Step S33: Select the optimal number of model topics, the optimal topic value K, by calculating the perplexity. 0 It is obtained at the inflection point where the perplexity decreases.
[0024] Step S34: Reset the number of topics to K 0 , and input the dataset into the LDA model again to obtain the final document-topic probability distribution and the word probability distribution under each topic.
[0025] In step 4, based on the longest text length in the dataset, the specific steps for filling the word vector matrix using the LDA filling method under the topic with the highest probability are as follows:
[0026] Step S41: Find the longest text length L in all documents in the dataset max , as the benchmark length of the word vector;
[0027] Step S42: Perform a filling operation on the documents in the dataset where each text length L is less than L max :
[0028] Step S43: Find the topic with the highest probability topic i in the document corresponding document-topic matrix
[0029] Step S44: Through the word probability distribution of topic i, select the first L max - L words in order from largest to smallest according to the word probability;
[0030] Step S45: Map the L max -L word tokens into L max -L n-dimensional word vectors using the dictionary obtained through training in Step 3;
[0031] Step S46: Use the word vectors to fill the word vector matrix in sequence until the length of the word vectors of the current document is equal to L max .
[0032] Step S47: Repeat S43 - S46 until the lengths of all documents are L max ;
[0033] Step S48: An equi - length word vector matrix is obtained.
[0034] Step 5 adopts the Text - CNN model, and its main feature is to perform convolution operations on the input text data. The specific steps in the Text - CNN model are as follows:
[0035] Step S51: Input the word vector matrix with the same length into the convolution layer. Use multiple shared convolution kernels and receptive fields to perform convolution operations, extract local features, and perform non - linear operations through the activation function to obtain the feature matrix;
[0036] Step S52: The feature matrix passes through the pooling layer. Under the action of max - pooling, select the maximum value of the feature matrix, concatenate it with the maximum values of other channels, and combine them into the global feature vector T.
[0037] In Step 6, the Multi - Head Attention model is used, which introduces the multi - head attention mechanism. It is an improvement of the attention mechanism and can better capture key features in different aspects and long - distance dependencies. The specific steps within the Multi - Head Attention model are as follows:
[0038] Step S61: Use the global feature vector T as the input of the Multi - Head Attention model to obtain the query matrix Q, key matrix K, and value matrix V respectively:
[0039] Q = T * W i Q
[0040] K = T * W i K
[0041] V = T * W i V
[0042] where W i Q , W i K , Wi V They are the weight matrices of the query matrix Q, the key matrix K, and the value matrix V respectively.
[0043] Step S62: Multiply the query matrix Q by the transposed matrix K of the key matrix K, calculate the dot product to obtain scores, and use the SoftMax function in the self-attention mechanism to calculate the similarity scores between each q in the query matrix Q and each v in the value matrix V, and perform weighted summation. Finally, obtain the self-attention sequence head of a single head: T i i i i head i :
[0044]
[0045] head i = A i (Q, K, y)
[0046] where is the scaling factor, which plays a role in adjustment;
[0047] Step S63: The above steps obtain the operation result of a single-head self-attention. By changing the weights of the weight matrix and repeating the operation multiple times, multiple single-head operation result matrices can be horizontally concatenated to obtain the operation result of multi-head self-attention:
[0048] MA(Q, K, V) = Concat(head 1 , …, head l )W o .
[0049] where, W 0 is the additional weight matrix;
[0050] In step 7, a fully connected network and the SoftMax function are used for classification to predict the sentiment tendency classification probability of the text and output the result. Brief Description of the Drawings
[0051] Figure 1 is the flowchart of the method in the embodiment of the present invention;
[0052] Figure 2 is the flowchart of the maximum probability LDA filling method in the embodiment of the present invention;
[0053] Figure 3 is the multi-head attention mechanism model diagram in the embodiment of the present invention. Detailed Embodiments
[0054] The present invention will be further illustrated below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications made by those skilled in the art fall within the scope defined by the appended claims of this application.
[0055] As Figure 1 shown, a text sentiment classification method based on LDA maximum probability filling and multi-head attention mechanism includes the following steps:
[0056] Step S10: Obtain a text dataset and perform text preprocessing. The dataset is divided into three parts: a training set, a validation set, and a test set in a ratio of 8:1:1. Each document has a corresponding sentiment polarity label: 0 or 1, representing positive emotion and negative emotion respectively; the positive and negative samples are evenly distributed. Text preprocessing can make the text conform to the input format and eliminate factors that have little impact on the results. The specific steps of step S10 include S11 - S14:
[0057] Step S11: Clean, remove special symbols in the text, such as line breaks and redundant whitespace, and convert traditional and simplified Chinese characters.
[0058] Step S12: Tokenize, use a tokenization tool to convert the sentence into a combination of words, and separate them with delimiters.
[0059] Step S13: Stopword removal, filter out meaningless function words and words that have little impact on the results according to the stopword list, such as conjunctions in the text.
[0060] Step S14: Truncate, truncate the text whose length exceeds the set length, and truncate the overly long text content to avoid the word vector matrix generated in subsequent steps being too large and affecting the training speed and effect.
[0061] Step S20: Input the dataset into the Word2Vec model according to the format to train word vectors, perform feature mapping, and obtain a word-word vector mapping dictionary and the corresponding word vector matrix A of the dataset. The Word2Vec model represents the semantic information of words in the form of word vectors by learning the text. In this embodiment, the Word2Vec model adopts the skip-gram algorithm.
[0062] The specific steps of step S20 include S21 - S25:
[0063] Step S21: Set parameters such as the window size of the predicted context and the dimension of the word vector, and input the dataset according to the format.
[0064] Step S22: In the input layer, the model first forms a vocabulary from the words in the dataset and uses the one-hot word representation method to convert the document into a vector.
[0065] Step S23: In the hidden layer, the model adopts the skip-gram algorithm, that is, it learns the hidden layer weights by predicting the probability of the context words appearing in a preset window given the center word.
[0066] Step S24: Save the model to obtain a word-word vector mapping dictionary with a given dimension.
[0067] Step S25: Use the model to generate word vector matrix embedding for the text in the dataset to obtain a word vector matrix mapped from the dataset.
[0068] Step S30: Input the dataset into the LDA topic model according to the format, select the optimal value of the topic number parameter K, and obtain the probability distribution of text-topic-word.
[0069] Among them, the LDA topic model can be used to infer the topic distribution of documents. Through unsupervised learning, it gives the topic of each document in the document set in the form of a probability distribution, and can perform topic clustering or text classification according to the topic distribution.
[0070] The specific steps of Step S30 include S31 - S33:
[0071] Step S31: Set the initial value of the topic number K and other initial parameters of the model.
[0072] Step S32: Input the dataset into the LDA model, and the preliminary topic probability distribution of the document and the vocabulary distribution of the topic can be obtained.
[0073] Step S33: Select the optimal number of model topics by calculating the perplexity. The calculation of the training perplexity is as follows:
[0074]
[0075] where M represents the number of texts in the training set, N d represents the size of the d-th document, and p(w d ) represents the text probability, that is, the product of the distribution values of the word in all topics and the topic probability distribution of the text where the word is located.
[0076] As the value of K is gradually increased, the training perplexity shows a curve that first decreases and then increases, and the optimal value of K 0 is obtained at the inflection point where the training perplexity decreases.
[0077] Step S34: Reset the number of topics to K 0 , input the dataset into the LDA topic model again, and after training, obtain the final document-topic probability distribution and the word probability distribution under each topic.
[0078] Step S40: Based on the longest text length in the dataset, use the LDA filling method under the maximum probability topic to fill the word vector matrix A, and obtain a word vector matrix B with the same length.
[0079] Among them, the word vector matrix is mapped by the word vector model, and there is a problem of different matrix lengths due to different text lengths in the document. In the current mainstream processing methods, the zero-padding method and the circular method are usually used for filling, resulting in problems of sparsity and semantic confusion in the word vector matrix. Using the maximum probability topic LDA filling method not only makes it more convenient to batch process data but also solves the problem of sparse short text data and enriches the text feature information.
[0080] Such as Figure 2 shown, the specific steps of step S40 include S41 - S48:
[0081] Step S41: Find the longest text length L in all documents in the dataset max , as the benchmark length of the word vector.
[0082] Step S42: Perform a filling operation for each document in the document set with a text length L less than L max :
[0083] Step S43: Find the maximum probability topic topic i in the document - topic matrix corresponding to the document;
[0084] Step S44: Through the word probability distribution of topic i, select the top L max - L words in order from largest to smallest word probability;
[0085] Step S45: Map the L max - L words into L max - L n - dimensional word vectors through the dictionary trained in step 3;
[0086] Step S46: Use the word vectors to fill the word vector matrix in turn until the word vector length of this document is equal to L max ;
[0087] Step S47: Repeat S43 - S46 until the length of all documents is L max ;
[0088] Step S48: Obtain an equi - length word vector matrix;
[0089] For example, currently L max is 10, and the length of the word matrix corresponding to the document to be operated on currently is only 8, so two lengths of word vectors need to be filled. Find the topic with the maximum probability corresponding to this document as topic i, arrange the subject words under this subject in ascending order of probability, which are "criticism", "school", "examination" respectively. Then select "criticism" and "school", obtain the corresponding word vectors through the mapping dictionary, and fill them into the word vector matrix in turn. At this time, the corresponding length of the document reaches L max .
[0090] For example, currently L max is 10, and the length of the word matrix corresponding to the document to be operated currently is only 8. Then two lengths of word vectors need to be filled. Find that the topic with the highest probability corresponding to this document is topic i , arrange the subject words under this subject in ascending order of probability, which are "criticism", "school", "examination" respectively. Then select "criticism" and "school", obtain the corresponding word vectors through the mapping dictionary, and fill them into the word vector matrix in turn. At this time, the corresponding length of the document reaches L max .
[0091] Step S51: At the input layer, input the word vector matrix A with the same length, and enter the convolutional layer; in the convolutional layer, use multiple shared convolutional kernels to perform convolutional operations with the receptive field to extract local feature information T i , as shown in the following formula:
[0092] T i = f(h·A i:i+n-l + b)
[0093] where n is the number of words corresponding to the convolutional kernel, h is the weight matrix, b is the bias value, and A i:i+n-1 is the sub-matrix obtained by intercepting the i-th to i + n - 1-th rows from the word vector matrix A. f() is the activation function. In this embodiment, the ReLU function is used for activation, and through non-linear operations, the feature matrix is obtained.
[0094] Step S52: In the pooling layer, select the maximum value of the feature matrix T i under the action of max-pooling, splice it with the maximum values of other channels to form a global feature vector T, and output:
[0095] T i = max{T i}
[0096] T = [T 1 , T 2 , T 3 , …, T j-n+1 .
[0097] Step S60: Use the Multi-Head Attention model, take the global feature vector T as the input, introduce the multi-head attention mechanism, and obtain the feature vector G composed of the splicing of multiple sub-spaces.
[0098] Among them, the attention mechanism is an imitation of human attention. In the field of NLP, it mainly calculates the attention of words in the text. The larger the value, the greater the role played in the task. The multi-head attention mechanism is an improvement of the attention mechanism, which can better capture key features and long-distance dependencies in different aspects.
[0099] Such as Figure 3 As shown, the specific steps of step S60 include S61 - S63:
[0100] Step S61: The global feature vector T is used as the input of the Multi-Head Attention model, and the query matrix Q, key matrix K, and value matrix V are obtained respectively:
[0101] Q = T * W i Q
[0102] K = T * W i K
[0103] y = T * W i V
[0104] Among them, W i Q , W i K , W i V are the linear transformation weight matrices of the query matrix Q, key matrix K, and value matrix V respectively.
[0105] Step S62: The query matrix Q and the key matrix K are dot-product calculated to obtain scores, and the SoftMax function in the self-attention mechanism is used to calculate the similarity scores between each q i in the query matrix Q and each v i in the value matrix V, and then weighted and summed. Finally, a single-head self-attention sequence head i is obtained:
[0106]
[0107] head i = A i (Q, K, V)
[0108] Among them, is the scaling factor, which plays a role in adjustment;
[0109] Step S63: The above steps obtain the result of a single-head self-attention operation after 1 time. In the present invention, 8 attention heads are set. By changing w iThe weight is used to repeat the single-head operation 8 times, and the result matrices of the 8 single-head operations are horizontally concatenated, followed by a linear operation to obtain the result of the multi-head self-attention operation:
[0110] MA(Q, K, V) = Concat(head 1 , …, head i )W o
[0111] where W o is the additional weight matrix.
[0112] In step S70, a fully connected network and the Softmax function are used for classification to predict the sentiment classification probability of the text and output the result. The calculation formula is as follows:
[0113] y = SoftMax(w MA MA + b MA )
[0114] where w MA is the weight coefficient, and b MA is the bias term. Through Softmax classification, the probability distribution of each text in each category is obtained, and the category with the maximum value is the predicted category. In this embodiment, it is a two-polar classification, that is, a binary classification task.
Claims
1. A text sentiment classification method based on LDA maximum probability filling and multi-head attention mechanism, characterized in that, it includes the following steps: Step 1: Obtain a text dataset and perform text preprocessing; Step 2: Input the dataset into the word2vec model according to the format to train word vectors, and obtain a word-word vector mapping dictionary and a word vector matrix A corresponding to the dataset; Step 3: Input the dataset into the LDA topic model according to the format, select the optimal number of topic parameter values K, and obtain the probability distribution of text-topic-word; Step 4: Based on the longest text length in the dataset, use the LDA filling method under the maximum probability topic to fill the word vector matrix A to obtain a word vector matrix B with the same length; Step 5: Input the word vector matrix B into the Text-CNN model to extract local context features; through convolution and pooling operations, obtain the global feature vector T; Step 6: Use the Multi-Head Attention model, take the global feature vector T as the input, introduce the multi-head attention mechanism, and obtain a feature vector G composed of multiple subspaces spliced together; Step 7: Use a fully connected network and a SoftMax classifier, take the feature vector G as the input, perform classification, obtain the classification probability of the sentiment polarity and output it; In the said Step 4, based on the longest text length in the dataset, the specific steps of using the LDA filling method of the word vector matrix under the maximum probability topic are as follows: Step S41: Find the longest text length L among all the documents in the dataset max , which is used as the reference length of the word vector. Step S42: Perform a padding operation on the documents in the dataset where the length L of each text is less than L max : Step S43: Find the maximum probability topic topic i in the document corresponding document-topic matrix, Step S44: According to the word probability distribution of topic i, select the top L max -L words in descending order of word probability; Step S45: Map the L max -L words into L max -L n-dimensional word vectors through the dictionary trained in Step 3; Step S46: Use the word vectors to fill the word vector matrix in sequence until the word vector length of the current document is equal to L max ; Step S47: Repeat S43 - S46 until the lengths of all documents are L max ; Step S48: Obtain an equi-length word vector matrix; The said Step 5 adopts the Text-CNN model, and the main feature is to perform convolution operations on the input text data; the specific steps in the Text-CNN model are as follows: Step S51: Input the word vector matrix with the same length into the convolutional layer; use multiple shared convolutional kernels and receptive fields to perform convolution operations, extract local features, and perform non-linear operations through an activation function to obtain a feature matrix; Step S52: The feature matrix passes through the pooling layer, and under the action of max pooling, the maximum value of the feature matrix is selected and spliced with the maximum values of other channels to form the global feature vector T.
2. The text sentiment classification method based on LDA maximum probability filling and multi-head attention mechanism according to claim 1, characterized in that, the specific steps of performing text preprocessing in the said Step 1 include: Step S11: Text cleaning; Step S12: Word segmentation, divide the text into a combination of words, and use delimiters to separate; Step S13: Stop words removal, remove stop words; Step S14: Truncation, truncate the text with a length exceeding the set length.
3. The text sentiment classification method based on LDA maximum probability filling and multi-head attention mechanism according to claim 1, characterized in that, In step 2, the data set is input into the word2vec model to train word vectors. The skip-gram algorithm, which is a small three-layer neural network, is used to generate word vectors. Set the size of the prediction context window and the dimension of the word vectors. The words in the data set are formed into a vocabulary and converted into one-hot encoding. In the hidden layer, the weights of the hidden layer are learned by predicting the probability of the context words appearing in the given window by the central word. Finally, a word-word vector mapping dictionary with a given dimension is obtained. The data set is mapped into a word vector matrix B through the dictionary.
4. The text sentiment classification method based on LDA maximum probability filling and multi-head attention mechanism according to claim 1, characterized in that, the specific steps of inputting the data set into the LDA model in step 3 to obtain the probability distribution of text-topic-words are as follows: Step S31: Set the initial number of topics K and other initial parameters; Step S32: Input the data set into the LDA model to obtain the initial topic probability distribution of the document and the vocabulary distribution of the topic; Step S33: Select the optimal number of model topics, the optimal topic value K, by calculating the perplexity 0 It is obtained at the inflection point where the perplexity decreases; Step S34: Reset the number of topics to K 0 , and input the dataset into the LDA model again to obtain the final document-topic probability distribution and the word probability distribution under each topic.
5. The text sentiment classification method based on LDA maximum probability filling and multi-head attention mechanism according to claim 1, characterized in that, using the Multi-Head Attention model in step 6 and introducing the multi-head attention mechanism is a refinement of the attention mechanism and can better capture key features and long-distance dependence relationships in different aspects. The specific steps within the Multi-Head Attention model are as follows: Step S61: The global feature vector T is used as the input of the Multi-Head Attention model to obtain the query matrix Q, the key matrix K, and the value matrix V respectively: Q = T * W i Q Among which W i Q , are the weight matrices of the query matrix Q, the key matrix K, and the value matrix V, respectively; Step S62: Multiply the query matrix Q by the transposed matrix K of the key matrix K T to calculate scores through dot product, and use the SoftMax function in the self-attention mechanism to calculate the similarity scores between each q in the query matrix Q i and each v in the value matrix V i , and perform weighted summation; finally obtain the self-attention sequence head of a single head i : head i = A i (Q, K, V) Among them is a scaling factor that plays a role in adjustment; Step S63: The above steps obtain a single-head self-attention operation result. By changing the weights of the weight matrix and repeating the operation multiple times, multiple single-head operation result matrices can be horizontally concatenated to obtain the result of the multi-head self-attention operation: MA(Q, K, V) = Concat(head 1 , …, head l )W o ; Among them, W 0 is the additional weight matrix.
6. The text sentiment classification method based on LDA maximum probability filling and multi-head attention mechanism according to claim 1, characterized in that, in step 7, a fully connected network and the SoftMax function are used for classification to predict the sentiment tendency classification probability of the text and output the result.
Citation Information
Patent Citations
Power grid equipment defect text classification method based on multi-head attention mechanism and RCNN network
CN112199496A
Text sentiment classification method
CN112818123A