Long Text Classification Method Based on TextRank and Attention Mechanism
By introducing TextRank and attention mechanism into the text classification method, key sentences and feature information of long texts are extracted, which solves the problem of long text classification in the existing technology, and improves the classification accuracy and feature extraction ability.
Patent Information
- Application Number
- CN202211280953.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-19
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-10-19
AI Technical Summary
When existing text classification methods deal with long text, it is difficult to effectively extract feature information and key content, resulting in performance differences in classification models in the context of long text and short text.
The long text classification method based on TextRank and attention mechanism is adopted to calculate the key sentence sequence and keyword sequence of the long text through the TextRank layer, and feature information is extracted in combination with the BiGRU layer, and the key sentences are used as query vectors of the attention mechanism to calculate the attention score of the text, so as to pay more attention to parts similar to the semantics of the key sentences.
It improves the accuracy of long text classification tasks, can extract key feature information in long text more effectively, and reduces the performance differences of classification models in long text and short text contexts.
Smart Images

Figure CN115599915B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of long text feature extraction, and particularly relates to a long text classification method based on TextRank and attention mechanism. Background Art
[0002] The text classification task can be divided into short text classification and long text classification according to the text length. Compared with short text classification, the difficulty of long text classification lies in the extraction of feature information of longer sequences and the division of key content. Existing text classification methods do not make good improvements in methods for long texts, and do not fully consider the differences between long texts and short texts during application, which will lead to differences in the performance of classification models in long text contexts and short text contexts.
[0003] For example, a classification method combining multi-scale convolutional attention and GRU (Gated Recurrent Unit) is proposed in the literature. Although this classification method has achieved good classification performance, the datasets used in its experiments are all short text datasets, and the average text length of the longest dataset is only 45. The literature uses the SRU (Simple Recurrent Unit) and Attention methods to extract feature information, and ordinary Attention cannot fully extract the key feature information in long texts. Summary of the Invention
[0004] In order to overcome the above technical problems, the purpose of the present invention is to provide a long text classification method based on TextRank and attention mechanism. This method is applicable to the topic classification and sentiment analysis of long texts. For longer texts, it will trim the text according to the importance of words in the text, improving the quality of each paragraph of text. Secondly, this method will extract the key sentences of the current text as the query vector of the attention mechanism, and calculate the attention score of the text according to the key sentence vector, making the model pay more attention to the part semantically similar to the key sentence.
[0005] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0006] A long text classification method based on TextRank and attention mechanism, comprising the following steps;
[0007] Step1: Input the long text sequence into the TextRank layer to calculate the key sentence sequence and keyword sequence of the long text. The key sentence sequence and keyword sequence are sorted according to the weights. The closer the weight is to 1, the more important it is. Select the sentence with the weight closest to 1 in the key sentence sequence as the key sentence of the text. Perform data preprocessing operations on the long text sequence, crop or pad each text according to the set unified sample length. For longer texts, crop the keywords with lower weights, and for shorter texts, fill the tail with keywords with higher weights;
[0008] Step2: Input the text sequence processed by the TextRank layer into the Word Embedding layer to generate word vector representations;
[0009] Step3: Input the long text vector into the BiGRU layer. BiGRU will extract its feature information by combining the context of the text;
[0010] Step4: Calculate the attention of the text vector in combination with the key sentence of the text, obtain the attention scores corresponding to the key sentences in the text vector, and update the text feature vector according to the attention scores;
[0011] Step5: Input the updated text feature vector into the Linear and Softmax layers to obtain the classification result.
[0012] The TextRank uses a graph network to generate weighted graph nodes for each word. If two words are in a co-occurrence window, an edge is established between the two word nodes. During training, the weights of each node are continuously iteratively updated. The update formula for the weights of each node is as follows:
[0013]
[0014] Among them, WS{V i}, WS{V j} represent the weight values of the i-th word and the j-th word; V i , V j represent the nodes of the i-th word and the j-th word in the graph; InV i , OutV j respectively represent the in-degree set of V i and the out-degree set of V j ; d is the damping coefficient, usually set to 0.85, indicating that the probability of this point pointing to another node is 85%.
[0015] The key sentence sequence of the TextRank is based on the similarity between sentences. By constructing a weighted graph at the sentence level, the similarity weights between sentence nodes are updated, and then the key sentence sequence is arranged according to the similarity scores of each sentence. The similarity calculation formula between each sentence node is as follows:
[0016]
[0017] Where S i , S j is a two-sentence node, w k is the word between the two sentences, and the entire formula (2) is used to calculate the content repetition between the two sentences.
[0018] The processing steps of the TextRank layer are as follows:
[0019] Step 1: The first step is data preprocessing. The long text sequence is input into the TextRank layer, and the word segmentation is performed through the stop word list to filter out irrelevant words.
[0020] Step 2: Update the weight of each word node according to formula (1), sort each word into a keyword sequence according to the weight value, divide the sentence according to any punctuation mark indicating the end of the sentence in the long text, and calculate the key sentence according to formula (2);
[0021] Step 3: According to the set uniform text length, for longer texts, delete the less important keywords, and for shorter texts, add more important keywords at the end. This method ensures that the lengths of all samples are the same, while also retaining the important content of the longer samples and strengthening the feature information of the shorter samples.
[0022] Step 4: Take the sentence with the highest weight in the key sentence sequence as the key sentence of the current sample, and input the processed text into the next layer;
[0023] The function of the BiGRU layer is to extract the feature information of the input text, and fully consider the contextual relationship of the text through the forward GRU layer and the reverse GRU layer;
[0024] The core formula of the GRU network is as follows:
[0025] z t =σ(W z ·[h t-1 , x t ]) ⑶
[0026] r t =σ(W r ·[h t-1 , x t ]) (⑷)
[0027]
[0028]
[0029] Among them, formula (3) and (4) are the calculation formulas for the update gate and the reset gate, which are given by h t-1With the current input x t It is calculated that σ is the sigmoid function; Formula (5) is the calculation formula of the candidate memory unit at the current moment The information to be retained in h is screened out by the reset gate t-1 and combined with x t to form Formula (6) is the calculation formula of h at the current moment t z t decides how much information in h is to be discarded t-1 and how much information in is to be retained.
[0030] The BiGRU is a bidirectional GRU. The text sequence is input forward into the GRU to obtain forward features, and the text sequence is input backward into the GRU to obtain backward features. The forward features and backward features are combined as the overall context features of the text sequence;
[0031] The sum of the forward output and the backward output is used as the content vector H of the long text. The formula is as follows:
[0032]
[0033] The key sentence is input into the BiGRU, and the outputs of the last time step of all hidden layers are added together as the summary vector of the key sentence. The formula is as follows:
[0034]
[0035] where num_layers is the number of hidden layers, h i is the output of the last time step of the i-th layer. K sen and H are input into the Attention layer together.
[0036] The Attention layer assigns weight values to the long text according to the importance of the content, and combines the key sentence with the attention mechanism;
[0037] The key sentence vector K sen is used as the Query of the attention mechanism, and the long text content vector H is used as the Key and Value of the attention mechanism. The calculation formula is as follows:
[0038]
[0039] where d is the convergence factor, usually the dimension of the word vector. The product of Q and K T results in a score matrix of the text vector relative to the key sentence. After dividing by the convergence factor, it is normalized by the softmax function to obtain a text vector weight matrix. The text vector V is updated through the weight matrix to obtain the vector C, and the vector C is input into the last layer to obtain the classification result.
[0040] Advantages of the present invention
[0041] The present invention provides a new idea for the existing text classification method, and designs a theme classification model suitable for long texts. A text preprocessing method based on TextRank and an attention calculation method based on key sentences are proposed, which improves the accuracy of the long text classification task. It provides a practical solution for the theme classification of long articles and news in daily life and the sentiment classification of long comments on social platforms. Description of the drawings
[0042] Figure 1 Is the long text classification method of the present invention.
[0043] Figure 2 Is a schematic diagram of the GRU network structure.
[0044] Figure 3 Is a schematic diagram of the BiGRU network structure. Specific implementation manners
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0046] Embodiment:
[0047] Step 1: Input a long text, and the long text is as follows:
[0048] Long text Label This movie is very good. I like it very much. The hero in... positive
[0049] Combine the TextRank algorithm to select the keywords in the long text, and the keyword sequence is as follows [movie, like, good, hero,...]. If the text length is 480 and the unified text length is 500, then the 20 most important keywords need to be selected from the keyword sequence and filled to the end of the long text.
[0050] Step 2: Input the long text processed by TextRank into the GloVe model to generate a vector representation. The vector shape of the long text is [1, 500, 100], where 1 is the number of texts, 500 is the length of the text, and 100 is the size of the word vector.
[0051] Step 3: Input the long text vector into the BiGRU model to extract the feature information of the long text according to the context semantics.
[0052] Concatenate the output of the first time step with the output of the last time step as the summary vector of the current long text and input it into the attention layer.
[0053] Step 4: The attention layer uses "positive" as the query vector Query and the current long text as the vector to be queried Key. Through the dot product attention calculation method, attention weights are assigned to each word in the long text. The calculation formula is as follows:
[0054]
[0055] where V positive is the word vector representation of "positive", K T is the transposed K vector, d is the dimension of the word vector, used to scale the value after the dot product, and softmax is the normalization function.
[0056] Step 5: Apply the vector C of the current long text to the linear layer and the softmax layer to obtain the classification result.
[0057] Two long text datasets are selected for the experiment: IMdb and Yelp. Both IMdb and Yelp are binary classification datasets. Filter out the samples in the IMdb dataset with a length less than 400. After filtering, there are 3370 samples in the IMdb dataset as the training set and 3147 samples as the test set. Filter out the samples in Yelp with a length less than 400. Among them, the training set is set to 20000 and the test set is set to 5000.
[0058] The average sample length of the IMdb dataset is 590, and the average sample length of the Yelp dataset is 545. The information of each dataset is set as shown in the following table:
[0059] Table 1 Dataset Information
[0060]
[0061] Experimental Parameter Settings
[0062] In this experiment, a comparative experiment method is adopted. The comparative models selected are LSTM, GRU, Bi GRU, BiLSTM, TextCNN, BiGRU-Att, CNN-BiGRU, and TextRank-Bi GRU-Att. The word embedding model of all models in this paper is the GloVe (Global Vectors) model, the optimization function is Adam, the dimension of the word vector is 100. The learning rate is 1e-4. The batch sizes on the IMd b and Yelp datasets are 128 and 64 respectively, the number of hidden layers is 100, the number of training iterations is 10, the convolutional kernel size of the CNN is [3, 4, 5], and the number of channels is 100.
[0063] 4.3 Experimental Evaluation Metrics
[0064] The evaluation metrics used in this experiment are precision, recall, and F1-score, and their calculation formulas are as follows:
[0065]
[0066]
[0067]
[0068] Among them, TP is the number of positive classes predicted as positive classes; FP is the number of negative classes predicted as positive classes; FN is the number of positive classes predicted as negative classes.
[0069] Experimental Results and Analysis
[0070] The experimental results of each model on the Imdb and Yelp datasets are shown in the following table:
[0071] Table 2 Experimental Results of Each Model on the Imdb Dataset (%)
[0072]
[0073]
[0074] Table 3 Experimental Results of Each Model on the Yelp Dataset (%)
[0075]
[0076] As shown in Table 2, the precision of the method in this paper on the Imdb dataset is 74.52%, the recall rate is 80.06%, and the F1 value is 77.44%. As shown in Table 3, the precision of the method in this paper on the Yelp dataset is 87.01%, the recall rate is 87.64%, and the F1 value is 87.32%. The experimental results of the method in this paper on the two long-text datasets are better than those of the comparison model. The F1 value is 3.03% and 8.13% higher than that of the TextRank-BiGRU-Att model, indicating that the method in this paper combines key sentences to calculate the attention of long texts, which can enhance the feature extraction ability of the model and highlight the important feature information in long texts. When the text is long, the ordinary attention mechanism can only find out the relatively important content within the long text. Such a feature extraction method has too wide a scope and lacks pertinence. Generally, the key sentences of long texts contain the theme of the text. Using the theme to calculate the attention can strengthen the pertinence of feature extraction, so that the part of the content closer to the key sentences obtains a higher weight. Comparing the TextRank-BiGRU-Att with the BiGRU-Att model, the F1 value is 0.17% and 2.33% higher, which proves that the data preprocessing based on TextRank can fully retain the important information of long texts while ensuring the consistency of the sample length and enhance the feature information of short texts.
Claims
1. A long text classification method based on TextRank and attention mechanism, characterized in that, it includes the following steps; Step1: Input the long text sequence into the TextRank layer. The TextRank model will calculate the key sentence sequence and keyword sequence of the long text with weights in the range of [0-1]. The closer the weights of sentences and words are to 1, the greater the importance coefficient. Then select the sentence with the weight closest to 1 in the key sentence sequence as the key sentence of the text. Perform data preprocessing operations on the long text sequence, crop or pad each text according to the set unified sample length. For longer texts, crop the keywords with lower weights, and for shorter texts, pad the keywords with higher weights at the end; Step2: Input the text sequence processed by the TextRank layer into the Word Embedding layer to generate word vector representations; Step3: Input the long text vector into the BiGRU layer. The BiGRU will combine the context of the text to extract its characteristic information; Step4: Calculate the attention of the text vector in combination with the key sentence of the text, obtain the attention scores corresponding to the key sentence in the text vector, and update the text feature vector according to the attention scores; Step5: Input the updated text feature vector into the Linear and Softmax layers to obtain the classification result; The BiGRU is a bidirectional GRU. Input the text sequence forward into the GRU to obtain forward features, input the text sequence backward into the GRU to obtain backward features, and combine the forward features and backward features as the overall context features of the text sequence; Add the forward output and the backward output as the content vector H of the long text. The formula is as follows: Input the key sentence into the BiGRU, and add the outputs of the last time step of all hidden layers as the summary vector of the key sentence. The formula is as follows: where num_layers is the number of hidden layers, and h i is the output at the last time step of the i-th layer. Input K sen and H into the Attention layer together; The Attention layer assigns weight values according to the importance of the content in the long text, and combines the key sentence with the attention mechanism; Take the key sentence vector K sen As the Query of the attention mechanism, take the long text content vector H as the Key and Value of the attention mechanism. The calculation formula is as follows: where d is the convergence factor, usually the dimension of the word vector, Q and K T are multiplied to obtain the score matrix of the text vector relative to the key sentence. After dividing by the convergence factor, it is normalized by the softmax function to obtain the text vector weight matrix. The text vector V is updated through the weight matrix to obtain the vector C, and the vector C is input into the last layer to obtain the classification result.
2. The long text classification method based on TextRank and attention mechanism according to claim 1, characterized in that, The TextRank uses a graph network to generate weighted graph nodes for each word. If two words are in a co-occurrence window, an edge is established between the two word nodes. During training, continuously iterate and update the weights of each node. The update formula for the weights of each node is as follows: Among them, WS{V i}, WS{V j} represent the weight values of the i-th word and the j-th word; V i , V j represent the nodes of the i-th word and the j-th word in the graph; InV i , OutV j represent the in-degree set of V i and the out-degree set of V j respectively; d is the damping factor, usually set to 0.85, indicating that the probability of this point pointing to another node is 85%.
3. The long text classification method based on TextRank and attention mechanism according to claim 1, characterized in that, The key sentence sequence of the TextRank is based on the sentence similarity. By constructing a weighted graph at the sentence level, update the similarity weights between sentence nodes, and then arrange the key sentence sequence according to the similarity scores of each sentence. The similarity calculation formula between each sentence node is as follows: Where S i and S j are two sentence nodes, and w k is the word between the two sentences. The entire formula (2) is calculating the content repetition degree between the two sentences.
4. The long text classification method based on TextRank and attention mechanism according to claim 1, characterized in that, The processing steps of the TextRank layer are as follows: Step1: The first thing to do is data preprocessing. Input the long text sequence into the TextRank layer, segment it using the stop word list and filter out irrelevant words; Step2: Update the weights of each word node according to Equation (1), sort each word into a keyword sequence according to the weight value, split sentences according to any punctuation marks indicating the end of a sentence in the long text, and calculate the key sentences according to Equation (2); Step3: According to the set unified text length, delete the less important keywords in the longer text and add more important keywords to the end of the shorter text. In this way, the lengths of all samples are made the same, while the important content of the longer samples is retained and the feature information of the shorter samples is strengthened; Step4: Use the sentence with the highest weight value in the key sentence sequence as the key sentence of the current sample, and input the processed text into the next layer.
5. The long text classification method based on TextRank and attention mechanism according to claim 1, characterized in that the role of the BiGRU layer is to extract the feature information of the input text, and fully consider the context relationship of the text through the forward GRU layer and the backward GRU layer; The core formula of the GRU network is as follows: z t = σ(W z · [h t-1 , x t ) (3) r t = σ(W r · [h t-1 , x t ) (4) Among them, equations (3) and (4) are the calculation formulas for the update gate and the reset gate, which are calculated from h t-1 and the current input x t ; σ is the sigmoid function. Equation (5) is the calculation formula for the candidate memory unit at the current time, which is obtained by screening out the information to be retained in h and combining it with x t-1 to form t . Equation (6) is the calculation formula for h at the current time. z t determines how much information in h t is to be discarded and how much information in t-1 is to be retained.
Citation Information
Patent Citations
BiGRU judgment result tendency analysis method based on attention mechanism
CN111027313A
Text sentiment classification method and system
CN111881291A