A Two-step Lightweight Text Classification Method Based on Attention Mechanism

Through a two-step lightweight text classification method based on attention mechanism, a lightweight recurrent neural network combines self-attention and channel attention mechanisms, the gradient explosion and disappearance of long text data is solved, and the accuracy and efficiency of text classification is improved, which is suitable for edge settings.

CN115687627BActive Publication Date: 2025-07-04NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211577299.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-09
Publication Date
2025-07-04
Estimated Expiration
2042-12-09

AI Technical Summary

Technical Problem

Existing text classification methods are prone to gradient explosion and disappearance problems when processing long text data, and it is difficult to effectively capture the hierarchical information of text data, resulting in low classification accuracy, especially poor performance in emerging categories predicting without labeled training data.

Method used

A two-step lightweight text classification method based on attention mechanism is adopted, and a lightweight recurrent neural network is used to combine self-attention and channel attention mechanisms to learn the relationship between text data through a stacked structure, avoid gradient disappearance and explosion, and improve model training accuracy.

Benefits of technology

While efficient classification is achieved in edge settings, the accuracy and efficiency of text classification are improved, the problem of blurred boundaries of the model is overcome, and the lightweight and accuracy of the model is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115687627B_ABST
    Figure CN115687627B_ABST
Patent Text Reader

Abstract

The present invention discloses a two-step lightweight text classification method based on an attention mechanism, which relates to the technical field of text classification and is applicable to being deployed in edge settings. A stacked lightweight recurrent neural network is utilized. This network is a special recurrent neural network that can comprehensively learn the relationships between the input text data; while ensuring the accuracy of the model, it also ensures the lightweight nature of the model. On the one hand, the lightweight recurrent neural network is used to explore the relationships of text data, avoiding the occurrence of gradient vanishing and gradient explosion problems; at the same time, the self-attention mechanism and the channel attention mechanism are also utilized, combined with the lightweight recurrent neural network to further explore the relationships between text data, overcoming the problem of the fuzzy boundary of the model to a certain extent. Therefore, this text classification method has higher classification efficiency and higher classification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text classification, and in particular to a two-step lightweight text classification method based on an attention mechanism. Background Art

[0002] Text classification is one of the basic tasks in various natural language processing (NLP) applications, such as sentiment analysis, topic tagging, and question answering. Although it has been proven that a variety of methods have achieved success in supervised text classification, they tend to fail when applied to predicting incremental emerging categories without labeled training data; the standard paradigm of text classification relies on supervised learning, and it is well known that the size and quality of labeled data will strongly affect its performance.

[0003] Recurrent Neural Network (RNN) has the ability to model variable-length sequential data and has been widely used to solve text classification problems. When applying RNN to classify the semantics of text data, there are two key technical challenges.

[0004] First, the length of text varies from dozens to thousands of words. For long text data, due to the problems of gradient explosion and vanishing, the effectiveness of RNN will be affected; second, text data is usually hierarchical in structure, and understanding its actual semantics requires integrating information from text components at different granularities, namely words, phrases, and sentences; although explicitly modeling the hierarchical information of the original text will have a beneficial impact on the classification accuracy, RNN essentially involves an ordinary structure arranged in sequence, so it is limited in capturing the hierarchical information in text data.

[0005] To address the first challenge, various methods have been proposed to capture the long-term dependencies between words in long texts. One attempt is the threshold mechanism used in Long Short-Term Memory (LSTM) and Gate Recurrent Unit (GRU). Compared with ordinary RNN, the gates enable the recurrent architecture to maintain relatively long-term memory, thus facilitating the learning of long-term dependencies; another attempt is to try to modify the connection topology between different steps. The key idea is to add skip connections from early steps to later steps in order to achieve better information and gradient flow by bypassing intermediate steps; in practice, using a gradient norm clipping strategy can greatly overcome the explosive gradient problem, but the gradient vanishing problem still remains to be solved.

[0006] The emergence of Transformer-based pre-trained language models, such as the BERT model (Bidirectional Encoder Representation from Transformers), has reshaped the landscape of natural language processing, leading to a significant improvement in the performance of most natural language processing tasks, including text classification; these models typically rely on pre-training on a large-scale heterogeneous corpus in a general masked language modeling (MLM) task, that is, predicting the words masked in the original text.

[0007] The most popular recent text classification methods are graph-based models, such as TextGCN, which first induces a synthetic word-document co-occurrence graph on the corpus and then applies a graph neural network (GNN) to perform the classification task; in addition to TextGCN, there are subsequent works such as HeteGCN, TensorGCN, and HyperGAT, which we collectively refer to as graph-based models.

[0008] When classifying text types, if the computer takes too long to process each piece of text, it will lead to too low efficiency, and the time for analyzing text types will not show the advantages of the computer in analyzing text; currently, the classification accuracy achieved by most computer-based text classifications is not high enough. For many similar text types, computer models are very likely to make mistakes in judgment, resulting in a low accuracy rate. Summary of the Invention

[0009] To solve the above technical problems, the present invention provides a two-step lightweight text classification method based on the attention mechanism, including the following steps

[0010] S1. Preprocess the text data and convert the text data into word vectors X = {X i , i = 1, 2,..., n}, where X i represents the word vector of each piece of text data;

[0011] S2. Shuffle all word vectors and their corresponding labels, and divide the preprocessed data;

[0012] S3. Build a lightweight text classification model and randomly initialize the model parameters;

[0013] S4. Set the hyperparameters of the lightweight text classification model, train the lightweight text classification model to obtain the optimal model parameters, and test the lightweight text classification model after retaining the optimal model parameters;

[0014] S5. Input the text data of unknown categories into the lightweight text classification model to achieve automatic classification;

[0015] Step S3 specifically includes the following steps

[0016] S3.1. Divide the word vectors of each text segment into m equal-length segments with a length of k. Each segment with a length of k corresponds to a recurrent neural network model, and each segment is used as the input of the corresponding recurrent neural network model. The recurrent neural network model has two layers in total. The first layer includes three recurrent neural network models, and a self-attention mechanism for improving the training accuracy of the model is provided between adjacent recurrent neural network models in the first layer. The second layer includes one recurrent neural network model;

[0017] S3.2. Aggregate the output results of all recurrent neural network models in the first layer into the recurrent neural network model in the second layer; the output of the recurrent neural network model in the first layer is set as follows,

[0018]

[0019] Among them, represents the recurrent neural network model in the first layer, and β1,i represents the output result of each text segment after ;

[0020] S3.3. Input the output results of all recurrent neural network models in the first layer into the recurrent neural network model in the second layer; the output of the recurrent neural network model in the second layer is set as follows,

[0021]

[0022] Among them, represents the recurrent neural network model in the second layer, represents the self-attention mechanism, and β 2,i represents the output result of the recurrent neural network model in the second layer;

[0023] S3.4. Input the output result of the recurrent neural network model in the second layer into the channel attention mechanism, and finally obtain the output result. The output is set as follows,

[0024]

[0025] Among them, σ represents the channel attention mechanism, and Out represents the output result;

[0026] S3.5. Input the output result of the channel attention mechanism into the classifier for classification.

[0027] The further limited technical solution of the present invention is:

[0028] Further, in step S1, the text is processed by the word embedding method Embedding to convert the text data into word vectors.

[0029] For the aforementioned two-step lightweight text classification method based on the attention mechanism, step S2 includes the following steps

[0030] S2.1. Shuffle the dataset composed of the word vectors of each text segment and their corresponding labels.

[0031] S2.2. Divide the dataset into a training set, a test set, and a validation set.

[0032] For the aforementioned two-step lightweight text classification method based on the attention mechanism, in step S2.2, the proportion of the training set, the test set, and the validation set in the dataset is set to 3:1:1.

[0033] For the aforementioned two-step lightweight text classification method based on the attention mechanism, in step S3.1, the length of the word vector of each text segment is set to 320, m is set to 10, and k is set to 32.

[0034] For the aforementioned two-step lightweight text classification method based on the attention mechanism, in step S3.1, the number of neurons in the recurrent neural network model is set to 320.

[0035] For the aforementioned two-step lightweight text classification method based on the attention mechanism, in step S3.5, the classifier is composed of three fully connected layers, and the dropout rate set for the dropout layer in each fully connected layer is 0.3.

[0036] For the aforementioned two-step lightweight text classification method based on the attention mechanism, step S4 includes the following steps

[0037] S4.1. Set the relevant hyperparameters of the lightweight text classification model, set the number of model training epochs Epoch to 10, and set the model training batch size batch_size to 256.

[0038] S4.2. Input the data of the training set into the built lightweight text classification model for training, and use the validation set to detect the text classification accuracy of the lightweight text classification model. The validation set is used to observe whether the lightweight text classification model has problems of overfitting or underfitting; finally, obtain the optimal parameters of the lightweight text classification model.

[0039] S4.3. After training, retain the model parameters and input the test set for testing.

[0040] For the aforementioned two-step lightweight text classification method based on the attention mechanism, in step S4, the optimizer used during the training of the lightweight text classification model is set to the Adam optimizer.

[0041] For the aforementioned two-step lightweight text classification method based on the attention mechanism, in step S4, the loss function in the lightweight text classification model is set to the sparse categorical crossentropy loss function.

[0042] The beneficial effects of the present invention are as follows:

[0043] (1) In the present invention, a lightweight text classification method is designed, which is suitable for deployment in edge settings. A stacked lightweight recurrent neural network is used. This network is a special recurrent neural network that can comprehensively learn the relationships between the input text data; while ensuring the accuracy of the model, the lightweight nature of the model is also guaranteed.

[0044] (2) In the present invention, on the one hand, a lightweight recurrent neural network is used to explore the relationships of text data, avoiding the occurrence of gradient vanishing and gradient explosion problems; at the same time, the self-attention mechanism and the channel attention mechanism are also used, combined with the lightweight recurrent neural network to further explore the relationships between text data, overcoming the problem of fuzzy boundaries of the model to a certain extent. Therefore, this text classification method has higher classification efficiency and higher classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is a schematic structural diagram of the lightweight text classification method in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] A two-step lightweight text classification method based on the attention mechanism provided in this embodiment is as Figure 1 shown, and includes the following steps

[0047] S1. Since the lengths of different texts are inconsistent, in the data preprocessing stage, it is necessary to use the word embedding method Embedding to process the text and convert the text data into word vectors X = {X i , i = 1, 2,..., n}, where X i represents the word vector of each piece of text data.

[0048] S2. Shuffle all the word vectors and their corresponding labels, and divide the preprocessed data; step S2 specifically includes the following sub-steps

[0049] S2.1. Shuffle the data set composed of the word vectors of each piece of text and their corresponding labels to prevent the problem of overfitting during the training of the model;

[0050] S2.2. Divide the dataset into a training set, a test set, and a validation set, with the ratio of the three being 3:1:1.

[0051] S3. Build a lightweight text classification model and randomly initialize the model parameters; Step S3 specifically includes the following sub-steps

[0052] S3.1. Divide the word vectors of each text segment into m equal-length segments of length k. Each segment of length k corresponds to a recurrent neural network model (Recurrent Neural Networks, RNN), and each segment is used as the input to the corresponding recurrent neural network model; the length of the word vectors of each text segment is set to 320, m is set to 10, k is set to 32, and the number of neurons in the recurrent neural network model is set to 320;

[0053] The recurrent neural network model has a total of two layers. The first layer includes three recurrent neural network models, and the second layer includes one recurrent neural network model; between the recurrent neural network models, a self-attention mechanism is used to improve the training accuracy of the model;

[0054] S3.2. Set the output size of all recurrent neural network models in the first layer to 32, and converge the output results to the recurrent neural network model in the second layer; the output of the recurrent neural network models in the first layer is set as follows,

[0055]

[0056] Among them, represents the recurrent neural network model in the first layer, and β 1,i represents the output result of each text segment after passing through ;

[0057] S3.3. Input the output results of all recurrent neural network models in the first layer into the recurrent neural network model in the second layer; the output of the recurrent neural network model in the second layer is set as follows,

[0058]

[0059] Among them, represents the recurrent neural network model in the second layer, represents the self-attention mechanism, and β 2,i represents the output result of the recurrent neural network model in the second layer;

[0060] S3.4. Input the output result of the recurrent neural network model in the second layer into the channel attention mechanism, and finally obtain the output result. The output is set as follows,

[0061]

[0062] Among them, σ represents the channel attention mechanism, and Out represents the output result;

[0063] S3.5. Input the output result of the channel attention mechanism into the classifier for classification. The classifier consists of three fully connected layers, and the dropout rate set in the dropout layer of each fully connected layer is 0.3.

[0064] S4. Set the hyperparameters of the lightweight text classification model, train the lightweight text classification model to obtain the optimal model parameters, and test the lightweight text classification model after retaining the optimal model parameters; Step S4 specifically includes the following sub-steps

[0065] S4.1. Set the relevant hyperparameters of the lightweight text classification model. Set the number of model training epochs Epoch to 10, set the model training batch size batch_size to 256, set the optimizer used during training to the Adam optimizer, and set the loss function to the sparse categorical crossentropy loss function;

[0066] S4.2. Input the data of the training set into the built lightweight text classification model for training, and use the validation set to detect the text classification accuracy of the lightweight text classification model. The validation set is used to observe whether the lightweight text classification model has problems of overfitting or underfitting; Finally, obtain the optimal parameters of the lightweight text classification model;

[0067] S4.3. Retain the model parameters after training is completed and input the test set for testing.

[0068] S5. Input the text data of unknown categories into the lightweight text classification model to achieve automatic classification.

[0069] The present invention uses a lightweight recurrent neural network model combined with an attention mechanism to establish a two-step lightweight text classification method. In this method, the lightweight recurrent neural network is mainly a stacked recurrent neural network. The two-step method is mainly reflected in: First, use the lightweight recurrent neural network combined with the self-attention mechanism for model training; Second, combine the channel attention mechanism for classification.

[0070] Thus, the relationships between the input text data can be comprehensively learned; while ensuring the accuracy of the model, the lightweight nature of the model is also guaranteed; on the one hand, a lightweight recurrent neural network is used to explore the relationships of the text data, avoiding the occurrence of gradient vanishing and gradient explosion problems; at the same time, the self-attention mechanism and the channel attention mechanism are also used, combined with the lightweight recurrent neural network to further explore the relationships between the text data, overcoming the problem of the fuzzy boundary of the model to a certain extent. Therefore, the text classification method in this article has higher classification efficiency and higher classification accuracy.

[0071] In addition to the above embodiments, the present invention can also have other implementation manners. Any technical solutions formed by equivalent replacement or equivalent transformation fall within the protection scope required by the present invention.

Claims

1. A two-step lightweight text classification method based on the attention mechanism, characterized in that: It includes the following steps S1. Preprocess the text data and convert it into word vector X={X i ,i=1,2,…,n}, where X i Represent the word vector for each piece of text data; S2. Shuffle all word vectors and their corresponding labels, and partition the preprocessed data; S3. Build a lightweight text classification model and randomly initialize the model parameters; S4. Set the hyperparameters of the lightweight text classification model, train the lightweight text classification model to obtain the optimal model parameters, and test the lightweight text classification model after retaining the optimal model parameters; S5. Input the text data of unknown categories into the lightweight text classification model to achieve automatic classification; Step S3 specifically includes the following steps S3.

1. Divide the word vectors of each text segment into m segments of equal length k. Each segment of length k corresponds to a recurrent neural network model, and each segment is used as the input of the corresponding recurrent neural network model. The recurrent neural network model has two layers in total. The first layer includes three recurrent neural network models, and there is a self-attention mechanism for improving the model training accuracy between adjacent recurrent neural network models in the first layer. The second layer includes one recurrent neural network model; S3.

2. Aggregate the output results of all recurrent neural network models in the first layer into the recurrent neural network model in the second layer. The output of the recurrent neural network model in the first layer is set as follows, Among them, represents the recurrent neural network model of the first layer, and β 1,i represents the output result after each piece of text passes through ; S3.

3. Input the output results of all recurrent neural network models in the first layer into the recurrent neural network model in the second layer. The output of the recurrent neural network model in the second layer is set as follows, Among them, represents the recurrent neural network model of the second layer, represents the self-attention mechanism, and β 2,i represents the output result of the recurrent neural network model of the second layer; S3.

4. Input the output result of the recurrent neural network model in the second layer into the channel attention mechanism, and finally obtain the output result. The output is set as follows, where, σ represents the channel attention mechanism, and Out represents the output result; S3.

5. Input the output result of the channel attention mechanism into the classifier for classification.

2. The two-step lightweight text classification method based on the attention mechanism according to claim 1, wherein: In step S1, the text is processed by the word embedding method Embedding to convert the text data into word vectors.

3. A two-step lightweight text classification method based on the attention mechanism according to claim 1, characterized in that: Step S2 includes the following steps S2.

1. Shuffle the dataset composed of the word vectors of each text segment and their corresponding labels; S2.

2. Partition the dataset into a training set, a test set, and a validation set.

4. A two-step lightweight text classification method based on the attention mechanism according to claim 3, characterized in that: In step S2.2, the proportions of the training set, the test set, and the validation set in the dataset are set to 3:1:

1.

5. A two-step lightweight text classification method based on the attention mechanism according to claim 1, characterized in that: In step S3.1, the length of the word vectors of each text segment is set to 320, m is set to 10, and k is set to 32.

6. The two-step lightweight text classification method based on the attention mechanism according to claim 5, wherein: In step S3.1, the number of neurons in the recurrent neural network model is set to 320.

7. A two-step lightweight text classification method based on the attention mechanism according to claim 1, characterized in that: In step S3.5, the classifier is composed of three fully connected layers, and the dropout rate set for the dropout layer in each fully connected layer is 0.

3.

8. A two-step lightweight text classification method based on the attention mechanism according to claim 1, characterized in that: Step S4 includes the following steps S4.

1. Set the relevant hyperparameters of the lightweight text classification model, set the number of model training epochs Epoch to 10, and set the model training batch size batch_size to 256; S4.

2. Input the data of the training set into the built lightweight text classification model for training, and use the validation set to detect the text classification accuracy of the lightweight text classification model. The validation set is used to observe whether the lightweight text classification model has problems of overfitting or underfitting; finally, obtain the optimal parameters of the lightweight text classification model. S4.

3. After training, retain the model parameters and input the test set for testing.

9. A two-step lightweight text classification method based on the attention mechanism according to claim 8, characterized in that: In step S4, the optimizer used during the training of the lightweight text classification model is set to the Adam optimizer.

10. A two-step lightweight text classification method based on an attention mechanism according to claim 8, characterized in that: In step S4, the loss function in the lightweight text classification model is set to the sparse categorical crossentropy loss function.

Citation Information

Patent Citations

  • A method of constructing a circulating neural network model of multi-layer attention mechanism

    CN109408633A

  • Text classification method based on deep learning

    CN112163064A