BERT-TextRank text classification method based on RAdam and cosine annealing

Through the BERT-TextRank model combined with RAdam and cosine annealing learning rate scheduler, the problems of time-consuming and labor-intensive and insufficient semantic capture are solved, and efficient and accurate text classification effect is achieved.

CN120470120APending Publication Date: 2025-08-12GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510348541.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Traditional text classification methods are time-consuming and labor-intensive and difficult to accurately capture the deep semantics of text. Machine learning algorithms require manual feature engineering and ignore semantic relationships, resulting in inefficiency in classification and limited accuracy.

Method used

The BERT-TextRank model is used to combine the RAdam optimizer and the cosine annealing learning rate scheduler, and the text is compressed through the TextRank algorithm, RAdam is used to improve the model stability and convergence, and the parameters are updated through the Lookahead loop, and the learning rate is dynamically adjusted using the cosine annealing learning rate scheduler to avoid local optimization.

Benefits of technology

Efficient and accurate text classification is achieved, the stability and convergence of the model are improved, local optimization is avoided, and the recall and accuracy of classification are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470120A_ABST
    Figure CN120470120A_ABST
Patent Text Reader

Abstract

A BERT-TextRank text classification method based on RAdam and cosine annealing comprises the following six steps: S1, data preprocessing: performing cleaning, word segmentation and data set division on web text data; s2, compressing the text: compressing the data obtained in the step S1 to obtain text content with a large amount of information; s2, improving the stability and convergence of the model: after the text in the S2 is obtained, training by using a BERT model, and improving the stability and convergence of the model by using an RAdam optimizer; s4, circularly updating the parameters: when the BERT model is trained in the S3, circularly updating the parameters by using a Lookahead optimizer; s5, dynamically adjusting the learning efficiency to avoid falling into local optimum: in the step S4, dynamically adjusting the learning rate by using a cosine annealing learning rate scheduler to avoid falling into local optimum; and S6, model performance verification. According to the method, the stability and convergence of the model are improved by compressing the long text, and the learning rate is dynamically adjusted to avoid local optimum, so that the text classification performance of the BERT model is enhanced, and the classification task is more excellent in performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of natural language processing and deep learning, and in particular to a BERT-TextRank text classification method based on RAdam and cosine annealing. Background Art

[0002] With the development of the internet and artificial intelligence, we have entered a digital age. This has generated a vast amount of text data, posing a challenge for managing and classifying this data. Traditional text classification relies on manual labor, a process that consumes significant time and resources, resulting in high costs and extremely low efficiency. Therefore, achieving automated and efficient text classification has become an urgent task, aiming to improve efficiency and reduce costs. While machine learning algorithms have advanced the automation of text classification to some extent, they still have significant drawbacks. They require manual feature engineering, which is both time-consuming and labor-intensive. More critically, these algorithms often overlook the semantic relationships and context within the text when processing it, making it difficult to accurately capture the text's deeper semantics, significantly limiting the accuracy and depth of classification. In recent years, deep learning technology has emerged and gained widespread application. Text classification algorithms have also gradually shifted from traditional machine learning methods to deep learning-based neural network models and pre-trained language models. Compared to traditional machine learning models, deep learning models demonstrate significant advantages. They can automatically learn and extract semantically rich text features, minimizing manual intervention and significantly reducing labor costs, providing a practical path for efficient and accurate text classification. The present invention uses BERT as the basic model for text classification. First, the TextRank algorithm is used for text compression. Secondly, RAdam is used to improve the convergence and stability of the model. Then, the Lookahead loop is used to update the parameters. Finally, the cosine annealing learning rate scheduler is used to dynamically adjust the learning efficiency to avoid falling into the local optimum. Finally, the test set is tested to obtain the classification results, and finally a text classification method based on a pre-trained language model is obtained. Summary of the Invention

[0003] The purpose of this paper is to provide a BERT-TextRank text classification method based on RAdam and cosine annealing to classify text data and detect the classification effect through recall rate and precision.

[0004] In order to achieve the above objectives, the technical solution adopted by the present invention includes the following steps:

[0005] Step S1. Data preprocessing: cleaning, word segmentation, and data set division of network text data;

[0006] Step S2. Compress the text: compress the data obtained in step S1 to obtain text content with a large amount of information;

[0007] Step S3. Improve model stability and convergence: After obtaining the text of S2, use the BERT model for training and use the RAdam optimizer to improve the model stability and convergence;

[0008] Step S4. Loop update parameters: When training the BERT model in S3, use the Lookahead optimizer to loop update parameters.

[0009] Step S5. Dynamically adjust the learning efficiency to avoid falling into the local optimum: In S4, use the cosine annealing learning rate scheduler to dynamically adjust the learning rate to avoid falling into the local optimum.

[0010] Step S6: Model performance verification.

[0011] The data preprocessing method in step S1 is as follows:

[0012] Step S11: Remove special characters and HTML tags, process stop words and punctuation marks, and delete missing values and duplicate data.

[0013] Step S12: WordPiece algorithm is used for word segmentation. WordPiece algorithm is a subword-based word segmentation algorithm. It will merge common character combinations into subwords based on the frequency of occurrence of character combinations in the corpus, thereby controlling the size of the vocabulary while retaining semantic information.

[0014] Step S13: Divide the data set into a training set and a validation set.

[0015] The step S2 compresses the text, and the Textrank algorithm formula is as follows:

[0016]

[0017] Where, is the damping factor, N is the total number of sentences, for A collection of connected sentences (usually calculated by similarity), for The set of all connected sentences, For sentences arrive The edge weight of .

[0018] Step S3 improves the stability and convergence of the model. The core formula of the Radam algorithm is as follows:

[0019] , are the updated model parameters, are the model parameters of the current step, represents the learning rate,

[0020] is the correction factor, is the bias-corrected first-order moment estimate (gradient mean), is the square root of the bias-corrected second moment estimate (the mean of the squared gradients), is a constant.

[0021] The step S4 cyclically updates the parameters, and the core formula of LookAhead is as follows:

[0022] Step S41, slow weight maintenance: , is the slow weight (the historical average of the accumulated fast weight), is the fast weight (the parameter after the basic optimizer is updated), and m is the current cumulative number of steps.

[0023] Step S42: Regularly synchronize fast weights: , is the step size hyperparameter (usually 0.5).

[0024] Step S5. Dynamically adjust the learning efficiency to avoid local optimality. The formula of the cosine annealing learning rate scheduler is as follows:

[0025] , is the learning rate for the t-th training step, is the maximum value of the learning rate, usually the initial learning rate, is the minimum value of the learning rate. When the learning rate drops to this value, it will no longer drop. is the current training step number, is the total number of steps required to complete one cosine cycle.

[0026] Step S6, model performance verification: import the test data into the trained BERT model to generate the model prediction recall and precision.

[0027] The present invention has the following beneficial effects and advantages:

[0028] (1) Compared with the traditional BERT data processing method, the present invention compresses text through the TextRank algorithm to generate text summaries, solving the problem that the traditional BERT truncates long texts and cannot keep the main idea of the content relatively intact.

[0029] (2) The introduction of Radam and Lookahead can maintain the stability and convergence of the model and automatically update the parameters.

[0030] (3) Introduce the cosine annealing learning rate scheduler to dynamically adjust the learning rate to avoid falling into the local optimum. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is a step diagram of the BERT-TextRank text classification method based on RAdam and cosine annealing in the present invention. DETAILED DESCRIPTION

[0032] Example:

[0033] like Figure 1 As shown, the technical solution of the present invention includes six steps: data preprocessing, text compression, improving model stability and convergence, cyclic parameter updating, dynamic adjustment of learning efficiency, and model performance verification.

[0034] Step S1: Data preprocessing: cleaning, word segmentation, and dividing the network text data into data sets (training set and validation set);

[0035] The step S2 compresses the text: the training set is compressed using the TextRank algorithm to generate a text summary.

[0036] Step S3 improves model stability and convergence: receiving the text generated by S2, putting it into the BERT model for training, and using the RAdam optimizer to improve the stability and convergence of the BERT model;

[0037] The step S4 cyclically updates the parameters: accepts the parameters generated by Radam, maintains the fast weights and slow weights, and periodically synchronizes the parameters;

[0038] The step S5 dynamically adjusts the learning rate to avoid local optimum: obtain the learning rate of Radam, adjust it periodically, and proceed to step S6 if it meets the standard, otherwise continue to step S3.

[0039] Step S6, model performance verification: load the verification set into the trained model, perform performance evaluation, and obtain the accuracy, recall, and F1 value.

[0040] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A BERT-TextRank text classification method based on RAdam and cosine annealing, characterized in that: Step S1: data preprocessing, Step S2: compress text, Step S3: improve model stability and convergence, Step S4: cyclic update parameters, Step S5: dynamically adjust learning efficiency to avoid local optimality, Step S6: model performance verification.

2. A BERT-TextRank text classification method based on RAdam and cosine annealing according to claim 1, characterized in that: Step S1: Data preprocessing: remove special characters and HTML tags, process stop words and punctuation, delete missing values and duplicate data, and use the WordPiece algorithm for word segmentation.

3. The BERT-TextRank text classification method based on RAdam and cosine annealing according to claim 1 is characterized in that: Step S2, text compression, generates a text summary. First, the text is divided into individual sentences, and then a word vector is represented for each sentence. The similarity between sentence vectors is calculated. The similarity matrix is then converted into a graph structure with sentences as nodes and similarity scores as edges for sentence TextRank calculation. Finally, the highest-ranked sentence is selected to form the final summary.

4. The BERT-TextRank text classification method based on RAdam and cosine annealing according to claim 1 is characterized in that: In step S3, we improve the model stability and convergence by first using Radam to increase the learning rate in the initial stage to avoid the cold start problem of the BERT deep network. Then, we dynamically calculate the unbiased estimate of the second-order moment to solve the problem of insufficient learning rate attenuation in the later stage. Then, we use Radam’s variance correction mechanism to prevent model oscillation caused by gradient fluctuations during BERT training.

5. The BERT-TextRank text classification method based on RAdam and cosine annealing according to claim 1 is characterized in that: In step S4, the parameters are updated cyclically. The main optimizer (RAdam) performs fast updates, while the secondary optimizer Lookahead performs periodic slow updates. After every 5-10 main updates, the main parameters are overwritten with the slow updates, forming a parameter averaging effect. This can also effectively alleviate the parameter fluctuation problem of the multi-head attention mechanism in the Transform architecture.

6. A BERT-TextRank text classification method based on RAdam and cosine annealing according to claim 1, characterized in that: In step S5, the learning efficiency is dynamically adjusted to avoid local optimality. Combining the periodic learning rate decay of cosine annealing, the convergence path is further optimized based on RAdam. The learning rate decreases from the peak to the minimum value in each cycle, complementing the adaptive mechanism of RAdam to avoid falling into local optimality. If the model effect is not ideal, execute step S3; otherwise, execute step S6.

7. The BERT-TextRank text classification method based on RAdam and cosine annealing according to claim 1, characterized in that: Step S6: Model performance verification. Load the verification set into the trained model and perform multi-classification. Use precision, recall, and harmonic mean F1 as evaluation indicators. The evaluation indicators of the classification algorithm are specifically defined as follows: Accuracy: (1) Recall: (2) F1 value: (3) Among them, Percision is the accuracy rate, Recall is the recall rate, TP indicates the number of positive classes predicted as positive classes, Fp indicates the number of negative classes predicted as positive classes, and F N It indicates that the negative class is predicted as the number of negative classes. The higher the F1 value, the better the classification performance of the model.