Chinese literature classification method based on incremental learning
By introducing incremental learning methods in the field of Chinese literature classification, using TextCNN network and decoupled distillation loss function, combined with weight alignment method, the problem of frequent retraining of the models in the existing technology is solved, and an efficient and flexible Chinese literature classification system is realized.
Patent Information
- Application Number
- CN202510022875.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-07
AI Technical Summary
The existing Chinese literature classification method relies on batch learning, which leads to the need for retraining whenever new literature data arrives, which increases the consumption of computing resources and makes it difficult to adapt to the continuous growth and changes of data.
Using an incremental learning-based method, a Chinese literature classification system that can dynamically update and optimize the model is constructed by introducing attention layer and decoupled distillation loss function into the TextCNN network model, combined with the weight alignment method.
It realizes dynamic adaptation to new data and new knowledge without retraining the entire model, improves the flexibility and scalability of the classification system, reduces computing resource consumption, and significantly improves classification accuracy and stability.
Smart Images

Figure CN119939428A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of Chinese document classification, and in particular relates to an effective Chinese document classification method based on incremental learning. Background Art
[0002] Documents are important carriers of knowledge and information, and are of great significance to academic research, knowledge dissemination, and cultural heritage. With the continuous deepening of scientific research and the increasing prominence of interdisciplinary phenomena, the number of documents has shown an explosive growth trend. Document classification refers to the activities and methods of organizing and managing documents according to the characteristics of their content and form. This classification method helps to systematically reveal and organize documents, making it easier for users to find and use them.
[0003] At present, there are two ways to classify documents: manual classification and automatic classification. Manual document classification relies on the knowledge and experience of experts to ensure the accuracy and pertinence of classification. However, in the face of such a huge amount of data, the traditional manual classification method also has some disadvantages. Manual classification not only requires a lot of time and human resources, but also is costly. Due to human factors, errors or omissions may occur in the classification process, making it difficult to ensure the consistency and accuracy of the classification. Therefore, the development of efficient automatic classification technology has become an urgent need.
[0004] As an important carrier of core research topics, scientific literature is often highly professional, complex, and technically profound in its language expression. This complexity makes automatic understanding and classification of these texts a very challenging task. With the vigorous development of machine learning and deep learning, the introduction of these technologies to achieve automatic literature classification has important practical significance for improving classification efficiency and accuracy.
[0005] Traditional Chinese literature classification methods mainly use machine learning methods. Yang Min et al. used TF-IDF as the feature representation of documents and built a bibliographic automatic classification system based on support vector machine (SVM). They also proved the feasibility of applying machine learning algorithms to the automatic classification of information resources in library application practice. Li Xiangdong et al. constructed word frequency, word count, etc. as the features of documents, and used the K nearest neighbor (KNN) algorithm to automatically classify journal articles belonging to different categories, and achieved certain results in the subject classification of journals with a small number of categories. Although traditional machine learning has good performance, there are still some limitations: the encoding method of machine learning algorithms is difficult to accurately capture the deep meaning and key features in complex texts, resulting in a decrease in classification performance; machine learning algorithms perform poorly on large data sets and are unable to cope with the processing of the growing amount of massive literature data. In recent years, with the rise of deep learning technology, more and more new technologies and large pre-trained models have gradually been applied to the research field of automatic literature classification. Deng Sanhong et al. first used word embedding to represent the text of Chinese books, and then used the long short-term memory network (LSTM) model to perform multi-label classification by building multiple binary classifiers. Luo Pengcheng et al. used pre-trained models such as BERT and ERNIE (enhanced representation through knowledge integration) to compare with traditional machine learning and typical deep learning algorithms on 21 first-level discipline literature data sets in humanities and social sciences. The comparison results verified the superiority of the pre-trained models. However, Chinese literature classification methods, whether based on machine learning or deep learning, currently rely on batch learning. This means that every time new literature data arrives, the model needs to be retrained, which further increases the consumption of computing resources.
[0006] Based on the above background, the present invention considers using a novel method to solve the emerging Chinese literature data. Incremental learning, as a machine learning paradigm that can continuously learn new knowledge, is innovatively introduced into the field of Chinese literature classification in the present invention. Incremental learning allows the model to continuously learn new sample data while retaining existing knowledge, thereby gradually improving classification performance. This method has the following advantages: incremental learning can dynamically adapt to new data and new knowledge without retraining the entire model, thereby improving the flexibility and scalability of the classification system; compared to batch learning, incremental learning is more efficient in processing new data because it only needs to update part of the model's parameters instead of retraining the entire model; as new data is continuously added, the incremental learning model can continuously optimize its classification performance and improve the accuracy and stability of classification.
[0007] In summary, the Chinese literature classification method based on incremental learning is an innovative solution proposed to cope with the challenges of document text complexity, explosive growth of data volume, and limitations of traditional classification methods. It is expected to bring revolutionary changes to the field of document classification, significantly improve classification efficiency and accuracy, thereby more effectively promoting the management and utilization of documents and providing important basic information support for basic scientific research. In the future, in order to classify Chinese documents more comprehensively, accurately and efficiently, relevant researchers need to conduct in-depth research on Chinese literature classification methods. Summary of the invention
[0008] In order to overcome the deficiencies of the above-mentioned prior art, the purpose of the present invention is to provide a Chinese document classification method based on incremental learning, aiming to form a benchmark data set using strict filtering conditions, ensure data accuracy and reduce system errors. The TextCNN network model is used to effectively learn the local features of the text sequence. The innovative introduction of the attention layer in the TextCNN network can significantly improve the model's attention to key information. A model based on incremental learning is constructed, and the following two methods are used to effectively overcome "catastrophic forgetting" so that the results output by the classification model have high accuracy and reliability. One is the weight alignment method (Weight Aligning, referred to as WA), which corrects the deviation of the fully connected layer output to the new and old categories due to data imbalance by updating the weight of the fully connected layer output as a new category; the other is to use a decoupled distillation loss function, the purpose of which is to help the model retain the knowledge and memory of the old tasks and achieve a balance between new and old knowledge. Therefore, the research method of Chinese document classification based on incremental learning has the advantages that it can adapt to the continuous growth of document data, efficiently update the model to capture new features, avoid catastrophic forgetting, and save computing resources. This helps to classify Chinese documents in real time and accurately, providing strong support for technological innovation and scientific research.
[0009] To achieve the above purpose, the main technical solution adopted by the present invention is: a Chinese document classification method based on incremental learning, comprising the following steps:
[0010] Step S1, construct a benchmark data set for the Chinese document classification problem; Step S2, construct a benchmark model for Chinese document classification based on incremental learning; Step S3, use a flock selection algorithm to select representative data from the old category, and merge it with the new category data to construct an incremental learning data set; Step S4, input the incremental learning data set into the benchmark model, and train it under the constraint of the incremental learning loss function; Step S5, use the weight alignment method to update the weight of the new category of the output of the fully connected layer in the benchmark model; Step S6, input the Chinese document to be classified into the model after the weight alignment of the fully connected layer, perform sequence information recognition, and obtain the document classification result.
[0011] The beneficial effects of the present invention are:
[0012] Since the present invention crawls data from Baidu Academic System based on the existing data set as a reliable data source, strict filtering conditions are used to ensure the accuracy of the data; the Chinese pre-training model is used to obtain the embedded representation of the Chinese document text sequence, which can more effectively capture the semantic information in the text; the incremental learning method is applied to the field of Chinese document classification, two methods for solving "catastrophic forgetting" are used and combined with the TextCNN network model to identify the Chinese document sequence information, so that the classification results of the model are more accurate and reliable. Therefore, the present invention has the advantages of novel scheme and accurate results.
[0013] In the present invention, data is crawled from the Baidu academic system on the basis of existing data sets, and pre-processing operations such as strict screening are performed on the data to construct a benchmark data set for the classification of Chinese documents. The Chinese document data set used in the present invention is a valuable resource contributed by Mr. Li Ronglu of the Natural Language Processing Group of the International Database Center of the School of Computer Science and Technology of Fudan University. Baidu Academic is a powerful and resource-rich academic resource search platform that provides comprehensive, convenient and efficient academic services for scientific researchers. The present invention obtains data from the Baidu academic system, and these data lay the foundation for the research of document classification. The document data obtained using crawler technology is merged with the document data in the Chinese text classification data set of Fudan University by category, and then the number of statistics and strict screening are performed. In short, through these pre-processing operations, a high-quality benchmark data set for the classification of Chinese documents is constructed.
[0014] The present invention designs an efficient benchmark model for Chinese literature classification based on incremental learning. The model consists of a coding layer, a convolutional layer, an activation layer, a maximum pooling layer, an attention layer and a fully connected layer, which can comprehensively and systematically identify the sequence information of Chinese literature. In particular, the present invention innovatively introduces an attention layer in the model, which can assign different weights to each feature after maximum pooling, ensuring that key features receive higher attention, thereby significantly improving the classification performance of the model.
[0015] In the present invention, the model is trained using a decoupled distillation loss combined with a cross entropy loss. By adjusting the parameters of the two parts after the decoupling of the original distillation loss, the model can better retain old knowledge while learning new knowledge. The present invention also introduces a weight alignment method to overcome "catastrophic forgetting", updating the weights of the fully connected layer output for the new category, and keeping the weights of the fully connected layer output for the old category unchanged, so that the output result of the classification model has higher accuracy and reliability.
[0016] The Chinese document classification method based on incremental learning proposed by the present invention uses a flock selection algorithm to select representative data from the old category. The new category data is learned together with the representative old category data. The incremental learning model learns knowledge under the constraints of the incremental learning loss function, which can not only learn the information of the new category, but also maintain the classification performance of the old category test data. In this way, a more accurate, efficient and low-storage Chinese document classification method is achieved.
[0017] The present invention can adapt to the characteristics of the continuous growth and change of document data, optimize the classification model through incremental learning of new data, and improve classification accuracy and efficiency. This is of great significance for quickly identifying and protecting knowledge achievements and promoting the effective use of document information. These aspects or other aspects of the present application will be more concise and easy to understand in the description of the following embodiments. It should be understood that the above general description and the detailed description below are only exemplary and explanatory and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or the prior art description. Obviously, the drawings described below are only some embodiments of the present application. Among them:
[0019] Figure 1 It is a schematic diagram of the steps of the Chinese document classification method based on incremental learning of the present invention;
[0020] Figure 2 It is a schematic diagram of the specific sub-steps of step S1 in the present invention;
[0021] Figure 3 It is a bar chart showing the number of data contained in each category in the Chinese literature data set downloaded and used by the present invention;
[0022] Figure 4 It is a columnar statistical chart of the number of original data and the number of filled data of the categories that need to be filled with data when constructing the data set of the present invention;
[0023] Figure 5 It is a structural schematic diagram of a benchmark model of Chinese document classification based on incremental learning of the present invention;
[0024] Figure 6 It is a flow chart of the Chinese document classification method based on incremental learning of the present invention;
[0025] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0026] Below, the present application is further described in conjunction with the accompanying drawings and specific implementation methods. It should be noted that, under the premise of no conflict, the various embodiments or technical features described below can be arbitrarily combined to form a new embodiment.
[0027] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0028] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0029] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present application, rather than to limit the present application. It should also be noted that, for ease of description, only the parts related to the present application, rather than all structures, are shown in the accompanying drawings.
[0030] It should be mentioned before discussing the exemplary embodiments in more detail that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the steps as sequential processes, many of the steps therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but there can also be additional steps not included in the accompanying drawings.
[0031] The purpose of the present invention is to solve the shortcomings in the above background and propose a Chinese document classification method based on incremental learning. To achieve the above purpose, the present invention provides the following technical solutions, which are further described in detail below in conjunction with the accompanying drawings and examples.
[0032] The present invention provides a Chinese literature classification method based on incremental learning, such as Figure 1 As shown, the method comprises the following steps:
[0033] Step S1, constructing a benchmark dataset for Chinese document classification problem.
[0034] This step crawls data from the literature search system based on the existing Chinese literature dataset, and then performs strict preprocessing operations such as screening on the data to construct a high-quality benchmark dataset for Chinese literature classification problems, providing strong support for the development and application of Chinese literature classification technology.
[0035] like Figure 2As shown, step S1 may specifically include:
[0036] Step S11: Download the Chinese literature classification data set and statistically analyze the data to determine whether data needs to be filled and determine the category of data that needs to be filled.
[0037] Specifically, the Chinese document dataset used in the present invention is constructed by Professor Li Ronglu of the Natural Language Processing Group of the International Database Center of the School of Computer Science and Technology of Fudan University. The present invention uses it as the original Chinese document dataset. It can be obtained from the following link: https: / / gitcode.com / open-source-toolkit / 6a679. This dataset is designed for Chinese text classification tasks and contains a wealth of Chinese documents, involving 20 categories of Chinese document data, and contains a total of 19,636 data. Figure 3 As shown in the figure, the number of data in 20 categories is shown. Among them, there are 1481 data belonging to art, 67 data belonging to literature, 120 data belonging to education, 89 data belonging to philosophy, 934 data belonging to history, 1282 data belonging to physical space, 65 data belonging to energy, 55 data belonging to electronics, 52 data belonging to communication, 2715 data belonging to computers, 67 data belonging to minerals, 116 data belonging to transportation, 2435 data belonging to environment, 2043 data belonging to agriculture, 3201 data belonging to economy, 103 data belonging to law, 104 data belonging to medicine, 150 data belonging to military, 2050 data belonging to politics, and 2507 data belonging to sports.
[0038] From the above statistical results, we can see that the number of documents in the 11 categories of literature, education, philosophy, energy, electronics, communications, minerals, transportation, law, medicine and military is too small. After calculation, the category with the least number is communications, accounting for 0.2648% of the total, and the category with the most number is economy, accounting for 16.302% of the total. The data distribution is extremely unbalanced. Therefore, for these 11 categories, data needs to be filled to balance the data in each category.
[0039] Step S12: Using the category that needs to be filled with data obtained in step S11 as a search condition, search in an academic document search system to obtain a corresponding Chinese document list as a search result.
[0040] This invention chooses to use Baidu Academic as the academic literature retrieval system. Baidu Academic (https: / / xueshu.baidu.com) is a free academic resource search platform under Baidu, covering various academic journals, conference papers and other academic resources, aiming to provide the best scientific research experience for scholars at home and abroad. Baidu Academic System is a comprehensive academic platform that integrates academic resource retrieval, academic services, academic analysis and other functions, providing scientific researchers with comprehensive, accurate, fast and novel academic resources and services.
[0041] Specifically, access the Baidu Academic system, use the name of the category to be filled with data obtained in step S11 as the search keyword, and enter these keywords in the search box of the Baidu Academic system. The system will return a list of documents related to the search keywords. Browse the search results, pay attention to the key information such as the title, author, and abstract of the document, to determine whether the retrieved document is related to the data category to be filled.
[0042] Step S13: Use Selenium-based program code to automatically crawl the titles, abstracts and keyword data of the Chinese literature list in the search results obtained in step S12 to obtain the literature data that needs to be filled in the data category.
[0043] Selenium is a tool for automated testing of web applications. It can run directly in the browser to simulate user operations in the browser, such as clicking, inputting, scrolling, etc., thereby realizing automated access to web pages and data capture.
[0044] Specifically, the present invention realizes the automatic crawling of data by writing a Selenium script. In the Selenium script, WebDriver (a tool for automated testing of Web applications) is used to access the page of the search results determined in step S12. Code is written to locate and extract the title, abstract and keyword elements in the document list. A loop logic is written to traverse each item in the document list and grab its title, abstract and keyword data. The captured data is stored in a text file with a suffix of .xsv. In this way, the document data that needs to be filled with data categories is obtained.
[0045] Step S14: filter the document data that needs to be filled with data categories crawled in step S13, and use the filtered data to fill the data in the original Chinese document data set by category to obtain a filled Chinese document classification data set.
[0046] Specifically, select the literature with complete content information and clear category information from the literature data of the data categories to be filled crawled in step S13. According to the experience of predecessors in constructing incremental data sets, the present invention fixes the number of literatures in each literature category at 600. Then, use the filtered literature data to fill the original Chinese literature data set by category, as Figure 4 shown. This figure shows the data situation of the categories that need to be filled with data. Randomly select 600 pieces of literature data from the categories in the original Chinese literature data set whose number exceeds 600 to ensure that the total number of each category is 600. Therefore, the data set for Chinese literature classification constructed by the present invention has a total of 20 literature categories, and each category contains 600 pieces of data.
[0047] Step S15: Perform operations of removing spaces and stop words on the Chinese literature data in the filled data set obtained in step S14 to obtain a benchmark data set for Chinese literature classification problems.
[0048] Removing redundant spaces in the text (including leading and trailing spaces, unnecessary spaces in the middle, etc.) is an important step in data cleaning. These spaces may be generated during data entry or format conversion. They do not carry any useful information but will increase the storage space and processing time of the data. Stop words are words that frequently appear in the text but lack practical meaning, such as "de", "le", "he", etc. These words do not contribute much to the meaning analysis of the text but will occupy a large amount of storage space and computing resources. By removing stop words, the keywords in the text can be made more prominent, which helps to extract and understand the theme and key information of the text. This is particularly important for literature data analysis because literature data usually contains a large number of technical terms and key information. For example, there is a Chinese literature with the title information "The Core Issues of Vocational Education Teaching Reform in the New Era". The result after removing spaces and stop words from this title information is: "The Core Issues of Vocational Education Teaching Reform in the New Era".
[0049] The method for Chinese literature classification based on incremental learning described in the present invention further includes the following steps:
[0050] Step S2: Construct a benchmark model for Chinese literature classification based on incremental learning.
[0051] As Figure 5 shown, in the present invention, the benchmark model of the method for Chinese literature classification based on incremental learning is composed of six parts: an encoding layer, a convolutional layer, an activation layer, a max pooling layer, an attention layer, and a fully connected layer. The specific construction method is as follows:
[0052] Step S2 specifically includes:
[0053] Step S21: Set the encoding layer and use the BERT-wwm-ext, Chinese pre-trained language model to obtain the embedded representation of Chinese literature text data as the input of the subsequent model.
[0054] BERT-wwm-ext, Chinese is a Chinese pre-trained language model released by the Harbin Institute of Technology iFlytek Joint Laboratory (HFL). It is an upgraded version of BERT-wwm. By increasing the pre-training data set and the number of training steps and adopting the full-word masking technology and Chinese word segmentation tools, the model performs well in Chinese NLP tasks. Its wide range of application scenarios and significant advantages make it an important technological achievement in the field of Chinese NLP. By using the BERT-wwm-ext, Chinese model, the present invention can convert Chinese literature text data into a low-dimensional, dense numerical vector representation, which is called an embedded representation. This embedded representation can capture the semantic information in the text, enabling the model to better understand the text content.
[0055] Specifically, step S21 includes:
[0056] Step S211: Acquire Chinese document data in the benchmark data set.
[0057] Step S212: Input the Chinese document data into the BERT-wwm-ext, Chinese pre-trained language model, and obtain the hidden state of the last layer in the model as the embedded representation vector of the Chinese document.
[0058] A Chinese document sequence is recorded as S. The Chinese document is input into the BERT-wwm-ext, Chinese pre-trained language model. The model uses the word segmentation algorithm to segment the text. The result after word segmentation is recorded as g(S). The formula is as follows:
[0059] g(S)={a1, a2, a3,…,a m , …, a L};
[0060] Among them, a m It represents the mth token after word segmentation (the smallest unit into which text data is divided when being processed by the model). m represents the subscript, which means the number. L represents the number of tokens, which is 100 and is controlled by the max_length and padding parameters in the model.
[0061] In the present invention, max_length is set to 100 and padding is set to True, which means that regardless of the length of the original text, it will be truncated or padded (depending on the original length of the text) to ensure that the length of the final input sequence is 100. For example of how to perform word segmentation, for the sequence "Core issues in the teaching reform of vocational education in the new era", the word segmentation results are: "new", "era", "vocational", "education", "teaching", "reform", "core", "issues".
[0062] After that, the hidden state of the last layer in the model is obtained as the embedding representation vector of the Chinese literature. Denote the embedding representation vector of a Chinese literature as f(S), and the formula is as follows:
[0063] f(S) = {e1, e2, e3, …, e m , …, e L};
[0064] Among them, e m represents the embedding representation corresponding to the m-th token, whose dimension is 768, and it is set by the hidden_size parameter in the model. In BERT-wwm-ext, Chinese, hidden_size is set to 768, which means the output dimension size of the hidden layer is 768. L represents the number of tokens. e L represents the embedding representation corresponding to the L-th token. For example, in the sequence "Core issues in the teaching reform of vocational education in the new era", the embedding representation vector of "education" is [0.16714, -0.33311, -0.03177, 0.36868, -0.20583,......, 0.03909, -0.43855, 0.01807], and its dimension is 768.
[0065] Step S22: Set up a convolutional layer, and use convolutional kernels of different sizes to slide on the text vector to capture local features of different lengths.
[0066] Specifically, the text data is a one-dimensional sequence, so the number of input channels of the convolutional layer is 1, and at the same time, the number of output channels of the convolutional feature map is set to 100. In the present invention, the convolutional layer slides on the text embedding matrix through different convolutional kernels (also called filters) to extract local features of different lengths. For text data, the convolutional kernel can be regarded as a series of weights, which slide on the text and perform weighted summation on the data within each window to generate new features. Denote the output after convolution as y(x), and the formula is as follows:
[0067] y(x) = A * x + b;
[0068] Among them, x is the text embedding matrix, and its dimension is (batchsize, max_length, hidden_size). Batchsize is the batch size, that is, the number of document samples contained in each batch, which is set to 128. max_length is the sequence length of 100, and hidden_size is the hidden layer output dimension size of 768. A is the weight of the convolution kernel, and its dimension is defined as (h, d). h is the size of the convolution kernel, which defines the number of words covered by the convolution kernel. d is the same as the dimension of the text embedding, which ensures that the convolution kernel can be fully matched with the corresponding word vector in the text embedding matrix when sliding, so as to perform the convolution operation. b is the bias term. In the formula of the convolution operation, the multiplication involved in the symbol * is not the traditional matrix multiplication, but the process of multiplying the corresponding positions and then summing them.
[0069] Step S23: Set the activation layer and use the ReLU activation function to activate the features obtained after convolution.
[0070] Activation functions play a vital role in neural networks. Their main role is to introduce nonlinear factors, improve the expressiveness of the model, and enable neural networks to learn and simulate complex nonlinear relationships.
[0071] In the present invention, the ReLU function is used as the activation function. The ReLU function can alleviate the gradient vanishing problem and speed up the training of the model. The result after the ReLU activation function is recorded as R(x), and the formula is as follows:
[0072] R(x)=max(0,y(x));
[0073] When the convolution value y(x) is greater than 0, the value itself is output; when the convolution value y(x) is less than or equal to 0, the output is 0. x is the text embedding matrix.
[0074] Step S24: define a maximum pooling layer, perform dimensionality reduction processing on the features, and extract the most important features.
[0075] The maximum pooling layer extracts the maximum value in each pooling window as the representative feature of the window, which can effectively extract the most significant and representative feature information in the text and achieve dimensionality reduction. This helps to reduce the amount of calculation and the number of parameters of the model, and improve the training speed and generalization ability of the model. The result after maximum pooling is recorded as M(x), and the formula is as follows:
[0076] M(x) = maxpooling(R(x));
[0077] Among them, M(x) means extracting the most important features from the feature map generated by each convolution kernel, R(x) is the result after the ReLU activation function, and x is the text embedding matrix.
[0078] Step S25: Set the attention layer, obtain the weight of each feature, and then multiply it with the original feature.
[0079] The attention mechanism can give different weights to each feature in the result after the maximum pooling, so that the features that are more critical to the classification task receive higher attention, thereby improving the classification performance of the model. In the present invention, the attention weight is calculated by a linear layer and a softmax function. Let the weight of the linear layer be R, the bias be θ, and the weight and bias are randomly initialized during implementation. The calculated attention weight is recorded as η, and the formula is as follows:
[0080] η=σ(B·M(x)+θ);
[0081] Among them, σ is the softmax activation function, which is used to limit the output between 0 and 1 as the weight of each feature. M(x) is the result after maximum pooling, and x is the text embedding matrix.
[0082] Afterwards, the attention weight is multiplied by the output feature after the maximum pooling obtained in step S24 to obtain the weighted feature denoted as Q(x), and the formula is as follows:
[0083] Q(x) = η⊙M(x);
[0084] The symbol ⊙ represents element-wise multiplication. η is the attention weight, M(x) is the result after maximum pooling, and x is the text embedding matrix.
[0085] For example, the feature vector of the sequence "Core issues of vocational education and teaching reform in the new era" after the encoding layer, convolution layer, activation layer and maximum pooling layer is: [0.213, 0.324, 0.425, -0.736, 0.167, ..., 0.847, -0.625, 0.789], with a dimension of 100. This feature is input into the attention layer, and the attention weight vector obtained is: [0.013, 0.015, 0.023, 0.004, 0.011, ..., 0.045, 0.005, 0.037], with a dimension of 100, and the sum of all elements is 1. The attention weight vector is multiplied by the corresponding position of the feature vector after the maximum pooling layer to obtain the weighted result: [0.002769, 0.00486, 0.009775, -0.002944, 0.001837, ..., 0.038115, -0.003125, 0.029193], with a dimension of 100.
[0086] Step S26: define a fully connected layer to obtain the classification result of the model.
[0087] In the classification task, the fully connected layer is responsible for mapping the global features into a specific category space and outputting the prediction results for each category. This is achieved by calculating the linear combination between the input features and the weight matrix.
[0088] In the present invention, three convolution functions with different convolution kernels are defined. The Chinese literature embedding matrix x is subjected to three convolution functions with different convolution kernels respectively, and then the three different convolution layer results are respectively passed through the activation layer, the maximum pooling layer and the attention layer in turn, and the results of the attention layer are recorded as Q1(x), Q2(x), and Q3(x), respectively. These three attention layer results are spliced together by column and recorded as φ(x). Finally, φ(x) is input into the fully connected layer to obtain the prediction result for each category. Specifically, in the present invention, the bias term of the fully connected layer function is set to 0 in order to better apply the weight alignment method mentioned later. The output of the fully connected layer is recorded as o(x), and the formula is as follows:
[0089] o(x)=W T φ(x);
[0090] Among them, the dimension of o(x) is (batchsize, classes_num), classes_num is the number of output categories, which is 20. x is the text embedding matrix, W is the weight matrix of the fully connected layer, whose size is related to the dimension of the input feature vector and the number of output categories, T represents transpose, W T Represents the transpose of the weight matrix. Each element in the weight matrix represents the strength of the association between the input feature and the output category.
[0091] The Chinese document classification method based on incremental learning of the present invention further includes the following steps:
[0092] Step S3: Use the flock selection algorithm to select representative samples from the old category and merge them with the new category data to construct an incremental learning dataset.
[0093] In incremental learning, the model does not forget the knowledge learned before, which helps the model better maintain its adaptability to previous tasks. Incremental learning only involves the update of new data, without the need to retrain the entire model, thus saving storage space and enabling the model to use computing resources more efficiently. The present invention adopts an incremental learning method, by mixing representative data from old document categories with new document category data, and incorporating them into the model for training together. This method can effectively utilize old data while also ensuring that the model adapts to new data.
[0094] Step S3 may specifically include:
[0095] Step S31: Use the flock selection algorithm to select representative training data in the old category.
[0096] The herding selection algorithm is a data screening method based on group behavior, which selects representative training data by simulating or utilizing the herding effect. Specifically, in the training data of each old category, a sorted list of samples of the category is generated according to the distance between the sample and the mean sample of the category to which it belongs. In this sorted sample list, the first u samples in the list are selected. These samples are most representative of the category to which they belong based on the mean. In the present invention, a sample set is set to store the data of the old category, and the size is fixed to 2000. The current number of old categories is recorded as v, and the following relationship exists:
[0097]
[0098] in, To round down the symbol, u represents the number of samples selected from each old category for partial training data.
[0099] Step S32: The representative training data of the old category obtained in step S31 together with all the training data of the new category constitute a training data set for incremental learning.
[0100] The Chinese document classification method based on incremental learning of the present invention further includes the following steps:
[0101] Step S4, input the incremental learning data set into the benchmark model and perform training under the constraints of the incremental learning loss function.
[0102] Specifically, the incremental learning dataset is input into the benchmark model, and the model is trained under the constraints of the incremental learning loss function, which includes: the cross entropy loss function with the true label and the decoupled distillation loss function between the old model output and the new model output.
[0103] In the present invention, the fully connected layer output of a training sample is recorded as z, and the formula is:
[0104] z=[z1,z2,z3,…,z i , …, z C ];
[0105] Among them, z i Represents the value corresponding to the i-th category in the output of the fully connected layer, and C is the number of categories.
[0106] z i The value after the softmax function is recorded as p i, which means the probability of the output being the i-th class, and the formula is:
[0107]
[0108] Among them, exp( ) is the exponential function, exp(z i ) represents the z of e i j represents the data category, and its value range is from 1 to C.
[0109] When the class label of a sample is t, the probability that the sample is output as t after being input into the model is called the probability of the target class, denoted by p t ; The probability of other categories instead of t is called the probability of non-target class, denoted by p \t . p t and p \t The formula is as follows:
[0110]
[0111] Among them, z t Represents the value corresponding to the target class in the output of the fully connected layer. k represents the data category, and k represents the remaining categories except the target class. j represents the data category, and the value range is 1 to C. k Represents the value corresponding to the kth class in the output of the fully connected layer, z j Represents the value corresponding to the jth class in the output of the fully connected layer.
[0112] In order to model the probability independently between non-target classes (not considering t classes), the present invention defines The formula is as follows:
[0113]
[0114] Among them, the value range of j is j∈{1, 2,…, t-1, t+1,…, C}, that is, j is the remaining categories except the target class.
[0115] The distillation loss formula used in the present invention is the decoupled distillation loss, denoted as l DKD , the formula is:
[0116] l DKD =αl TCKD +βl NCKD ;
[0117] Among them, α and β are the decoupled coefficients used to balance l TCKD and l NCKD The importance of TCKD It is called the target class knowledge distillation loss, which represents the similarity between the binary probabilities of the new and old models in the target class. NCKDIt is called the non-target class knowledge distillation loss, which represents the similarity between the probabilities of the new and old models in the non-target classes. TCKD and l NCKD The specific formula is as follows:
[0118]
[0119] in, and Representing the old model and the new model respectively. represents the probability that the fully connected layer outputs the target class in the old model, Represents the probability that the fully connected layer output is the target class in the new model. represents the probability that the fully connected layer outputs a non-target class in the old model, Represents the probability that the fully connected layer output is a non-target class in the new model. Indicates that the Indicates that the new model is calculated
[0120] In the present invention, the incremental learning loss function used in the training model includes the decoupled distillation loss and cross entropy loss. The incremental learning loss function is denoted as l, and the formula is:
[0121] l=λl DKD +(1-λ)l CE ;
[0122] Among them, λ is the coefficient used to balance the decoupled distillation loss and cross entropy loss. DKD is the decoupled distillation loss, l CE is the cross entropy loss, and the formula is:
[0123]
[0124] Among them, δ t=i is the indicator function, p i is the probability that the output of the fully connected layer is the i-th category.
[0125] The Chinese document classification method based on incremental learning of the present invention further includes the following steps:
[0126] Step S5, using the weight alignment method to update the weight of the new category of the output of the fully connected layer in the baseline model.
[0127] like Figure 6 As shown, the method used in the present invention consists of two stages. In the first stage ( Figure 6The present invention trains a new model on new data and representative old data, where the loss function is a combined loss including the decoupled distillation loss l DKD and the cross entropy loss l CE In the second phase ( Figure 6 The present invention uses a weight alignment method to correct the bias weights in the training model. Represent the fully connected layer outputs of the current new model and the old model, o corrected Represents the fully connected layer output after correction using the weighted approach (WA) method.
[0128] Step S5 specifically includes the following steps:
[0129] Step S51: Calculate the norms of the old category and the new category weights in the fully connected layer weights respectively.
[0130] In the present invention, the fully connected layer weight W is defined as the following formula:
[0131] W=(W old , W new );
[0132] Among them, W old and W new Represent the weight of the old category and the weight of the new category in the fully connected layer weight. old and W new It can be expressed as the following formula:
[0133]
[0134] Among them, C old Represents the number of old categories, C new Represents the number of newly added categories. {1, 2, ..., C old} indicates the old category, {C old +1, C old +2,…,C old +C new} represents a new category. w1 refers to the weight of the first category in the fully connected layer weight, and w2 refers to the weight of the second category in the fully connected layer weight. Refers to the weight of the Cth in the fully connected layer old The weight of each category, Refers to the weight of the Cth in the fully connected layer old +1 category weight, Refers to the weight of the Cth in the fully connected layer old +C new The weight of each category.
[0135] In the present invention, the norm of the old category weight in the fully connected layer weight is denoted as Norm old . The norm of the new category weight in the fully connected layer weight is recorded as Norm new . Norm old and Norm new The formula is as follows:
[0136]
[0137] Among them, || || represents the norm operation. Refers to the weight of the Cth old The weight of each category is the norm, Refers to the weight of the Cth old +1 category weights are taken as norms, Refers to the weight of the Cth old +C new The weight of each category is the norm.
[0138] Step S52: performing an averaging operation on the norms obtained in step S51.
[0139] In the present invention, the average value of the norm of the old category weights in the fully connected layer weights is recorded as Mean (Norm old ). The average value of the norm of the new category weight in the fully connected layer weight is recorded as Mean (Norm new ).
[0140] Step S53: Update the new category weights in the fully connected layer weights.
[0141] In the present invention, the Mean (Norm old ) and Mean(Norm new ) is recorded as γ, and the formula is as follows:
[0142]
[0143] The updated new category weight in the fully connected layer weight is recorded as The formula is as follows:
[0144]
[0145] Among them, W new is the weight of the new category in the fully connected layer weight.
[0146] In the present invention, the original fully connected layer output o(x) of the trained model can also be expressed as:
[0147]
[0148] Among them, old (x) indicates that the original fully connected layer outputs the value of the old category, o new (x) represents the value of the original fully connected layer output as the new category. and Represent the transposition of the old category weight and the transposition of the new category weight in the fully connected layer weight. φ(x) is the value of the Chinese document embedding matrix x after passing through three convolution functions with different convolution kernels, and then passing through the activation layer, the maximum pooling layer and the attention layer in sequence, and concatenating the obtained attention layer results by column.
[0149] The corrected fully connected layer output is recorded as o corrected (x), the formula is as follows:
[0150]
[0151] in, is the updated new category weight in the fully connected layer weight. As can be seen from the above formula, the final effect of the weight alignment method is to rescale the fully connected layer output to the value of the new category through the ratio coefficient γ.
[0152] The Chinese document classification method based on incremental learning of the present invention further includes the following steps:
[0153] Step S6: input the Chinese document to be classified into the model after the weight alignment of the fully connected layer, perform sequence information recognition, and obtain the document classification result.
[0154] For example, "The core issues of vocational education and teaching reform in the new era. Since the new era, my country's vocational education reform has entered a critical stage of connotation improvement. In the overall picture of vocational education reform,..." is a Chinese document of unknown category. This document is input into the model as the data to be classified. It passes through the encoding layer, convolution layer, activation layer, maximum pooling layer, attention layer and fully connected layer in turn, and the output category is "politics". Subsequently, the weight alignment method is applied, and this document is input into the model after the weight alignment of the fully connected layer. At this time, the output document category is "education". After verification, the document does belong to the education category. Therefore, it can be seen that the model's category recognition of the Chinese document sequence is successful, which reflects the effectiveness of the weight alignment method in correcting bias.
[0155] The beneficial effects of the present invention are:
[0156] Since the present invention crawls data from Baidu Academic System based on the existing data set as a reliable data source, strict filtering conditions are used to ensure the accuracy of the data; the Chinese pre-training model is used to obtain the embedded representation of the Chinese document text sequence, which can more effectively capture the semantic information in the text; the incremental learning method is applied to the field of Chinese document classification, two methods for solving "catastrophic forgetting" are used and combined with the TextCNN network model to identify the Chinese document sequence information, so that the classification results of the model are more accurate and reliable. Therefore, the present invention has the advantages of novel scheme and accurate results.
[0157] In the present invention, data is crawled from the Baidu academic system on the basis of existing data sets, and pre-processing operations such as strict screening are performed on the data to construct a benchmark data set for the classification of Chinese documents. The Chinese document data set used in the present invention is a valuable resource contributed by Mr. Li Ronglu of the Natural Language Processing Group of the International Database Center of the School of Computer Science and Technology of Fudan University. Baidu Academic is a powerful and resource-rich academic resource search platform that provides comprehensive, convenient and efficient academic services for scientific researchers. The present invention obtains data from the Baidu academic system, and these data lay the foundation for the research of document classification. The document data obtained using crawler technology is merged with the document data in the Chinese text classification data set of Fudan University by category, and then the number of statistics and strict screening are performed. In short, through these pre-processing operations, a high-quality benchmark data set for the classification of Chinese documents is constructed.
[0158] The present invention designs an efficient benchmark model for Chinese literature classification based on incremental learning. The model consists of a coding layer, a convolutional layer, an activation layer, a maximum pooling layer, an attention layer and a fully connected layer, which can comprehensively and systematically identify the sequence information of Chinese literature. In particular, the present invention innovatively introduces an attention layer in the model, which can assign different weights to each feature after maximum pooling, ensuring that key features receive higher attention, thereby significantly improving the classification performance of the model.
[0159] In the present invention, the model is trained using a decoupled distillation loss combined with a cross entropy loss. By adjusting the parameters of the two parts after the decoupling of the original distillation loss, the model can better retain old knowledge while learning new knowledge. The present invention also introduces a weight alignment method to overcome "catastrophic forgetting", updating the weights of the fully connected layer output for the new category, and keeping the weights of the fully connected layer output for the old category unchanged, so that the output result of the classification model has higher accuracy and reliability.
[0160] The Chinese document classification method based on incremental learning proposed by the present invention uses a flock selection algorithm to select representative data from the old category. The new category data is learned together with the representative old category data. The incremental learning model learns knowledge under the constraints of the incremental learning loss function, which can not only learn the information of the new category, but also maintain the classification performance of the old category test data. In this way, a more accurate, efficient and low-storage Chinese document classification method is achieved.
[0161] The present invention can adapt to the characteristics of the continuous growth and change of literature data, optimize the classification model by incrementally learning new data, and improve the classification accuracy and efficiency. This is of great significance for quickly identifying and protecting knowledge achievements and promoting the effective use of literature information. It should be understood that the above general description and detailed description in this application are only exemplary and explanatory, and cannot limit this application.
Claims
1. A Chinese document classification method based on incremental learning, characterized in that: include: Step S1, constructing a benchmark dataset for Chinese document classification problem; Step S2, constructing a benchmark model for Chinese literature classification based on incremental learning; Step S3, using the flock selection algorithm to select representative samples from the old category and merge them with the new category data to construct an incremental learning data set; Step S4, inputting the incremental learning data set into the benchmark model and performing training under the constraints of the incremental learning loss function; Step S5, using a weight alignment method to update the weight of the new category of the output of the fully connected layer in the baseline model; Step S6: input the Chinese document to be classified into the model after the weight alignment of the fully connected layer, perform sequence information recognition, and obtain the document classification result.
2. A Chinese document classification method based on incremental learning as claimed in claim 1, characterized in that: The step S1 specifically includes: Step S11: Download the Chinese literature classification data set and statistically analyze the data to determine whether data needs to be filled and determine the categories of data that need to be filled; Step S12: using the category of data to be filled obtained in step S11 as a search condition, searching in an academic literature search system, and obtaining a corresponding Chinese literature list as a search result; Step S13: using a Selenium-based program code to automatically crawl the title, abstract and keyword data of the Chinese literature list in the search results obtained in step S12 to obtain the literature data that needs to be filled in the data category; Step S14: screening the document data that needs to be filled with data categories crawled in step S13, and using the screened data to fill the data in the original Chinese document data set by category to obtain a filled Chinese document classification data set; Step S15: Remove spaces and stop words from the Chinese document data in the filled data set obtained in step S14 to obtain a benchmark data set for the Chinese document classification problem.
3. A Chinese document classification method based on incremental learning as claimed in claim 1, characterized in that: The step S2 specifically includes: Step S21: Set the encoding layer and use the BERT-wwm-ext, Chinese pre-trained language model to obtain the embedded representation of Chinese literature text data as the input of the subsequent model; Step S22: setting a convolution layer, using convolution kernels of different sizes to slide on the text vector to capture local features of different lengths; Step S23: setting an activation layer and using a ReLU activation function to activate the features obtained after convolution; Step S24: define a maximum pooling layer, perform dimensionality reduction processing on the features, and extract the most important features; Step S25: Set the attention layer, obtain the weight of each feature, and then multiply it with the original feature; Step S26: define a fully connected layer to obtain the classification result of the model.
4. A Chinese document classification method based on incremental learning as claimed in claim 3, characterized in that: The step S21 specifically includes: Step S211: Acquire Chinese document data in the benchmark data set; Step S212: Input the Chinese document data into the BERT-wwm-ext, Chinese pre-trained language model, and obtain the hidden state of the last layer in the model as the embedded representation vector of the Chinese document.
5. A Chinese document classification method based on incremental learning as claimed in claim 4, characterized in that: In step S212: A Chinese document sequence is recorded as S. The Chinese document is input into the BERT-wwm-ext, Chinese pre-trained language model. The model uses the word segmentation algorithm to segment the text. The result after word segmentation is recorded as g(S). The formula is as follows: <h2 style=";text-align:left;direction:ltr">g(S) = {a1, a2, a3,..., a<h2 style=";text-align:left;direction:ltr"> m <h2 style=";text-align:left;direction:ltr"> ,...,a<h2 style=";text-align:left;direction:ltr"> L <h2 style=";text-align:left;direction:ltr">}; Among them, a m Indicates the mth token after word segmentation (the smallest unit into which text data is segmented during model processing), where m is a subscript, meaning the number; L represents the number of tokens, which is 100 and is controlled by the max_length and padding parameters in the model; After that, the hidden state of the last layer in the model is obtained as the embedding representation vector of the Chinese document. The embedding representation vector of a Chinese document is recorded as f(S), and the formula is as follows: f(S)={e1,e2,e3,...,e m ,...,e L }; Among them, e m represents the embedding representation corresponding to the mth token, with a dimension of 768, which is set by the hidden_size parameter in the model. In BERT-wwm-ext, Chinese, hidden_size is set to 768, which means that the hidden layer output dimension size is 768. L represents the number of tokens. e L Represents the embedding representation corresponding to the Lth token.
6. A Chinese document classification method based on incremental learning as claimed in claim 5, characterized in that: The step S22 specifically includes: For text data, the convolution kernel can be regarded as a series of weights that slide over the text and perform weighted summation on the data in each window to generate new features. The output after convolution is recorded as y(x), and the formula is as follows: y(x)=A*x+b; Where x is the text embedding matrix, whose dimension is (batchsize, max_length, hidden_size), batchsize is the batch size, which is set to 128; max_length is the sequence length, which is 100, and hidden_size is the hidden layer output dimension size, which is 768; A is the weight of the convolution kernel, and its dimension is defined as (h, d); h is the size of the convolution kernel, which defines the number of words covered by the convolution kernel; d is the same dimension as the text embedding; b is the bias term; in the formula for the convolution operation, the multiplication involved in the symbol * is the process of multiplying the corresponding positions and then summing them.
7. A Chinese document classification method based on incremental learning as claimed in claim 6, characterized in that: The step S23 specifically includes: Use the ReLU function as the activation function and record the result after the ReLU activation function as R(x). The formula is as follows: R(x)=max(0,y(x)); Among them, when the convolution value y(x) is greater than 0, the value itself is output; when the convolution value y(x) is less than or equal to 0, the output is 0; x is the text embedding matrix.
8. A Chinese document classification method based on incremental learning as claimed in claim 7, characterized in that: The step S23 specifically includes: The result after maximum pooling is recorded as M(x), and the formula is as follows: M(x) = maxpooling(R(x)); Among them, M(x) means extracting the most important features from the feature map generated by each convolution kernel, R(x) is the result after the ReLU activation function, and x is the text embedding matrix.
9. A Chinese document classification method based on incremental learning as claimed in claim 1, characterized in that: The step S3 specifically includes: Step S31: Use the flock selection algorithm to select representative training data in the old category; Step S32: The representative training data of the old category obtained in step S31 together with all the training data of the new category constitute a training data set for incremental learning.
10. The Chinese document classification method based on incremental learning as claimed in claim 1, characterized in that: The step S5 specifically includes: Step S51: Calculate the norms of the old category and the new category weights in the fully connected layer weights respectively; Step S52: performing an average operation on the norms obtained in step S51; Step S53: Update the new category weights in the fully connected layer weights.
Citation Information
Patent Citations
Incremental relation extraction method based on knowledge distillation
CN115203404A
Target detection incremental learning method and device for finite model space and medium
CN117372819A
Prompt augmented generative replay via supervised contrastive training for lifelong intent detection
US20240013094A1
Few shot incremental learning for named entity recognition
US20240362419A1