A Small-Sample Text Classification Method Based on Domain Template Pretraining
Through the methods of domain template pre-training and prompt learning, the limitations of data volume and computing resources in text classification are solved, and efficient small sample text classification is achieved, which improves classification accuracy and model adaptability.
Patent Information
- Application Number
- CN202211598846.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-14
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-12-14
AI Technical Summary
The existing text classification methods in the field of natural language processing require a large amount of labeled data and high-performance computing resources, and there is a domain barrier between the pre-trained language model and the target task, resulting in complex training and poor results.
A combination of domain template pre-training and prompt learning is adopted to build templates using data sets related to the target task, further pre-training is performed through pre-training language models, and a small sample text classification is used to reduce data requirements and hardware requirements.
It improves the accuracy and training speed of text classification, reduces the dependence on computing resources, and enhances the adaptability and classification effect of pre-trained models on target tasks.
Smart Images

Figure CN115840820B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and specifically to a small sample text classification strategy based on domain template pre-training and improved Prompt. Background Art
[0002] With the continuous development of natural language processing, model algorithms for text classification tasks emerge in an endless stream, from probability-based machine learning models to deep learning models composed of deep neural networks. Although these model methods have gradually improved the classification accuracy, these models generally start training directly from scratch on the task dataset, requiring a large amount of labeled data and high-performance processors, and also requiring a large amount of training time. In addition, the models trained have poor adaptability to new tasks, and for new tasks, it is often necessary to re-label data and train the model again. The small sample learning method based on the pre-trained model has developed rapidly in recent years and can well solve the above problems. Applying the small sample learning method based on the pre-trained model to text classification is of research value. The small sample learning method based on the pre-trained model can well obtain general common language representation knowledge and model initialization parameters from a large number of unlabeled datasets, and then use very little data for training in the target task to achieve very good results.
[0003] At present, the methods of text classification in the field of natural language processing are mainly divided into two categories: those based on deep learning models and those based on pre-trained language models. Classic methods based on deep learning models include, for example, the recurrent neural network (RNN) based on neurons (Mikolov T, Karafiát M, Burget L, et al. Recurrent neural network based language model [C]. Interspeech. 2010, 2(3): 1045-1048); the improved long short-term memory neural network (LSTM) (Zhang, Yanbo. Research on Text Classification Method Based on LSTM Neural Network Model. 2021 IEEE Asia-Pacific Conference on Image Processing, Electronics and Computers (IPEC) (2021): 1019-1022); and the text convolutional neural network model (TextCNN) (Kim, Yoon. Convolutional Neural Networks for Sentence Classification [C]. Empirical Methods in Natural Language Processing, 2014: 1746-1751). For training corpora with a large number of annotations, these models can repeatedly adjust model parameters through training and achieve good classification results. However, methods based on deep learning all require training the model from scratch and need a large amount of training data sets to establish the mathematical mapping relationship between input variable X and output variable Y. In actual application scenarios, due to privacy and security issues or the cost of collecting annotations, a large amount of data cannot be obtained to train and learn the model, or it is difficult to train a deep learning model due to limitations in computer hardware levels. Therefore, in the case of resource constraints, the performance and effects of such methods are not satisfactory.With the emergence of Transformer (Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[J]. Neural Information Processing Systems, 2017, (30): 6000 - 6010) and Bert (Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[J]. Neural Information Processing Systems, 2017, (30): 6000 - 6010), the development of the new round of natural language field has been accelerated. The methods based on pre - trained language models can be divided into the fine - tuning strategy of making the model adapt to the task and the prompting learning strategy of making the task adapt to the model. The pre - training and fine - tuning solutions first design training objects for downstream tasks based on pre - trained language models and fine - tune the models to obtain the semantic information of the corpus and the initialization parameters of the pre - trained models, enabling the models to adapt to various downstream tasks. However, due to the inconsistent goals between pre - trained language models and downstream tasks, there are often barriers between domains, structural deviations between input and output, complex fine - tuning designs, and high optimization costs. The prompting learning method based on pre - trained models can fully exploit the potential of pre - trained language models. By adding a prompt description to data reconstruction, the task is transformed into a cloze task familiar to pre - trained language models. Without the need to redesign the classifier, only by designing different Prompts, the target task can be adapted to the pre - trained language model, and good classification results can be demonstrated.
[0004] Although the current natural language processing field is developing rapidly and there are a large number of excellent algorithms for text classification task research, there are still some unsolved problems. For example, the designs of pre - trained language models and fine - tuning methods are becoming increasingly complex, the difference between the target task and the pre - trained language task domain is too large, it is difficult for pre - trained language models to learn domain - specific knowledge, there are structural deviations between the input and output of pre - trained language models and the target task, semantic information is lost in the data processing of the dataset, redundant information, etc. How to reasonably obtain domain data information and fully improve the performance of pre - trained models on target tasks is still one of the key issues to be studied in the field of small - sample text classification. Summary of the Invention
[0005] The object of the present invention is to provide a small-sample text classification method based on domain template pre-training in view of the deficiencies of the prior art. The method combines domain template pre-training and prompt learning to perform the small-sample text classification task. The method constructs templates using the target data set, trains the pre-trained language model for the MLM task, and through the construction of a hybrid template and multi-label mapping of the obtained in-domain information, it can achieve better classification with less data, not only shortening the training time of the target task, but also reducing the requirements for the computer hardware performance. This method constructs templates using the in-domain data set related to the target task, then further pre-trains the pre-trained language model with the constructed data, constructs a hybrid template for the target task data set, and preprocesses the data of the target data set. Then, it uses the further pre-trained model to train and validate the target task to obtain the predicted words, and uses a label word mapper to map the predicted words to the final target labels. The method is simple, has a faster training speed, lower requirements for hardware performance, makes better use of the pre-trained language model, greatly improves the classification accuracy of the target task, and provides technical support for the technical development of related fields.
[0006] The specific technical solution for achieving the object of the present invention is: a small-sample text classification method based on domain template pre-training, which is characterized by further pre-training the pre-trained language model through domain template construction, and performing a construction of improved prompt learning on the target task to classify small-sample text. The method mainly includes the following steps:
[0007] Step 1: Construct a prompt template through the in-domain data set related to the target task. If the input data is X, and fprompt is a prompt function used to add prompt information, it is constructed into x defined by the following formula (a):
[0008] x = fprompt(x) (a).
[0009] Wherein, x is the text data of the in-domain data set; fprompt is a template construction function; x' is the data of the in-domain data set after template construction.
[0010] Step 2: Use the data after template construction in Step 1 to further pre-train the selected pre-trained language model for the MLM task, so that the pre-trained language model obtains in-domain information related to the target task.
[0011] Step 3: Take the same number of data samples for each category of the target task, and perform head and tail truncation processing on the long texts in the data set to obtain summary semantic information at the head and tail, and perform dynamic padding on the short texts to reduce a large amount of useless padding.
[0012] Step 4: Construct a prompt-mixed template for the target dataset by using a human-understandable natural language template and a machine-understandable encoded language template. Using the training data Xtarget of the target task, the template is designed as {soft:This}topic{soft:is}{mask}{Xtarget}, where soft is an adjustable template understandable by the machine and is initialized according to the task, topic is the natural language template that will be converted into the corresponding Embedding form, mask is the value to be predicted, and Xtarget is the original input sequence.
[0013] Step 5: Use the mixed template constructed from the target task dataset in Step 4 to train and predict the further pre-trained language model generated in Step 2), and adjust parameters such as the learning rate. At this time, the probability of the pre-trained language predicting the token at the Mask position is shown in the following formula (b):
[0014]
[0015] where X’ is the input data constructed through the template; Y f is the final output probability; Y is the current predicted word; Z(X) is the answer space; Y’ is the answer space excluding the current predicted word.
[0016] Step 6: Obtain the predicted answer with the highest probability through the argmax function, and then according to the answer space Z, the label word mapper obtains the final output label required for the target task calculated by the following formula (c):
[0017] Ylabel = Z(argmaxY ∈% (P(Yf|X’)))(c).
[0018] where Y label is the final output label, P(Yf|X’) is the probability of the predicted word, and Z is the answer label mapper that maps the predicted word to the final output label.
[0019] The present invention has the following remarkable technical progress and beneficial technical effects compared with the prior art:
[0020] 1) The present invention constructs prompt information by using data in the field and further pre-trains the pre-trained language model, fully obtaining semantic information and domain knowledge related to the target field.
[0021] 2) The present invention transforms the target task to adapt to the pre-trained language model. When constructing a template for the target task, a mixed template method is used to give full play to the different advantages of the prompt template and reduce the cost of template construction.
[0022] 3) In the research of small-sample text classification of the present invention, it has higher accuracy and shorter training time compared with the deep learning method using a larger data set and the method using pre-training and fine-tuning. Description of the Drawings
[0023] Figure 1 It is a flow chart of the present invention. Detailed Implementation Manner
[0024] The present invention will be further described and explained in detail below with specific embodiments:
[0025] Refer to Figure 1 , the present invention performs small-sample text classification according to the following steps:
[0026] Step 1: Construct a prompt template through a domain data set related to the target task. If the input data is X, after passing through fprompt as a prompt function to add prompt information to construct it into X', for example, for the input sentence: I like this toy, after template construction, it is: It is [Mask], I like this toy.
[0027] Step 2: Use the data constructed by the template in Step 1 to further pre-train the selected pre-trained language model for the MLM task, so that the pre-trained language model obtains domain information related to the target task.
[0028] Step 3: Take the same number of data samples for each category of the target task, and perform head and tail truncation processing on the long texts in the data set to obtain summary semantic information at the head and tail, and perform dynamic padding on the short texts to reduce a large amount of useless padding.
[0029] Step 4: Construct a prompt hybrid template for the target data set by using a natural language template understandable by humans and an encoding language template understandable by machines. Use the training data Xtarget of the target task, which is designed as {soft:This}topic{soft:is}{mask}{Xtarget} through the template. soft is an adjustable template understandable by machines and is initialized according to the task. topic is a natural language template that will be converted into the corresponding Embedding form. mask is the value to be predicted, and Xtarget is the original input sequence.
[0030] Step 5: Use the template constructed from the target task dataset in Step 4 to train and predict the further pre-trained language model generated in Step 2, and adjust parameters such as the learning rate. The warm-up step is set to 0.1 of the total number of steps, the model update rate is 0.00002, and the weight decay is set to 0.01. At this time, the probability of the pre-trained language model predicting the token at the Mask position is calculated as shown in the following formula (b):
[0031]
[0032] where X’ is the input data after template construction; Y f is the final output probability; Y is the current predicted word, Z(X) is the answer space, and Y’ is the answer space excluding the current predicted word.
[0033] Step 6: Obtain the predicted answer with the highest probability through the argmax function, and then according to the answer space Z, the label word mapper obtains the output label required for the final target task calculated by the following formula (c):
[0034] Ylabel = Z(argmax ∈% (P(Y|X’)))(c).
[0035] where Y label is the final output label; P(Y|X’) is the probability of the predicted word; Z is the answer label mapper.
[0036] Map the predicted word to the final output label. For example, in the case of binary classification, the answer space and label words may be (positive: wonderful, great, interesting...; negative: bad, boring, terrible...).
[0037] The present invention constructs a template using a domain dataset related to the target task and further pre-trains the pre-trained language model, enabling the pre-trained model to better learn domain information related to the target task, greatly reducing the domain gap between the pre-trained language model and the target task. By using the Prompt method of template construction, an input-output structure similar to that during the training of the pre-trained language model is constructed, enabling the target task to adapt to the pre-trained language model, fully leveraging the potential of the pre-trained language model. Moreover, by constructing a mixed template using human-readable natural language and machine-understandable coding language, the embeddable words that can be optimized are enhanced, improving the representation ability of the template. In addition, by using an answer space mapper to expand the answer space and enhance the robustness of the model, the method of the present invention can achieve good results in small-sample text classification, greatly reducing the dependence on the amount of target task data and the requirements for computer performance.
[0038] The above are only specific embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A small-sample text classification method based on domain template pre-training, characterized in that A method for further training a pre-trained language model used with a domain dataset related to a target task, which performs small-sample text classification through data preprocessing, parameter processing, hybrid templates, and multi-label mapping, specifically including the following steps: 1) Use a dataset related to the target task domain to construct a prompt template to obtain domain data; 2) Use the pre-trained language model selected for the domain data with MLM as the target task for pre-training to generate a further pre-trained language model; 3) Perform class-balanced sampling on the training dataset, truncate the long text to the same length at the beginning and end, and perform dynamic padding on the short text; 4) Use the target dataset to construct a hybrid template combining discrete templates and continuous templates to construct a prompt hybrid template; 5) Use the generated further pre-trained language model to train and predict the target task, adjust the learning rate parameter, and obtain a predicted answer; 6) Use a multi-label mapper according to the predicted answer to perform label mapping conversion of the actual labels of the target task on the words predicted by the model according to the answer space to obtain the final output label, realizing small-sample text classification; The step 1) constructs a prompt template using a dataset related to the target task domain. If the input data is X, use as a prompting function to add prompt information, and construct it into the defined by the following formula (a): (a); Among them, x is the text data of the domain dataset; is the template construction function; x' is the data of the domain dataset after template construction; In step 4), a prompt hybrid template is constructed for the target dataset using a natural language template understandable by humans and an encoded language template understandable by machines. The training data Xtarget of the target task is designed through the template as {soft:This} topic {soft:is}{mask} {Xtarget}, where soft is an adjustable template understandable by machines and is initialized according to the task; topic is a natural language template in which the language is converted into the corresponding Embedding form; mask is the value to be predicted; Xtarget is the original input sequence.
2. The small-sample text classification method based on domain template pre-training according to claim 1, wherein In step 5), the generated further pre-trained language model is used to train and predict the target task. The warm-up step of the training is set to 0.1 of the total number of steps, the model update rate is 0.00002, and the weight decay is set to 0.01; the probability of predicting the token at the Mask position using the pre-trained language prediction in the prediction is calculated by the following formula (b): (b); Among them, X’ is the input data after template construction; Y f is the final output probability; Y is the current predicted word; Z(X) is the answer space; Y’ is the answer space excluding the current predicted word.
3. The small-sample text classification method based on domain template pre-training according to claim 1, wherein In step 6), the predicted answer with the highest probability is obtained through the argmax function, and then according to the answer space Z, the label word mapper obtains the output label required for the final target task calculated by the following formula (c): (c); Among them, Y label is the final output label; is the probability of the predicted word; Z is the answer label mapper that maps the predicted word to the final output label.
Citation Information
Patent Citations
Chinese short text classification method based on prompt learning
CN115169340A
Small sample text classification method for prompting learning
CN115455181A