An Open-Domain Small-Sample Text Learning Method Based on Active Learning

Through the integration of active learning and small sample learning, and using feature coding and clustering analysis, the problems of multiple and slow iterations in the text classification of small sample texts are solved, and the rapid iteration and efficient training of the model are achieved.

CN115344696BActive Publication Date: 2025-07-08TENTH RES INST OF TELECOMM TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210927182.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-03
Publication Date
2025-07-08
Estimated Expiration
2042-08-03

AI Technical Summary

Technical Problem

The existing technology requires a lot of manual annotation in the classification of small sample texts, which affects the efficiency of model online. The existing methods are slow in model training iteration, making it difficult to quickly implement and apply.

Method used

The open-domain small sample text learning method based on active learning is adopted, and through feature coding, clustering analysis and multiple iterations, active learning and small sample learning are integrated to reduce the number and number of manual annotations, and the advantages of small sample learning are used to accelerate model iteration.

Benefits of technology

It realizes the rapid iteration of required models in different fields, reduces manual annotation work, improves model training efficiency, and quickly achieves optimization results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115344696B_ABST
    Figure CN115344696B_ABST
Patent Text Reader

Abstract

The present invention discloses an open-domain small-sample text learning method based on active learning. First, the characteristics of the small-sample text data are encoded and the small-sample model is initialized; then the active learning algorithm is used to obtain the correct data set and candidate set data, and the candidate data set is encoded; then the encoded candidate data set is subjected to clustering analysis to obtain the optimal number of clustering clusters; the clustering clusters with the optimal number are reclustered to identify the optimal cluster; after annotation, new-category text data and small-sample text incremental data are generated, and the correct data set obtained by active learning, the new-category text data, and the small-sample text incremental data are added to the small-sample text data set; the above steps are repeatedly executed until a sufficient text data set is finally obtained. The present invention combines active learning with small-sample learning, takes advantage of small-sample learning, and through multiple iterations of active learning, reduces the quantity and frequency of manual annotation, so that the model can be quickly implemented and applied.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine learning, and particularly relates to an open-domain small-sample text learning method. Background Art

[0002] For text classification in natural language processing, in engineering applications, the data sources used specifically include texts with unknown categories and only a dozen text data to be classified. The most common method is to sample a part of the data, manually label the sampled data, and preliminarily classify the data source based on the sampled data. Through continuous manual labeling, when each type of data accumulates to about ten thousand, a relatively excellent text classification model can be trained. The advantage of this method is its ease of implementation, while the disadvantage is that it requires a large amount of manpower for manual labeling, and manual labeling will greatly affect the efficiency of model deployment. To solve this problem, the academic and industrial communities have adopted two methods. The academic community has adopted the method of few-shot learning, and through a series of few-shot learning algorithms, trained a model with relatively high accuracy as the application model; the industrial community has chosen to use the method of active learning, and through continuous rapid iteration of the model, enables the model to converge to an excellent model in a shorter time.

[0003] Background and learning process of active learning technology:

[0004] 1. Use a small amount of sample data to train an initial model;

[0005] 2. Use the initial model to predict the data to be labeled and calculate the data that most needs to be labeled;

[0006] 3. Manually label the predicted data;

[0007] 4. Combine the labeled data with the small-sample data to train a new model;

[0008] 5. The above process is iterated several times to make the model achieve the optimal effect.

[0009] 6. For how to calculate the data that most needs to be labeled, the commonly used algorithms at present are: query based on committee, voting entropy, average KL divergence, expected model change, expected error reduction, variance reduction, selection based on density weight.

[0010] Background of few-shot learning technology:

[0011] Few-shot learning methods can be divided into three categories: data augmentation, model training optimization, and gradient descent algorithm optimization.

[0012] The data augmentation method is a method of increasing sample data.

[0013] 1. In the field of natural language processing, samples can be increased by means of synonym replacement, insertion, and deletion.

[0014] 2. Search for a dataset similar to the small sample for supplementary replacement.

[0015] Model training optimization is to achieve learning of small sample data through the model structure.

[0016] 1. Multi-task learning: Integrate multiple small sample learning tasks into a multi-task learning with sufficient samples and share the parameters among them to achieve small sample learning.

[0017] 2. Representation learning: Learn samples through general prior knowledge and then apply specific representation methods to specific applications.

[0018] 3. Generative model method: Generate samples through a generative model to expand the small sample dataset.

[0019] Gradient descent algorithm optimization is to learn the parameter update algorithm based on gradient descent. Through the learning of the optimization algorithm, the model can achieve generalization with the shortest number of iterations under the problem of small samples, and reduce the problem of overfitting of the small sample model. The advantage of gradient descent algorithm optimization is that it does not require as high a requirement for the design of the training model as the model training optimization method.

[0020] Disadvantages of active learning:

[0021] Too few training samples will still increase the workload of manual annotation and affect the iterative efficiency of model training.

[0022] Disadvantages of small sample learning:

[0023] 1. In the case of extremely small samples, data augmentation methods such as using synonym replacement, insertion, and deletion to expand text content, or using similar datasets to expand text content. However, for special industries, it is very difficult to find similar datasets using the similar dataset method.

[0024] 2. The methods of model training optimization are not universal for different artificial intelligence applications and have a high technical threshold for model optimization. Secondly, the models obtained by model training optimization methods are not yet ready for practical application.

[0025] The optimization method of the gradient descent algorithm is to set a specific optimization function by analyzing the gradient of the loss value of small sample data during model training. This optimization method also requires a high technical threshold. The optimization of gradient descent can only ensure that the model reaches the optimal under the current training conditions, that is, the local optimum. It cannot reach the application level in real applications.

[0026] Currently, the most effective method for optimizing artificial intelligence models is to increase the number of training samples. However, a large amount of manual annotation work requires professional requirements for annotators and involves a huge volume of annotations. Summary of the Invention

[0027] To overcome the deficiencies of the prior art, the present invention provides an open-domain small-sample text learning method based on active learning. First, encode the feature of the small-sample text data and initialize the small-sample model; then use the active learning algorithm to obtain the correct data set and candidate set data, and encode the candidate data set; then perform clustering analysis on the encoded candidate data set to obtain the optimal number of clustering clusters; recluster the optimal number of clustering clusters to identify the optimal cluster; generate new category text data and small-sample text incremental data after annotation, and add the correct data set obtained by active learning, the new category text data, and the small-sample text incremental data to the small-sample text data set; repeat the execution to finally obtain a sufficient text data set. The present invention combines active learning with small-sample learning, utilizes the advantages of small-sample learning, and through multiple iterations of active learning, reduces the quantity and frequency of manual annotation, so that the model can be quickly applied in practice.

[0028] The technical solution adopted by the present invention to solve its technical problems includes the following steps:

[0029] Step 101: Encode the feature of the small-sample text data;

[0030] Encode the data of the small-sample text data set into feature vectors: If the classification model of the small-sample text data uses a classification model with a pre-trained model, use the pre-trained model of this classification model for feature vector encoding; if the classification model of the small-sample text data does not have a pre-trained model, randomly encode to generate feature vectors;

[0031] Step 102: Initialize the small-sample model;

[0032] Input the already encoded feature vectors into the classification model of the small-sample text data, and train to obtain the small-sample model;

[0033] Step 103: Obtain the correct data set and candidate set data;

[0034] Encode the unlabeled text data through the encoding method of Step 101, input it into the small-sample model, and obtain the correct data set and the candidate data set that needs to be manually annotated through the voting entropy active learning algorithm;

[0035] Step 104: Encode the candidate data set;

[0036] Encode the candidate data set through the encoding method of Step 101;

[0037] Step 105: Perform clustering analysis on the encoded candidate dataset; find the inflection point of the sum of squared errors of clusters through multiple iterative calculations of the sum of squared errors of clusters, and obtain the optimal number of clustering clusters;

[0038] Step 106: Re-cluster the optimal number of clustering clusters, predict the small-sample text data of the existing labels, and identify the optimal cluster by finding the cluster with the most known labels among the predicted clusters;

[0039] Step 107: Label the optimal cluster;

[0040] Step 108: After discriminating and labeling the optimal cluster, the labeled data will generate text data of new categories and incremental small-sample text data. Add the correctly learned dataset, text data of new categories, and incremental small-sample text data to the small-sample text dataset;

[0041] Step 109: Set the number of repetitions, and repeatedly execute Step 101 to Step 108;

[0042] Step 110: After the repeated execution of Step 109 ends, a sufficient text dataset is obtained.

[0043] Preferably, the classification model with the pre-trained model is the BERT model.

[0044] Preferably, the classification model without the pre-trained model is the TextCNN model.

[0045] Preferably, the clustering analysis adopts KMeans clustering.

[0046] The beneficial effects of the present invention are as follows:

[0047] The present invention combines active learning with small-sample learning, utilizes the advantages of small-sample learning, and through multiple iterations of active learning, reduces the quantity and frequency of manual labeling, so that the model can be quickly implemented and applied. The present invention is not limited to a certain field of artificial intelligence at the same time, and the invention can be used in different fields to quickly iterate the required model. Description of the Drawings

[0048] Figure 1 It is a schematic diagram of the method of the present invention. Detailed Embodiments

[0049] The present invention will be further described below with reference to the drawings and embodiments.

[0050] The objective of the present invention is to efficiently initialize the model of active learning by using few-shot learning, reduce the number of manual annotations, address the problems of inaccurate recognition in few-shot learning and slow iterative speed of model training, accelerate the iteration of the model, achieve the optimal model with the least amount of manual annotation, and solve the rapid iteration of model training for various tasks in the field of artificial intelligence.

[0051] An open-domain few-shot text learning method based on active learning, comprising the following steps:

[0052] Step 101: Encoding the features of few-shot text data: Encode the few-shot text data into feature vectors that can be recognized by a classification model, such as models like TextCNN, BERT, etc. If using a classification model with a pre-trained model, such as a pre-trained BERT model that has been trained, the pre-trained model of BERT can be used for feature vector encoding. If there is no pre-trained model, feature vectors need to be randomly generated.

[0053] Step 102: Initializing the few-shot model: If using the BERT pre-trained model, initialize the BERT model; if not using the pre-trained model, initialize the TextCNN model; input the encoded feature vectors into the BERT or TextCNN model to train the initialized few-shot model.

[0054] Step 103: Obtaining the correct dataset and candidate set data: After feature encoding the unlabeled text data, input it into the initialized few-shot model. Obtain the correct dataset and the candidate dataset that requires manual annotation through the active learning algorithm of voting entropy.

[0055] Step 104: Encoding the candidate dataset: If there is a BERT model for the text data of the candidate dataset, use the BERT pre-trained model to perform feature encoding on the data of the candidate set, and the selected pre-trained model is the same as the pre-trained model in Step 101. Otherwise, encode the data in a randomly initialized manner;

[0056] Step 105: Conducting clustering analysis on the encoded candidate set data, and the clustering method uses KMeans clustering. By calculating the sum of squared errors of the clusters through multiple iterations, find the inflection point of the sum of squared errors to obtain the optimal number of clustering clusters;

[0057] Step 106: Re-clustering according to the optimal number of clustering clusters, predicting the few-shot data with existing labels, and identifying the optimal cluster by finding the method that contains the most known labels in the predicted cluster;

[0058] Step 107: Judging and annotating the optimal cluster.

[0059] Step 108: After the discrimination, the labeled data will generate new category data and small sample incremental data. The correct recognition data, new category data, and small sample incremental data learned actively will be added to the small sample data set.

[0060] Step 109: Repeat steps 101 to 108 several times;

[0061] Step 110: After step 109, a data set of sufficient samples will be obtained, and the data set of sufficient samples can be used to complete a specific task.

[0062] At present, artificial intelligence applications face the pain points of inaccurate small sample recognition and slow model iteration. The industry uses active learning to iterate and label data multiple times to make the model achieve the optimal effect, and the academic community prefers to use small sample learning to make the model optimal. Although active learning is a great improvement compared to blindly labeling data, it still requires a large amount of labeled data multiple times. Although small sample learning makes full use of small sample data to train a model with a certain effect, there is still a certain gap from the implementation of the model. The present invention integrates active learning with small sample learning, utilizes the advantages of small sample learning, and reduces the number and frequency of manual labeling through multiple iterations of active learning, so that the model can be quickly implemented. The present invention is not limited to a certain field of artificial intelligence at the same time, and different fields can use the invention to quickly iterate the required model.

[0063] The key point of the present invention is how to integrate active learning and small sample learning, and use small sample learning to initialize the active learning model. The algorithm used in small sample learning can be selected according to the specific artificial intelligence application.

[0064] 1. The entire method flow of integrating active learning and small sample learning;

[0065] 2. Active learning operator matching method;

[0066] 3. Optimal number of clusters selection method;

[0067] 4. Optimal cluster identification method.

[0068] The most direct implementation scheme for the small sample problem in the existing artificial intelligence implementation schemes is to use a large amount of manpower for manual annotation, or to use the method of active learning for manual annotation, and to use small sample learning to optimize at the model level. The implementation scheme of the present invention unifies and integrates the separate implementation schemes, combining the advantages of each scheme. For different artificial intelligence applications, both active learning and small sample learning require algorithm selection for different sample data. The present invention integrates different artificial intelligence applications, adopts a general active learning method and a small sample learning method, making the application more general. In addition to having very good effects on the small sample text classification problem, the present invention also has very good performance in solving the image classification problem and other classification problems in the field of artificial intelligence.

Claims

1. An open-domain small-sample text learning method based on active learning, characterized in that, It includes the following steps: Step 101: Encoding the feature of small sample text data; Encoding the data of the small sample text data set into feature vectors: If the classification model of the small sample text data adopts a classification model with a pre-trained model, use the pre-trained model of this classification model for feature vector encoding; If the classification model of the small sample text data does not have a pre-trained model, randomly encode to generate feature vectors; Step 102: Initializing the small sample model; Input the already encoded feature vectors into the classification model of the small sample text data, and train to obtain the small sample model; Step 103: Obtaining the correct data set and candidate set data; After encoding the unlabeled text data by the encoding method in Step 101, input it into the small sample model, and obtain the correct data set and the candidate data set that needs to be manually labeled through the voting entropy active learning algorithm; Step 104: Encoding the candidate data set; Encoding the candidate data set by the encoding method in Step 101; Step 105: Performing clustering analysis on the encoded candidate data set; By calculating the sum of squared errors of the clusters through multiple iterations, find the inflection point of the sum of squared errors to obtain the optimal number of clustering clusters; Step 106: Re-clustering the optimal number of clustering clusters, predicting the small sample text data with existing labels, and identifying the optimal cluster by finding the cluster with the most known labels in the predicted clusters; Step 107: Labeling the optimal cluster; Step 108: After discriminating and labeling the optimal cluster, the labeled data will generate text data of new categories and incremental data of small sample text. Add the correctly learned data set, text data of new categories, and incremental data of small sample text to the small sample text data set; Step 109: Set the number of repetitions, and repeat Steps 101 to 108; Step 110: After the repeated execution of Step 109 ends, obtain a sufficient text data set.

2. The open-domain small-sample text learning method based on active learning according to claim 1, characterized in that The classification model with a pre-trained model is the BERT model.

3. The method for open-domain small-sample text learning based on active learning according to claim 1, wherein The classification model without a pre-trained model is the TextCNN model.

4. A method for open-domain small-sample text learning based on active learning according to claim 1, characterized in that, The clustering analysis adopts KMeans clustering.

Citation Information

Patent Citations

  • Active learning based data automatic marking method

    CN107067025A

  • Small sample learning method in text classification

    CN112115265A