Active learning method, electronic equipment and storage medium
By combining the proactive learning methods of query committees and uncertainty criteria, screening and correcting labels that are prone to confusing samples, the problems of high labor costs and unstable training set quality in text label correction are solved, and efficient and accurate text label correction and model training are achieved.
Patent Information
- Application Number
- CN202311862329.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-04
AI Technical Summary
In the correction of text labels, existing active learning methods have problems such as high manual labeling costs, unreliable training set quality, large data scale and easy to overfit, making it difficult to correct text labels efficiently and accurately.
An active learning method is adopted, combined with the query committee and uncertainty criteria, by determining the voting entropy of samples with high consistency and confusing samples, using the query function to filter confusing samples and correct their labels, reducing the scale of the training set, using similarity matching to expand confusing samples, forming a new training set, and optimizing the text classification model.
It reduces the cost of manual tagging, improves the speed and accuracy of training models, improves the tolerance of the model to data quality, ensures the efficiency of computing time, and is compatible with multiple text classification models.
Smart Images

Figure CN120256606A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to an active learning method, an electronic device, and a storage medium. Background Art
[0002] Field retrieval is a commonly used content search means for portal websites, multimedia websites, etc. The background program determines which content to feedback or recommend to users as search results based on the tags of the fields.
[0003] The correction of field tags of text in website data is usually completed by combining active learning with manual annotation.
[0004] Currently, there are three types of active learning methods, namely, methods based on query committees, methods based on uncertainty criteria, and methods based on generalization error reduction.
[0005] The query committee method predicts and votes to select samples by forming a committee with pre-trained text classification models and submitting them to manual annotation. This method is relatively easy to implement, but to ensure more reliable prediction results of the classification model, a sufficiently large training set is required, which will increase time and labor costs.
[0006] The method based on uncertainty criteria selects annotation samples by evaluating the uncertainty of each sample. It is applicable to small-scale sample sets and highly dependent on the quality of the initial training set. When dealing with non-linear data, the complexity will increase sharply.
[0007] The method based on generalization error reduction is used to estimate the category of samples and measure the credibility of the model's estimation of the category to which the samples belong. This method has a large amount of calculation, requires evaluating the generalization error of the model before and after adding samples, is only applicable to binary classification, and is prone to overfitting problems.
[0008] Therefore, how to reduce the manual annotation cost, more accurately complete the correction of text tags, and at the same time be able to select samples with a large amount of information, reducing the data scale on which model training depends, is a technical problem to be solved urgently. Summary of the Invention
[0009] Aiming at the problem of how to improve the quality of the training set when the quality of the training set is not very reliable, the present invention proposes an active learning method for correcting text tags and supporting the annotation of a large amount of external data at the same time. This method combines the advantages of two types of methods, namely, query committee and uncertainty sample selection, and the quality of label correction is more accurate and reliable.
[0010] The present invention is realized through the following technical solutions:
[0011] An active learning method, comprising the following steps:
[0012] Determine the distribution data of the classification vote counts of the predicted labels of the samples in the text set to be tested;
[0013] Determine the samples with high consistency according to the distribution data and the voting entropy of the samples in the text set to be tested;
[0014] Input the voting entropy of each sample in the text set to be tested into the query function respectively to determine the query function value of each sample in the text set to be tested, and obtain the confusing samples;
[0015] Correct the labels of the samples with high consistency and the confusing samples.
[0016] Furthermore,
[0017] Determine the distribution data of the classification vote counts of the predicted labels of the samples in the text set to be tested, including:
[0018] Select multiple subsets from the training set, and obtain multiple text classification models through multiple rounds of training based on multiple known text classification models respectively;
[0019] Screen the multiple text classification models, and select the models with qualified classification performance to form a committee;
[0020] Use each of the qualified models in the committee to vote on the classification label of each sample in the text set to be tested, and obtain the distribution of the classification vote counts of the predicted labels of the samples.
[0021] Furthermore,
[0022] Determine the voting entropy of the samples in the text set to be tested through the following formula,
[0023]
[0024] In formula (1): k represents the number of the qualified model; i represents the number of the sample in the text set to be tested; H BAG (x i ) represents the voting entropy of the sample x i in the text set to be tested; K is the total number of the qualified models in the committee; y i represents the label category of the classification vote of the predicted label of the sample x i in the text set to be tested by the committee; l represents the label category number of the classification vote of the predicted label of the sample x i in the text set to be tested by the committee; C(y i = l|x i ) represents the set of the number of cases where the label category of the classification vote of the predicted label of the sample x i in the text set to be tested by the committee is y i , and the label category number of the classification vote is l; N i represents the committee's treatment of the sample x in the text set to be testedi The total number of label categories for the classification vote of the predicted labels; w k is the voting weight of the k-th compliant model of the committee.
[0025] Furthermore,
[0026] Calculate w according to the F1 score of the k-th compliant model on the text set to be tested k :
[0027]
[0028] In formula (2): F k represents the F1 score of the k-th compliant model of the committee on the text set to be tested; F min represents the minimum value of the F1 scores of all compliant models of the committee on the text set to be tested; F max represents the maximum value of the F1 scores of all compliant models of the committee on the text set to be tested.
[0029] Furthermore,
[0030] The query function is
[0031]
[0032] In formula (3): X EQB represents the query function value; U represents the text set to be tested; H BAG (x i ) represents the voting entropy of the sample x i in the text set to be tested.
[0033] Furthermore,
[0034] Correcting the labels of samples with high consistency and easily confused samples includes:
[0035] Using each compliant model in the committee to separately conduct a classification vote on the predicted labels of each easily confused sample to obtain the distribution data of the classification vote votes of the predicted labels of each easily confused sample;
[0036] Set the first threshold T1 for the committee vote,
[0037] Select the predicted labels in the predicted labels of the easily confused samples whose classification vote votes are greater than T1. If the classification vote votes of the predicted labels of multiple easily confused samples are greater than T1, set the predicted labels of the easily confused samples as easily confused labels, and add the corresponding new easily confused samples to the easily confused sample set; if only the classification vote votes of the predicted labels of one easily confused sample are greater than T1, then correct the label of the one easily confused sample to the predicted label whose classification vote votes are greater than T1, and add it to the label set to be corrected.
[0038] As a preferred means,
[0039] The method further includes:
[0040] Adding the misclassified samples with corrected labels to the training set.
[0041] As a further preferred means,
[0042] Adding the misclassified samples with corrected labels to the training set includes reducing the scale of the misclassified sample set; specifically, reducing the scale of the misclassified sample set is as follows:
[0043] Calculating the query function values of the samples in the misclassified sample set;
[0044] Setting a second threshold T2 of the query function, and sorting the query function values of the samples in the misclassified sample set from low to high;
[0045] If the query function value of a sample in the misclassified sample set ≤ T2, removing the sample from the misclassified sample set.
[0046] As a further preferred means,
[0047] Adding the misclassified samples to the training set to form a new training set;
[0048] Selecting multiple subsets from the new training set, and obtaining multiple new text classification models through multiple rounds of training based on multiple known text classification models respectively;
[0049] Screening the multiple new text classification models, and selecting the new text classification models with qualified classification performance to form a new committee;
[0050] Predicting labels and conducting classification voting on the samples in the text set to be tested based on the new committee.
[0051] As a further preferred means,
[0052] Expanding the number of misclassified samples in the new training set:
[0053] For each misclassified sample, performing similarity matching on the text set to be tested, and adding the sample texts similar to the misclassified sample to the new training set.
[0054] As a further preferred means,
[0055] The similarity matching adopts the Simhash-Hamming similarity matching algorithm:
[0056] First, use the Simhash algorithm for preliminary screening to form a small-scale candidate set, and then perform precise similarity matching in the candidate set.
[0057] As a further preferred means,
[0058] The Simhash-Hamming similarity matching algorithm specifically includes:
[0059] For the sample y j belonging to the training set, generate the Simhash representation of the sample y j
[0060] Set the set of Hamming distances HD = {0, 1, 3, 5}, set the first-step candidate set as S1, and limit its data scale to Size1; set the second-step candidate set as S2, and limit the data scale to Size2;
[0061] For the confusing sample x g belonging to the confusing sample set, generate the Simhash representation of the confusing sample x g
[0062] For the Hamming distance hd belonging to HD, add the index of the confusing sample x with a Hamming distance less than hd g to the first-step candidate set S1;
[0063] Calculate the size of the first-step candidate set S1. If it is less than Size1, generate the Simhash representation for other samples in the confusing sample set If the first-step candidate set S1 is greater than Size1, then traverse the first-step candidate set S1:
[0064] For the confusing sample x g whose index number idx belongs to the first-step candidate set S1,
[0065] find the corresponding text y of idx idx , and perform word segmentation on the confusing sample x g and y idx respectively to form word bags;
[0066] For the word bags of the confusing sample x g and the text y idx , calculate {Jaccard distance, Manhattan distance};
[0067] Calculate the sentence vectors of the confusing sample x g and the text y idx , and represent the weights of the word vectors weighted average as the TF-IDF values of the words;
[0068] Calculate the {Euclidean distance, cosine distance} of the sentence vectors of the confusing sample x g and the text y idx ;
[0069] Perform a weighted sum of the Euclidean distance and cosine distance of the sentence vectors to obtain a comprehensive similarity; for each text y idx Sort the comprehensive similarities of the texts and select the top Size2 texts to put into S2.
[0070] As a further preferred measure,
[0071] If the scale of S1 is not large, use the weighted average of word vectors to represent the sentence vectors of the text, and calculate the text similarity through the distance between the vectors;
[0072] If S1 is large, then use the Jaccard similarity coefficient to calculate. Compare the norm of the intersection of the word bags of two texts with the norm of the union. Denote the word bags of the two texts as q1 and q2 respectively, then the Jaccard similarity coefficient J(q1, q2) is:
[0073]
[0074] As a further preferred measure,
[0075] Update the training set with the similar samples of the said easily confused samples;
[0076] Redivide the updated training set and enter a new round of text classification model training;
[0077] Select the new round of text classification models that meet the standards and add them to the new round of committee;
[0078] Enter the samples in the new round of text to be tested for predicted labels and classification voting.
[0079] The present invention also provides an electronic device, including:
[0080] A memory, a processor, and a computer program stored on the memory and running on the processor;
[0081] When the processor executes the computer program, it implements any of the foregoing methods.
[0082] The present invention also provides a computer storage medium, in which computer executable instructions are stored, and when the computer executable instructions are executed, any of the foregoing methods is implemented.
[0083] In the technical solution adopted by the present invention, the models in the committee are compatible with the models trained by all text classification methods. The system can select samples with high uncertainty based on the prediction voting of the committee model, streamline the sample size required for training the model, reduce the cost of manual labeling, improve the speed of training the model, improve the quality of the training set through the improved active learning algorithm, improve the tolerance of the classification algorithm to data quality, and ensure the classification performance of the model while ensuring the time efficiency of the calculation.
[0084] The beneficial effects of the present invention at least include:
[0085] The method of the present invention is used to correct the sample labels in the text set to be tested. It is a model-independent method and can be compatible with any model trained based on machine learning and deep learning algorithms. This method can screen out samples with high uncertainty through the designed voting entropy, and use the distribution of the classification voting votes of the sample prediction labels of the voting committee to correct the labels, reducing the cost of manual labeling. These samples have a large amount of information, and the sample size required for training the model can be streamlined during model training, ensuring the time efficiency of the calculation.
[0086] Other features and advantages of the present invention will be described in the subsequent description, or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written description, claims, and drawings.
[0087] The technical solution of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings
[0088] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation to the present invention. In the drawings:
[0089] Figure 1 It is a schematic diagram of the logical framework of the active learning method in the embodiment of the present invention;
[0090] Figure 2 It is a schematic diagram of the structure of an electronic device based on the active learning method in the present invention. Detailed Embodiments
[0091] In order to more clearly illustrate the technical method of the embodiments of the present invention and be able to fully convey the scope of the present invention to those skilled in the art, the drawings used in the embodiments will be briefly introduced below. Some embodiments of the present invention are shown in the drawings. However, those skilled in the art should be able to apply the present invention to other similar scenarios and should not be limited by the embodiments.
[0092] An active learning method is provided in the embodiments of the present invention, and its overall logical framework is asFigure 1 As shown in the figure, the active learning method includes the following steps:
[0093] Step S1: Determine the distribution data of the classification vote counts of the predicted labels of the samples.
[0094] S11 Select multiple subsets from the training set, and respectively obtain multiple text classification models through multiple rounds of training based on different known text classification methods such as TextCNN, TextRNN, and BiLSTM;
[0095] The training set is selected based on the publicly available data with a large amount of data, such as the text data set retrieved by a portal website for fields. The training set contains a sufficient number of labeled sample data texts to meet the needs of training multiple text classification models.
[0096] S12 Screen the multiple text classification models, select the qualified models with qualified prediction performance, and form a committee;
[0097] S13 Use each qualified model in the committee to respectively conduct a classification vote on the predicted labels of each sample in the text set to be tested, and obtain the distribution data of the classification vote counts of the predicted labels of each sample.
[0098] The text set to be tested is composed of text data that needs to provide retrieval services.
[0099] Step S2: Determine the samples with high consistency.
[0100] According to the distribution data of the classification vote counts of the predicted labels of each sample, determine the voting entropy of each sample, and determine the samples with a voting entropy lower than the set threshold α as the samples with high consistency.
[0101] Specifically, the voting entropy of the samples in the text set to be tested is determined by the following formula
[0102]
[0103] In formula (1): k represents the qualified model number; i represents the sample number in the text set to be tested; H BAG (x i ) represents the voting entropy of the sample x i in the text set to be tested; K is the total number of qualified models in the committee; y i represents the label category of the classification vote of the predicted label of the sample x i in the text set to be tested by the committee; l represents the label category number of the classification vote of the predicted label of the sample x i in the text set to be tested by the committee; C(y i = l|x i ) represents the committee's classification vote on the sample x iThe label category of the classification vote for the predicted label is y i , and the number of label categories of classification voting is l; N i It represents the sample x in the test text set of the committee i The total number of label categories of the classification vote for the predicted label; w k is the voting weight of the kth qualified model of the committee.
[0104] w is calculated mainly based on the F1 score of the kth qualified model on the text set to be tested. k :
[0105]
[0106] F1 score is an evaluation indicator used in statistics to measure the recall and precision of classification models. The F1 score can be regarded as a harmonic average of the model's precision and recall, with a maximum value of 1 and a minimum value of 0.
[0107] In formula (2): F k represents the F1 score of the committee's kth qualified model on the test text set; F min represents the minimum F1 score of all qualified models of the committee on the test text set; F max It represents the maximum F1 score of all qualified models of the committee on the test text set.
[0108] According to the voting entropy h BAG (x i ) The samples with high consistency selected from the test text set have a very concentrated distribution of the classification votes of the committee for their predicted labels. If the distribution of the classification votes of the predicted labels is more concentrated, the voting entropy h BAG (x i ) is smaller, the higher the certainty of the sample with high consistency.
[0109] Step S3: Determine easily confused samples.
[0110] The voting entropy of each sample in the test text set is input into the query function to determine the query function value of each sample in the test text set. Samples with query function values greater than the set threshold β are determined as easily confused samples.
[0111] The characteristics of easily confused samples are large category divergence, high uncertainty, easy classification errors, and large amount of information. The key to screening easily confused samples is to meet the distribution of classification votes for the predicted labels obtained by the aforementioned committee.
[0112] Specifically, the query function is
[0113]
[0114] In formula (3): X EQB represents the query function value; U represents the text set to be measured; h BAG (x i ) represents the voting entropy of sample x i in the text set to be measured.
[0115] For the easily confused samples selected by the query function X EQB the distribution of the classified voting votes of the committee on their predicted labels is relatively scattered. If the divergence of the distribution of the classified voting votes of the predicted labels is greater, the query function value is greater, and the information volume of the easily confused samples is greater.
[0116] Step S4: Correct the labels of the samples with high consistency and the easily confused samples.
[0117] Specifically, it includes:
[0118] S41 Use each qualified model in the committee to conduct classified voting on the predicted labels of each easily confused sample respectively, and obtain the distribution data of the classified voting votes of the predicted labels of each easily confused sample.
[0119] S42 Set the first threshold T1 for the committee voting
[0120] Select the label with the classified voting vote of the predicted label greater than T1. If the classified voting votes of multiple predicted labels are greater than T1, set the sample label as the easily confused label, and add the corresponding obtained easily confused samples to the easily confused sample set;
[0121] If only the classified voting vote of one predicted label is greater than T1, the label of the easily confused sample is corrected to this predicted label and added to the label set to be corrected.
[0122] Extend the label correction method to other unlabeled sample sets, and automatic sample annotation can also be carried out.
[0123] As a further optimization scheme of the present invention, the candidate scale of the samples to be measured can be further reduced by the following method:
[0124] Step S5: Reduce the scale of the easily confused sample set. Specifically, it includes:
[0125] S51 Calculate the query function values of the samples in the easily confused sample set.
[0126] S52 Set the second threshold T2 of the query function, and sort the query function values of the samples in the easily confused sample set from low to high.
[0127] S53 If the sample query function value ≤ T2, remove this sample from the easily confused sample set.
[0128] As a further optimization solution of the present invention, the confusing samples after label correction are added to the training set, and data augmentation is performed on the training set to improve the generalization performance of the qualified model, thereby ensuring the reliability of the committee voting.
[0129] Step S6: Add the confusing samples after label correction to the training set.
[0130] After the confusing samples are added to the training set, in the same way as above, based on TextCNN, TextRNN, and BiLSTM, the next round of training of the text classification model is carried out, and new qualified models with qualified prediction performance are selected to form a new committee, and the samples in the text set to be tested are predicted and classified by the new committee.
[0131] Through two rounds of voting mechanisms, this algorithm can reduce the candidate scale of potential samples and improve the reliability of sample label correction. Using the committee to conduct classification voting on prediction labels can reduce the prediction deviation caused by accidental factors. Samples are selected through the query function, and the distribution of the classification voting votes of the modular committee is used for label annotation, and then supplemented to the training set.
[0132] As a further optimization solution, the method of the present invention further includes
[0133] Step S7: Discover more confusing samples.
[0134] Since active learning requires re-segmenting the training set, re-training the text classification model, and updating all qualified models of the committee for re-training in each iteration, the sample size is large and the time cost is high. Therefore, it is expected to discover more confusing samples when selecting training set samples, and this type of sample contributes greatly to the parameter learning of model training.
[0135] Specifically, to discover more confusing samples, retrieve in the training set and match the texts similar to the above-mentioned confusing samples to expand the number of confusing samples, including the following steps:
[0136] For each confusing sample, perform similarity matching on the training set to match similar sample texts. If directly using word vectors and combining Euclidean distance, cosine distance, etc. to calculate the similarity of samples, the calculation efficiency is very low. To improve the matching efficiency, the present invention designs a Simhash-Hamming similarity matching algorithm to improve the efficiency of similar text matching in a step-by-step manner.
[0137] The specific Simhash-Hamming similarity matching algorithm is determined as follows:
[0138] First, use the Simhash algorithm for preliminary screening to form a small-scale candidate set, and then perform precise similarity matching in the candidate set. Specifically:
[0139] For sample y j belonging to the training set, generate the Simhash representation of sample y j
[0140] Set the set of Hamming distances HD = {0, 1, 3, 5}, set the first-step candidate set as S1 with its data scale being Size1; set the second-step candidate set as S2 with the restricted data scale being Size2;
[0141] For the confusing sample x g belonging to the confusing sample set, generate the Simhash representation of sample x g
[0142] For the Hamming distance hd belonging to HD, add the index of the sample x with Hamming distance less than hd g to the first-step candidate set S1;
[0143] Calculate the size of S1. If it is less than Size1, then generate the Simhash representation for other samples in the confusing sample set If S1 is greater than Size1,
[0144] Traverse the candidate set S1. For the index number idx belonging to S1:
[0145] Find the corresponding text y of idx idx , and perform word segmentation on x g and y idx respectively to form their own word bags.
[0146] For the word bags of x g and y idx , calculate {Jaccard distance, Manhattan distance};
[0147] Calculate the sentence vectors of x g and y idx , represented by weighted average of word vectors with the weight being the TF-IDF value of the word;
[0148] Calculate the {Euclidean distance, cosine distance} of the sentence vectors of x g and y idx ;
[0149] Sum up the above distances with weights to obtain the comprehensive similarity; sort the comprehensive similarities of each y idx and select the top Size2 texts to put into S2.
[0150] Among them, the Simhash representation of the text is a fixed-length 0-1 sequence obtained based on the HashMap algorithm. It can use the term frequency–inverse document frequency (TF-IDF) of the features as weights. During similarity matching calculation, by fuzzily matching the Hamming distance of the text, the matching efficiency can be greatly improved.
[0151] The aforementioned algorithm preliminarily screens the set S1 of similar texts that meet the Hamming distance in a large-scale dataset, and then precisely matches in S1 to obtain the set S2 of similar texts. If the scale of S1 is not large, the weighted average of word vectors can be used to represent the sentence vector of the text, and the text similarity can be calculated through the distance between vectors; if S1 is large, the Jaccard similarity coefficient can be used for calculation, which is the ratio of the norm of the intersection of the bag-of-words of two texts to the norm of the union. Denote the bag-of-words of the two texts as q1 and q2 respectively, then the Jaccard similarity coefficient J(q1, q2) is:
[0152]
[0153] The comprehensive similarity is an idea for calculating text similarity. The similarity calculation methods include but are not limited to the Jaccard distance and the Manhattan distance. The sample degree of the text can also be calculated using methods such as SimBERT and SimCSE; similarly, the representation methods of the samples include but are not limited to the bag-of-words representation after sample word segmentation, and other representation methods can be used instead. For example, the sentence vector of the sample can be represented using BERT.
[0154] In this embodiment, it is possible to select confusing samples, streamline the sample size required for training the model, reduce the cost of manual labeling, improve the speed of training the model, improve the quality of the training set through the improved active learning algorithm, improve the tolerance of the classification algorithm to data quality, and ensure the classification performance of the model while ensuring the time efficiency of the calculation.
[0155] The present invention preferably improves the reliability of sample label correction based on two votes, and this method is model-independent and can be run using different text classification models. The present invention is compatible with machine learning, deep learning, and models such as Transformer and BERT.
[0156] A similar text matching method provided by the present invention can quickly match similar texts in a massive sample set for confusing samples, and discover more potential confusing samples. These samples can be used for data augmentation of the training set to improve the prediction ability of the committee model.
[0157] In the present invention, by selecting mislabeled samples based on the classification voting distribution of predicted labels of the committee and the query function values of samples, the quality of the training set samples is improved, the quality of model training is enhanced, the tolerance of the classification algorithm to data quality is increased, the classification performance of the model is ensured, and the time efficiency of the calculation is guaranteed.
[0158] The active learning method of the present invention can discover similar samples of confusing samples in a large amount of training data with a high accuracy rate, add them to the training set, and effectively reduce the cost of manual labeling. Update the training set with the similar samples of the confusing samples; re-partition the updated training set and enter a new round of text classification model training; select a new round of text classification models that meet the standards and add them to a new round of the committee; enter a new round of prediction labels and classification voting for the samples in the text set to be tested.
[0159] Those skilled in the art can change the above order without departing from the protection scope of the present invention.
[0160] Based on the same inventive concept, an embodiment of the present invention provides an electronic device, the structure of which is as Figure 2 shown and includes: a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, it implements the foregoing active learning method, or text model training method, or text set confusing sample recognition method, or text label automatic correction method, or text similarity matching method.
[0161] Based on the same inventive concept, an embodiment of the present invention provides a computer storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed, they implement the foregoing active learning method, or text model training method, or text set confusing sample recognition method, or text set label automatic correction method, or text similarity matching method.
[0162] Any modifications, supplements, equivalent replacements, etc. made within the principle scope of the present invention shall still fall within the patent coverage scope of the present invention.
Claims
1. An active learning method, characterized in that the method comprises the following steps: determining distribution data of classification vote counts of predicted labels of samples in a text set to be measured; determining samples with high consistency according to the distribution data and the voting entropy of samples in the text set to be measured; inputting the voting entropy of each sample in the text set to be measured into a query function respectively to determine the query function value of each sample in the text set to be measured, and obtaining samples that are prone to confusion; correcting the labels of samples with high consistency and samples that are prone to confusion.
2. The method according to claim 1, characterized in that determining distribution data of classification vote counts of predicted labels of samples in a text set to be measured includes: selecting multiple subsets from a training set, and respectively obtaining multiple text classification models through multiple rounds of training based on multiple known text classification models; screening the multiple text classification models, and selecting models with qualified classification performance to form a committee; using each qualified model in the committee to vote on the classification label of each sample in the text set to be measured respectively, and obtaining the distribution of classification vote counts of the predicted labels of the samples.
3. The method according to claim 2, characterized in that the voting entropy of samples in the text set to be measured is determined by the following formula In formula (1): k represents the model number that meets the standard; i represents the number of samples in the text set to be tested; H BAG (x i ) represents the voting entropy of sample x i in the text set to be tested; K is the total number of models that meet the standard in the committee; y i represents the label category of the classification vote on the predicted label of sample x i in the text set to be tested by the committee; l represents the label category number of the classification vote on the predicted label of sample x i in the text set to be tested by the committee; C(y i = l|x i ) represents the number set of the label category of the classification vote on the predicted label of sample x i in the text set to be tested by the committee, where the label category of the classification vote is y i , and the label category number of the classification vote is l; N i represents the total number of label categories of the classification vote on the predicted label of sample x i in the text set to be tested by the committee; w k is the voting weight of the k-th model that meets the standard in the committee.
4. The method according to claim 3, characterized in that Calculate w according to the F1 score of the k-th compliant model on the text set to be tested k : In formula (2): F k represents the F1 score of the k-th compliant model of the committee on the text set to be tested; F min represents the minimum value of the F1 scores of all compliant models of the committee on the text set to be tested; F max Represents the maximum F1 score of all compliant models of the committee on the text set to be tested.
5. The method according to claim 3, characterized in that the query function is In formula (3): X EQB represents the query function value; U represents the text set to be measured; H BAG (x i ) represents the voting entropy of the sample x i in the text set to be measured.
6. The method according to claim 3, characterized in that correcting the labels of samples with high consistency and samples that are prone to confusion includes: using each qualified model in the committee to vote on the predicted label of each sample that is prone to confusion respectively, and obtaining the distribution data of classification vote counts of the predicted labels of each sample that is prone to confusion; setting a first threshold T1 for committee voting, selecting the predicted label with a classification vote count greater than T1 among the predicted labels of samples that are prone to confusion. If the classification vote counts of the predicted labels of multiple samples that are prone to confusion are greater than T1, setting the predicted label of the sample that is prone to confusion as a label that is prone to confusion, and adding the corresponding new sample that is prone to confusion to the set of samples that are prone to confusion; if only the classification vote count of the predicted label of one sample that is prone to confusion is greater than T1, then correcting the label of the one sample that is prone to confusion to the predicted label with a classification vote count greater than T1, and adding it to the set of labels to be corrected.
7. The method according to claim 6, characterized in that the method further comprises: adding the sample that is prone to confusion with the corrected label to the training set.
8. The method according to claim 7, characterized in that adding the sample that is prone to confusion with the corrected label to the training set includes reducing the scale of the set of samples that are prone to confusion; specifically reducing the scale of the set of samples that are prone to confusion is: calculating the query function value of samples in the set of samples that are prone to confusion; setting a second threshold T2 for the query function, and sorting the query function values of samples in the set of samples that are prone to confusion from low to high; if the query function value of a sample in the set of samples that are prone to confusion ≤ T2, removing the sample in the set of samples that are prone to confusion from the set of samples that are prone to confusion.
9. The method according to claim 7 or 8, characterized in that the sample that is prone to confusion is added to the training set to form a new training set; Select multiple subsets from the new training set, and respectively obtain multiple new text classification models through multiple rounds of training based on multiple known text classification models; Screen the multiple new text classification models, and select the new text classification models with qualified classification performance to form a new committee; Based on the new committee, predict labels and conduct classification voting on the samples in the text set to be tested.
10. The method according to claim 9, wherein: Increase the number of easily confused samples in the new training set: For each easily confused sample, perform similarity matching on the text set to be tested, and match the sample texts similar to each easily confused sample and add them to the new training set.
11. The method according to claim 10, wherein: The similarity matching uses the Simhash-Hamming similarity matching algorithm: First, use the Simhash algorithm for preliminary screening to form a small-scale candidate set, and then perform accurate similarity matching in the candidate set.
12. The method according to claim 11, wherein: The Simhash-Hamming similarity matching algorithm specifically includes: For sample y j belonging to the training set, generate the j Simhash representation of sample y Set the set of Hamming distances HD = {0, 1, 3, 5}, set the first-step candidate set as S1, and limit its data scale to Size1; set the second-step candidate set as S2, and limit the data scale to Size2; For the confusing sample x g belonging to the confusing sample set, generate the Simhash representation of the confusing sample x g For a Hamming distance hd belonging to HD, add the indices of the confusing samples x with Hamming distance less than hd g to the first-step candidate set S1; Calculate the size of the candidate set S1 in the first step. If it is less than Size1, generate Simhash representations for other samples in the confusing sample set. If the candidate set S1 in the first step is greater than Size1, traverse the candidate set S1 in the first step: For the confusing sample x g the index idx belongs to the first-step candidate set S1 Find the text y corresponding to idx idx , and for the confusing samples x g and y idx perform word segmentation respectively to form their own word bags; For the confusing samples x g and the text y idx calculate the bag of words of {Jaccard distance, Manhattan distance}; Calculate the confusing sample x g and the text y idx of the sentence vectors, where the weighted average of the word vectors is used to represent the weight as the TF-IDF value of the word; Calculate the confusing sample x g and the text y idx of the Euclidean distance and cosine distance of the sentence vectors; Weighted sum of the Euclidean distance and cosine distance of sentence vectors is calculated to obtain the comprehensive similarity; for each text y idx Sort the comprehensive similarities, and select the top Size2 texts and put them into S2.
13. The method according to claim 12, wherein: If the scale of S1 is not large, use the weighted average of word vectors to represent the sentence vector of the text, and calculate the text similarity through the distance between vectors; If S1 is large, then use the Jaccard similarity coefficient for calculation, and compare the norm of the intersection of the word bags of the two texts with the norm of the union. Denote the word bags of the two texts as q1 and q2 respectively, then the Jaccard similarity coefficient J(q1, q2) is:
14. The method according to any one of claims 10 to 13, wherein: Use the similar samples of the easily confused samples to update the training set; Redivide the updated training set and enter a new round of text classification model training; Select the new text classification models with qualified effects in the new round and add them to the new round of committee; Enter a new round of predicting labels and conducting classification voting on the samples in the text set to be tested.
15. An electronic device, wherein: Comprising: A memory, a processor, and a computer program stored on the memory and running on the processor; When the processor executes the computer program, the method according to any one of claims 1 to 14 is implemented.
16. A computer storage medium, wherein: The computer storage medium stores computer-executable instructions, and when the computer-executable instructions are executed, the method according to any one of claims 1 to 14 is implemented.