Sample processing method, device and computer readable storage medium
By conducting classification prediction model training on unlabeled target samples and generating pre-labeled training samples, the problem of high sample annotation cost in the existing technology is solved and more efficient sample annotation is achieved.
Patent Information
- Application Number
- CN202111348688.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-11-15
AI Technical Summary
In the data classification, the sampling method is based on diversity and uncertainty, resulting in high sample annotation cost.
By determining the unlabeled target samples, input them into the classification prediction model to obtain probability distribution data, calculate the stability data, obtain pre-labeled training samples, and use these samples to train the classification prediction model until the model meets the preset stop training conditions.
The obtained pre-labeling training samples are more stable and more targeted, effectively reducing the sample labeling cost.
Smart Images

Figure CN114091595B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to, but are not limited to, the field of data processing technology, and in particular, to a sample processing method, device, and computer-readable storage medium. Background Art
[0002] In today's information explosion society, the amount of unlabeled data is usually very large, and the acquisition of labeled data is also very difficult, time-consuming and costly. Active learning methods can effectively select unlabeled data for annotation and training to obtain a model with good performance. In real life, data classification is also widely used, and data classification also requires a large amount of training data to obtain good classification results.
[0003] In data classification of related technologies, labeled samples are usually used to train an initial classification prediction model, and active learning methods are used to perform edge sampling on unlabeled samples to further manually label the sampled samples, and then the manually labeled samples are used to train the above classification prediction model to obtain a classification prediction model that meets the expectations. However, since the above sampling method is usually based on diversity and uncertainty, the cost of labeling samples is relatively high. Summary of the invention
[0004] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.
[0005] Embodiments of the present invention provide a sample processing method, device, and computer-readable storage medium, which can effectively reduce the labeling cost of samples.
[0006] In a first aspect, an embodiment of the present invention provides a sample processing method, comprising:
[0007] Determine unlabeled target samples;
[0008] Inputting the unlabeled target sample into a classification prediction model to obtain probability distribution data of classification prediction;
[0009] According to the probability distribution data, stability data is calculated;
[0010] Obtaining pre-labeled training samples according to the stability data;
[0011] The classification prediction model is trained using the pre-labeled training samples and a preset training set until the classification prediction model meets a preset stop training condition.
[0012] In a second aspect, an embodiment of the present invention further provides a sample processing device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the sample processing method as described in the first aspect above is implemented.
[0013] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the sample processing method as described in the first aspect above.
[0014] The embodiment of the present invention includes: determining an unlabeled target sample, inputting the unlabeled target sample into a classification prediction model, obtaining probability distribution data of classification prediction, then calculating stability data based on the probability distribution data, and then obtaining a pre-labeled training sample based on the stability data, and training the classification prediction model using the pre-labeled training sample and a preset training set until the classification prediction model meets a preset stop training condition. Compared with the related art, the pre-labeled training sample obtained by the embodiment of the present invention is more stable and more targeted, and can effectively reduce the labeling cost of the sample.
[0015] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings are used to provide a further understanding of the technical solution of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the technical solution of the present invention and do not constitute a limitation on the technical solution of the present invention.
[0017] Figure 1 is a flow chart of a sample processing method provided by an embodiment of the present invention;
[0018] Figure 2 is a schematic diagram of a process for determining an unlabeled target sample provided by an embodiment of the present invention;
[0019] Figure 3 is a schematic diagram of a flow chart of probability distribution data provided by an embodiment of the present invention;
[0020] Figure 4 is a schematic diagram of a stability data flow chart provided by an embodiment of the present invention;
[0021] Figure 5 is a schematic diagram of a flow chart of determining an unlabeled target sample provided by another embodiment of the present invention;
[0022] Figure 6 is a schematic diagram of a flow chart of probability distribution data provided by another embodiment of the present invention;
[0023] Figure 7 is a schematic diagram of a stability data flow chart provided by another embodiment of the present invention;
[0024] Figure 8 is a schematic diagram of a process of pre-labeling training samples provided by an embodiment of the present invention;
[0025] Fig. 9 It is a schematic diagram of a process for training a classification prediction model provided by an embodiment of the present invention;
[0026] Fig.10 It is a schematic diagram of a process for determining a classification prediction model provided by an embodiment of the present invention;
[0027] Fig.11 It is a schematic diagram of a flow chart of labeling a target sample to be labeled provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0029] It should be noted that, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0030] In today's information explosion society, the amount of unlabeled data is usually very large, but obtaining labeled data is also very difficult, time-consuming and costly. The active learning method used aims to effectively select unlabeled data for labeling and training, thereby reducing the labeling cost and obtaining a model with good performance. In real life, data classification is also widely used, and data classification also requires a large amount of training data to obtain good classification results.
[0031] In data classification of related technologies, a small number of labeled samples are usually used to train an initial classification prediction model, and active learning methods are used to select samples from unlabeled samples. Human experts are then asked to label the selected / sampled samples, and then the labeled samples are added to the original labeled training set to retrain the above classification prediction model; then the active learning method is used to re-screen samples, and this process is repeated until a classification prediction model that meets expectations is obtained. However, in the above scheme, since the active learning method usually selects samples based on diversity and uncertainty, it does not consider the stability of the sampled samples themselves for model training, which leads to a high labeling cost for the samples.
[0032] Based on this, the embodiments of the present invention provide a sample processing method, device and computer-readable storage medium, which can effectively reduce the sample annotation cost.
[0033] It is understandable that the embodiments of the present invention specifically relate to data classification, such as text classification, and text classification includes but is not limited to application scenarios such as news classification, sentiment analysis, and text review.
[0034] The embodiments of the present invention are further described below in conjunction with the accompanying drawings.
[0035] The first embodiment of the present invention specifically provides a sample processing method, such as Figure 1 As shown, Figure 1 1 is a flow chart of a sample processing method provided by an embodiment of the present invention. The sample processing method of the embodiment of the present invention includes but is not limited to the following steps:
[0036] Step S100, determining unlabeled target samples;
[0037] Step S200, inputting the unlabeled target sample into the classification prediction model to obtain probability distribution data of classification prediction;
[0038] Step S300, calculating stability data according to the probability distribution data;
[0039] Step S400, obtaining pre-labeled training samples according to stability data;
[0040] Step S500, training the classification prediction model using pre-labeled training samples and a preset training set;
[0041] Steps S100 to S500 are repeatedly executed until the classification prediction model meets the preset stop training condition.
[0042] It is understandable that, before determining the unlabeled target samples in step S100, the embodiment of the present invention also acquires original sample data, and then performs initialization processing on the original sample data to obtain initialization data, and then divides the initialization data into labeled samples and unlabeled samples.
[0043] It is understandable that the preset training set comes from the above-mentioned labeled samples.
[0044] In some embodiments, the labeled samples may be divided into a training set and a test set.
[0045] For the above initialization processing, a variety of initialization processing methods can be set accordingly. For example, a random sampling method is used to obtain a part of the original target sample data from all the original sample data, for example, the number of original target sample data obtained corresponds to 10% of the original sample data. In other embodiments, the number ratio corresponding to the randomly sampled original target sample data can also be set between 10%-30%, or adaptive adjustment can be made according to the number of original sample data, and the embodiment of the present invention does not specifically limit this. Afterwards, the original target sample data is annotated by experts to generate labeled samples, that is, the initialization data includes labeled samples and unlabeled samples, and then the labeled samples are divided into training sets and test sets. Specifically, assuming that there are 1,000 original sample data in total, a random sampling method is used to obtain 10% of the original sample data, that is, 100 original sample data are obtained, and the 100 original sample data are used as the original target sample data.
[0046] For another example, a clustering method is used to classify the original sample data to obtain classified sample data, and then a portion of the original target sample data is obtained from the classified sample data according to a preset ratio, and then the original target sample data is annotated by experts to generate annotated samples, so as to facilitate the division of the annotated samples into a training set and a test set. Clustering methods include but are not limited to k-means clustering method (k-means clustering algorithm), hierarchical clustering method, etc., wherein the distance metric can be word2vec word vector or edit distance, etc. Specifically, assuming that the original sample data is classified by a clustering method, three categories of classified sample data are obtained, and the number of samples corresponding to the three categories of classified sample data is 500, 300, and 200 respectively, then the original target sample data is obtained from the three categories of classified sample data at a ratio of 10%, that is, the number of samples corresponding to the three categories of original target sample data is 500*10%, 300*10%, and 200*10% respectively, which means that 50, 30, and 20 original target sample data are selected from the three categories of classified sample data respectively. It is understandable that if the number of samples calculated according to the ratio is a non-integer, the non-integer is rounded off, for example, by rounding off.
[0047] It is understandable that the embodiments of the present invention may also adopt other initialization processing methods to initialize the original sample data, and are not limited to the above embodiments, which will not be described in detail here.
[0048] It is understood that the specific application of the set of test sets can refer to Fig.10 shown. Fig.10 FIG. 5 is a flow chart of determining a classification prediction model provided by an embodiment of the present invention. That is, step S500 includes but is not limited to the following steps:
[0049] Step S510, training the classification prediction model using the pre-labeled training samples and the preset training set to obtain a candidate prediction model;
[0050] Step S520, inputting a preset test set into the candidate prediction model to obtain test data;
[0051] Step S530: When the test data meets the expected test results, it is determined that the classification prediction model meets the preset stop training condition.
[0052] It can be understood that the test set can be input into the candidate prediction model to determine whether the test data output by the candidate prediction model meets the expected test results. If it meets the expected test results, it can be determined that the current candidate prediction model, i.e., the classification prediction model, meets the preset stop training conditions.
[0053] Specifically, before determining the unlabeled target samples in step S100, the labeled samples can be divided into a training set and a test set, and the training set in the labeled samples can be used to train the classification prediction model. At this time, the test set in the labeled samples can be directly input into the classification prediction model to obtain the first test data. At this time, if the first test data meets the expected test results, it can be directly determined that the classification prediction model meets the preset stop training condition.
[0054] In some embodiments, an initial model such as XLNet or textcnn can be trained based on a training set in annotated samples to obtain a classification prediction model, such as a text classification model. It is understandable that the present invention does not specifically limit the type of classification prediction model to be trained. Since the classification prediction model is continuously iteratively updated, after each round of training of the classification prediction model is completed, the test set can be input into the above classification prediction model to obtain a second test data. When the second test data meets the expected test results, it is determined that the current classification prediction model meets the preset stop training condition.
[0055] It is understandable that data such as precision, recall, F1 value (i.e., the harmonic mean of precision and recall) can be used to characterize the expected test results. Taking precision as an example to characterize the expected test results, assuming that the second test data obtained is 82% precision, and the expected test result set is 85%, it means that the current classification prediction model does not meet the preset stop training conditions, and it is necessary to continue to use the sample processing method of the embodiment of the present invention to train the classification prediction model; or, if the expected test result is set to 80%, it means that the current second test data meets the expected test results, and then it is determined that the current classification prediction model meets the preset stop training conditions. It is understandable that the expected test results can be set according to the actual application scenario, and are not limited to the above embodiments, and will not be repeated here.
[0056] Reference Fig.11 As shown, it can be understood that after determining that the classification prediction model meets the preset stop training condition, the sample processing method of the embodiment of the present invention further includes:
[0057] Step S600, obtaining a target sample to be labeled;
[0058] Step S700 , labeling the target samples to be labeled according to the classification prediction model that meets the training stop condition.
[0059] After determining that the classification prediction model meets the preset stop training condition, the embodiment of the present invention can directly execute step S600 and step S700. The target sample to be labeled is the sample that actually needs to be labeled. The target sample to be labeled is input into the classification prediction model that meets the stop training condition (i.e., input into the classification prediction model that has completed training), and the probability distribution data to be labeled of the classification prediction can be obtained. Then, according to the probability distribution data to be labeled, the labeling attribute data corresponding to the target sample to be labeled is obtained, and according to the labeling attribute data, the labeling data corresponding to the target sample to be labeled is determined.
[0060] It can be understood that step S600 and step S700 of the embodiment of the present invention can be arranged after step S500, or can be arranged after step S530.
[0061] It should be noted that when the classification prediction model meets the preset stop training conditions, it will stop acquiring pre-labeled training samples. At this time, the classification prediction model that has completed training is used to label the target samples to be labeled, and the labeling data corresponding to the target samples to be labeled are reviewed by experts. It is understandable that the target samples to be labeled can be unlabeled samples other than the training samples to be labeled; or, the target samples to be labeled can also be other actually required samples to be labeled, which are not specifically limited here.
[0062] Reference Figure 2 It can be understood that step S100 includes but is not limited to the following steps:
[0063] Step S101, performing data perturbation processing on preset unlabeled samples to obtain perturbed samples;
[0064] Step S102: determine the disturbed samples and the unlabeled samples as unlabeled target samples.
[0065] The embodiment of the present invention performs data perturbation processing on the preset unlabeled sample A to obtain a perturbed sample. For example, unlabeled sample data can be selected from the preset unlabeled sample A (the unlabeled sample A can be a set of unlabeled sample data), and data perturbation processing is performed on each unlabeled sample data to obtain a perturbed sample corresponding to each unlabeled sample data. The perturbed sample and the unlabeled sample A are determined as unlabeled target samples.
[0066] It is understood that data perturbation processing includes but is not limited to the following methods:
[0067] 1. Synonym perturbation processing: Initialize a synonym word list, select an unlabeled sample data from the preset unlabeled sample A (unlabeled sample A can be a set of unlabeled sample data), perform word segmentation processing on the unlabeled sample data, obtain a number of unlabeled word data, randomly select an unlabeled word from the unlabeled word data, and perform synonym replacement processing on the selected unlabeled word to obtain synonym data. That is, search for synonyms in the synonym word list, and when a synonym corresponding to the unlabeled word is found, replace the unlabeled word with a synonym to obtain synonym data, that is, the synonym data can be used as one of the perturbation samples corresponding to the unlabeled sample data. When no synonym is found, randomly select another unlabeled word from the remaining unlabeled word data to perform synonym replacement processing. When all the unlabeled word data in the unlabeled sample data cannot find a synonym, give up using the synonym perturbation processing method to generate perturbation samples.
[0068] For example, the unlabeled sample data is “how to quickly learn to sing a song”. After word segmentation, several unlabeled word data, namely “how to quickly learn to sing a song”, are obtained. An unlabeled word is randomly selected from the unlabeled word data, such as “sing”. If a synonym for “sing” cannot be found in the synonym word list, then another unlabeled word is randomly selected from the remaining unlabeled word data, such as “how”. The synonyms of “how” found in the synonym word list include “how”, “how”, etc. Then a synonym is randomly selected from the synonym word list, such as “how”, to perform synonym replacement processing on the unlabeled word “how” to obtain synonym data, that is, the synonym data represents the perturbation sample corresponding to the unlabeled sample data, specifically “how to quickly learn to sing a song”.
[0069] 2. Translation disturbance processing: Select an unlabeled sample data from the preset unlabeled sample A (unlabeled sample A can be a set of unlabeled sample data), use the translation tool to perform language translation processing on the unlabeled sample data, obtain translation data, and then perform language translation processing on the translation data again to obtain one of the disturbance samples corresponding to the unlabeled sample data, wherein the language type corresponding to the disturbance sample is the same as the language type corresponding to the unlabeled sample data. That is, first translate the unlabeled sample data into other languages, and then translate it back to the source language. It can be understood that the unlabeled sample data can be text data, for example, a sentence can be used as an unlabeled sample data.
[0070] For example, if the language type corresponding to the unlabeled sample A is Chinese, the unlabeled sample data in Chinese can be translated into English translation data first, and then the translated translation data can be translated back into Chinese to obtain the perturbed sample in Chinese. It can be understood that the perturbed sample in this embodiment may be exactly the same as the unlabeled sample data before translation, so the Beamsearch method can be used to ensure that the perturbed sample is different from the unlabeled sample data before translation. The Beam search method is a public technology in the field of machine translation and will not be described in detail here. It can be understood that the unlabeled sample data can be processed by language translation with the help of a variety of different language types to generate a variety of perturbed samples. For example, the translation perturbation process can be embodied as: generating a perturbed sample through Chinese-English-Chinese; generating another perturbed sample through Chinese-Italian-Chinese. The Beam search method is also used to ensure that the generated perturbation samples are different.
[0071] 3. Pre-trained language model perturbation processing: A pre-trained language model is used to construct a perturbation sample, that is, a pre-trained language model is inputted with a preset unlabeled sample A into a preset pre-trained language model to obtain a perturbation sample. Specifically, the pre-trained language model can be trained using the MASK masking method through models such as BERT and ELECTRA. Taking the BERT model as an example, an unlabeled sample data is selected from the preset unlabeled sample A (the unlabeled sample A can be a set of unlabeled sample data), and some unlabeled word data in the unlabeled sample data is randomly set to MASK to obtain unlabeled mask data, and the unlabeled mask data is inputted into the BERT model, and the unlabeled mask data is predicted by the BERT model to output one of the perturbation samples corresponding to the unlabeled sample data. Since the result of the MASK predicted by the BERT model, that is, the perturbation sample, may be the same as the unlabeled sample data, at this time, the unlabeled word data in the unlabeled sample data can be combined and set into different unlabeled mask data and input into the BERT model. Set a maximum of 10 attempts. If no suitable perturbation sample is generated after 10 attempts, give up using this method to generate perturbation samples.
[0072] For example, the unlabeled sample data is “How to quickly learn to sing a song”, and the unlabeled sample data is segmented using BERT to obtain a number of unlabeled word data, namely “How to quickly learn to sing a song”. Some of the unlabeled word data in the unlabeled sample data are randomly set to MASK to obtain unlabeled masked data. For example, no more than 20% of the unlabeled word data are randomly set to MASK, for example, the unlabeled mask data is “How to MASK quickly learn MASK to sing a song”, and the unlabeled mask data is input into the BERT model to obtain a perturbation sample. When the perturbation sample is different from the unlabeled sample data, the perturbation sample is determined to be the corresponding perturbation sample. When the perturbation sample is the same as the unlabeled sample data, the unlabeled mask data is reacquired. It can be set to repeat up to 10 times. In other embodiments, other numbers of repetitions can also be set, which are not specifically limited here.
[0073] It is understandable that other data perturbation processing methods may be used instead of the above data perturbation processing method, which does not affect the training of the classification prediction model in the embodiment of the present invention, and is within the protection scope of the present application, and will not be described in detail here.
[0074] Afterwards, refer to Figure 3 It can be understood that step S200 includes but is not limited to the following steps:
[0075] Step S201: input the disturbed samples and the unlabeled samples into the classification prediction model to obtain the disturbed probability distribution data and the unlabeled probability distribution data.
[0076] In an embodiment of the present invention, after the disturbed sample and the unlabeled sample are determined as unlabeled target samples, the disturbed sample and the unlabeled sample A are input into a classification prediction model to obtain disturbed probability distribution data and unlabeled probability distribution data, wherein the disturbed probability distribution data is the probability distribution data corresponding to the classification prediction of the disturbed sample, and the unlabeled probability distribution data is the probability distribution data corresponding to the classification prediction of the unlabeled sample A.
[0077] In another embodiment, each unlabeled sample data can be input into the classification prediction model respectively to obtain the unlabeled probability distribution data corresponding to each unlabeled sample data; multiple perturbation samples corresponding to each unlabeled sample data can be input into the classification prediction model respectively to obtain the perturbation probability distribution data corresponding to each perturbation sample.
[0078] Reference Figure 4 It can be understood that step S300 includes but is not limited to the following steps:
[0079] Step S301, calculating stability data according to a first stability algorithm, disturbance probability distribution data and unlabeled probability distribution data.
[0080] That is, according to step S101, step S102, step S201 and step S301 of the embodiment of the present invention, the stability data corresponding to the unlabeled sample A is obtained.
[0081] Specifically, the classification prediction model is used to predict the probability distribution data of the classification prediction of the disturbed sample, i.e., the disturbed probability distribution data, and the probability distribution data of the classification prediction of the unlabeled sample A, i.e., the unlabeled probability distribution data; based on the first stability algorithm, the disturbed probability distribution data, and the unlabeled probability distribution data, the stability data corresponding to the unlabeled sample A is calculated. It can be understood that each unlabeled sample data in the unlabeled sample A corresponds to stability data.
[0082] It can be understood that, assuming that an unlabeled sample data in unlabeled sample A generates a total of N perturbation samples, the perturbation samples are specifically represented as R1, R2, ..., R N , plus the original unlabeled sample data, there are a total of N+1 samples, for example, N+1 sentences. Assuming it is an m-classification problem, m is the number of classification prediction categories output by the classification prediction model, then the unlabeled probability distribution data Shape A vector of length m, where j = 1, 2, ..., m; and the unlabeled probability distribution data In It can be understood as the probability data that the unlabeled sample data in unlabeled sample A belongs to category 1.
[0083] Record the perturbation probability distribution data The probability distribution data corresponding to the classification prediction of the perturbation sample, unlabeled probability distribution data is the probability distribution data corresponding to the classification prediction of the unlabeled sample A. It can be understood that each unlabeled sample data in the unlabeled sample A may correspond to unlabeled probability distribution data.
[0084] A first stability algorithm is defined to calculate a first stability parameter of the unlabeled sample A after data disturbance processing.
[0085] Specifically, the calculation formula of the first stability algorithm is:
[0086]
[0087] Among them, TDS is the first stability parameter, which is used to characterize the stability data, N is the number of disturbance samples, m is the number of classification prediction categories output by the classification prediction model, is the perturbation probability distribution data, is the unlabeled probability distribution data, i = 1, 2, ..., N, j = 1, 2, ..., m.
[0088] It can be understood that, the larger the TDS, ie, the first stability parameter, is, the more stable the corresponding unlabeled sample data in the unlabeled sample A is.
[0089] It should be noted that for synonym perturbation processing and pre-trained language model perturbation processing, effective perturbation samples may not be generated. However, for translation perturbation processing, perturbation samples can be effectively generated. For example, through translation perturbation processing, two different language types are selected to generate two perturbation samples respectively. Through the above-mentioned data perturbation processing method, it can be known that the number N of generated perturbation samples may take values of 2, 3, and 4. In order to solve the problem that the number of perturbation samples corresponding to different unlabeled sample data is not uniform, there is a 1 / N coefficient in the above-mentioned first stability algorithm that can be adjusted to ensure the accuracy of the data.
[0090] It is understandable that the corresponding stability data can be calculated for each unlabeled sample data in the unlabeled sample A. Figure 8 It can be understood that step S400 includes but is not limited to the following steps:
[0091] Step S410, sorting the unlabeled samples according to the stability data;
[0092] Step S420, screening the sorted unlabeled samples to obtain training samples to be labeled;
[0093] Step S430 , labeling the training samples to be labeled according to the classification prediction model to obtain pre-labeled training samples.
[0094] It is understandable that the embodiment of the present invention includes labeled samples and unlabeled samples. Since the number of unlabeled samples is usually large, if all unlabeled samples are directly labeled manually, the labeling cost will be high. Therefore, the embodiment of the present invention performs labeling processing on the unlabeled samples through the classification prediction model in the training process through steps S100 to S500, and then obtains pre-labeled training samples. It is understandable that in the labeling process, it is necessary to obtain the pre-labeled training samples after expert review and confirmation, that is, by querying the correct labeling data from the experts.
[0095] Specifically, an embodiment of the present invention performs data perturbation processing on preset unlabeled samples to obtain perturbed samples, and then inputs the perturbed samples and the unlabeled samples into a classification prediction model to obtain perturbed probability distribution data and unlabeled probability distribution data; then, according to a first stability algorithm, the perturbation probability distribution data and the unlabeled probability distribution data, the stability data corresponding to each unlabeled sample data in the unlabeled samples is calculated.
[0096] Specifically, the embodiment of the present invention selects unlabeled sample data with poor stability from the unlabeled samples, and then asks experts to label and confirm them.
[0097] The unlabeled samples are sorted by stability data, for example, the stability data is sorted in order from small to large or from large to small. It can be understood that the embodiment of the present invention needs to obtain unlabeled sample data with poor stability to realize the training of the classification prediction model. Therefore, it is necessary to select the top n unlabeled sample data with the smallest TDS, i.e., the first stability parameter (indicating the worst stability). It can be understood that the above n can be adjusted according to the actual situation. For example, it is 5% of the total data volume of the unlabeled samples. For example, if the total data volume of the unlabeled samples is 10,000, the sorted unlabeled samples are screened to obtain the training samples to be labeled, that is, the unlabeled sample data corresponding to 10,000*5% with the smallest TDS, i.e., the first stability parameter, is selected to obtain 500 data volumes of training samples to be labeled. Afterwards, the training samples to be labeled are labeled according to the classification prediction model to obtain pre-labeled training samples. Afterwards, the pre-labeled training samples can be reviewed by experts to save workload.
[0098] It is understandable that the embodiment of the present invention fully considers the stability of each unlabeled sample data in the unlabeled sample A to data disturbance, thereby making the collected pre-labeled training samples more valuable. And subsequent experts only need to confirm or modify the obtained high-value pre-labeled training samples, such as querying the correct labeled data by experts, thereby effectively reducing the labeling cost.
[0099] Reference Figure 5 It can be understood that step S100 includes but is not limited to the following steps:
[0100] Step S110, when the current training round of training the classification prediction model is greater than the preset round threshold, the preset unlabeled samples corresponding to the current training round are determined as the unlabeled target samples of the current training round, and the preset unlabeled samples corresponding to the next training round after the current training round are determined as the unlabeled target samples of the next training round.
[0101] The embodiment of the present invention calculates stability data by obtaining the last k rounds of training of the classification prediction model.
[0102] In the last k rounds, each round predicts the preset unlabeled samples corresponding to the current training round. Specifically, refer to Figure 6 It can be understood that step S200 includes but is not limited to the following steps:
[0103] Step S210, for the current training round, inputting the corresponding unlabeled target sample into the classification prediction model to obtain the current round probability distribution data of the current training round;
[0104] Step S220, for the next training round, input the corresponding unlabeled target sample into the classification prediction model to obtain the next round probability distribution data of the next training round;
[0105] Step S230, and so on, perform multiple rounds of training processing on the classification prediction model to obtain multiple current round probability distribution data and multiple next round probability distribution data.
[0106] It can be understood that the current round probability distribution data is the probability distribution data corresponding to the classification prediction of the current training round, and the next round probability distribution data is the probability distribution data corresponding to the classification prediction of the next training round.
[0107] Reference Figure 7 It can be understood that step S300 includes but is not limited to the following steps:
[0108] Step S310, calculating stability data according to the second stability algorithm, the current round probability distribution data and the next round probability distribution data.
[0109] Specifically, the embodiment of the present invention selects unlabeled sample data with poor stability from the unlabeled target samples, and then asks experts to mark and confirm. For example, steps S410 to S430 can be used to select unlabeled target samples, and the specific implementation steps and effects are the same as above, which will not be repeated here.
[0110] In the last k rounds of training the classification prediction model in an embodiment of the present invention, predictions are made on the unlabeled target samples corresponding to the current training round in each round to obtain the probability distribution data of the classification prediction of the current training round, i.e., the current round probability distribution data, and the probability distribution data of the classification prediction of the next training round, i.e., the next round probability distribution data.
[0111] It can be understood that training the classification prediction model once using all the training samples is called a training round, and the number of training rounds is calculated in this way.
[0112] For example, suppose the classification prediction model ultimately needs to be trained for 20 rounds.
[0113] Set the preset round threshold k to 10. When the current training round x is 11, it means that the current training round for training the classification prediction model is greater than the preset round threshold 10. Starting from the xth round, that is, the 11th round, the classification prediction model is used to predict the corresponding unlabeled target samples to obtain the corresponding classification prediction probability distribution data.
[0114] It is understandable that the unlabeled target samples corresponding to the current training round and the next training round may be the same.
[0115] It is understandable that the embodiment of the present invention needs to perform multiple rounds of training processing on the classification prediction model to obtain multiple current round probability distribution data and multiple next round probability distribution data. That is, for the current training round, the unlabeled target samples corresponding to the current training round are input into the classification prediction model to obtain the current round probability distribution data of the current training round; for the next training round, the unlabeled target samples corresponding to the next training round are input into the classification prediction model to obtain the next round probability distribution data of the next training round.
[0116] More specifically, by inputting the corresponding unlabeled target samples into the classification prediction model trained for 11 rounds, the current round probability distribution data corresponding to the 11th round can be obtained, which can be expressed as Q 11 Indicates that all unlabeled target samples are input into the classification prediction model for another round of training. The next training round after round 11 is round 12. Re-predict the probability distribution data of the next round corresponding to the unlabeled target samples in round 12, and use Q 12 In other words, we can finally get Q 20 .
[0117] For the last k rounds, assuming it is an m-classification problem, for Q x ,(x=r-k+1,r-k+2,...,r), that is (Q r-k+1 ,Q r-k+2 ,...,Q r ), which is recorded as the probability distribution data of the current round, that is, Then the probability distribution data of the next round is expressed as For example, for a classification prediction model trained for 11 rounds, the number of unlabeled sample data in an unlabeled target sample corresponds to It can be understood as: the probability data of the unlabeled sample data belonging to the jth category in the classification prediction model trained in the 11th round.
[0118] It can be seen that for the second stability algorithm, when the preset round threshold is 10, that is, for the last 10 rounds, in the 11th round, it is necessary to obtain Q 11j and Q 12j ; In the 12th round, you need to obtain Q 12j and Q 13j , and so on, in the 19th round, you need to obtain Q 19j and Q 20j .
[0119] A second stability algorithm is defined to calculate the second stability parameter of the unlabeled sample data in the unlabeled target sample for the last k rounds.
[0120] Specifically, the calculation formula of the second stability algorithm is:
[0121]
[0122] Among them, LKS is the second stability parameter, which is used to characterize the stability data, r is the total training rounds for the classification prediction model, x is the current training round, k is the preset round threshold, and m is the number of classification prediction categories output by the classification prediction model. is the probability distribution data of the current round, is the probability distribution data for the next round, x=r-k+1, r-k+2,...,r.
[0123] It can be understood that, the larger the LKS, i.e., the second stability parameter, is, the more stable the corresponding unlabeled sample data in the unlabeled target sample is.
[0124] It should be noted that the corresponding stability data can be calculated for each unlabeled sample data in the unlabeled target samples corresponding to each round.
[0125] The unlabeled samples (i.e., the unlabeled target samples of this embodiment) are sorted by stability data, for example, the stability data is sorted in the order from small to large or from large to small. It can be understood that the embodiment of the present invention needs to obtain unlabeled sample data with poor stability to realize the training of the classification prediction model. Therefore, it is necessary to select the top n unlabeled sample data with the smallest LKS, i.e., the second stability parameter (indicating the worst stability). It can be understood that the above n can be adjusted according to the actual situation. For example, it is 5% of the total data volume of the unlabeled target samples. For example, if the total data volume of the unlabeled target samples is 10,000, the sorted unlabeled target samples are screened to obtain the training samples to be labeled, i.e., the unlabeled sample data corresponding to 10,000*5% with the smallest LKS, i.e., the second stability parameter, is selected to obtain 500 data volumes of training samples to be labeled. Afterwards, the training samples to be labeled are labeled according to the classification prediction model to obtain pre-labeled training samples. Afterwards, the pre-labeled training samples can be reviewed by experts to save workload.
[0126] It is understandable that the embodiment of the present invention fully considers the stability of the training time, so that the collected pre-labeled training samples are more valuable and targeted. And the subsequent experts only need to confirm or modify the obtained high-value samples, which are the pre-labeled training samples in this embodiment, for example, by querying the correct labeling data through experts, thereby effectively reducing the labeling cost.
[0127] It is understandable that for each specific classification task, the total training rounds corresponding to the classification prediction model are different. The embodiment of the present invention can determine whether the third test data meets the expected test results by obtaining y consecutive rounds, such as 5 consecutive rounds of third test data. For example, when the third test data of y consecutive rounds does not exceed the test data of r selection, the r value can be determined. Assuming that the accuracy of the classification prediction model obtained by training the 22nd round is 95%, reaching a new high, and the accuracy of the subsequent 23rd to 27th rounds does not exceed 95%, then the total training round r is determined to be 22. The above-mentioned preset round threshold can be adjusted according to actual conditions. For example, it is recommended to take r / 2, and if r / 2 is not an integer, round it. For example, when the total training rounds are 22, the preset round threshold can be taken as 11.
[0128] Reference Fig. 9 It is understandable that step S500 also includes but is not limited to the following steps:
[0129] Step S501, inputting the preset pre-labeled samples into the classification prediction model to obtain probability distribution data corresponding to the pre-labeled samples;
[0130] Step S502, obtaining a first pre-labeled training sample according to the probability distribution data corresponding to the pre-labeled sample, wherein the confidence corresponding to the first pre-labeled training sample is greater than or equal to a preset confidence;
[0131] Step S503, training the classification prediction model using the first pre-labeled training sample, the training weight value corresponding to the first pre-labeled training sample and the training set, wherein the training weight value is the product of the confidence corresponding to the first pre-labeled training sample and the preset hyperparameter.
[0132] If the classification prediction model is trained using labeled samples and pre-labeled training samples, the remaining unlabeled samples that have not been manually reviewed and labeled cannot be effectively used. Therefore, the remaining unlabeled samples that have not been manually reviewed and labeled are defined as preset pre-labeled samples.
[0133] It is understandable that the classification prediction model of the embodiment of the present invention not only uses labeled samples, but also uses pre-labeled samples. The preset pre-labeled samples are input into the classification prediction model to obtain the probability distribution data corresponding to the pre-labeled samples, and the high-confidence samples are selected according to the probability distribution data corresponding to the pre-labeled samples. Specifically, according to the probability distribution data corresponding to the pre-labeled samples, the first pre-labeled training sample with a confidence greater than or equal to the preset confidence is obtained, that is, the first pre-labeled training sample is a high-confidence sample.
[0134] It should be noted that, for the pre-labeled samples, the label data corresponding to the first pre-labeled training samples further obtained are pseudo labels given by the classification prediction model. The embodiment of the present invention can further screen the pre-labeled samples according to the probability distribution data corresponding to the pre-labeled samples. Specifically, the first pre-labeled training samples are screened out, wherein the confidence corresponding to the first pre-labeled training samples is greater than or equal to the preset confidence. That is, the first pre-labeled training samples corresponding to the preset confidence are abandoned. Afterwards, the training weight value corresponding to the first pre-labeled training sample is obtained, wherein the training weight value is the product of the confidence corresponding to the first pre-labeled training sample and the preset hyperparameter, which can be expressed as w*σ. Among them, the confidence σ represents the maximum probability data in the probability distribution data corresponding to the pre-labeled sample, that is, the probability data corresponding to the pseudo label (first pre-labeled training sample); w is a preset hyperparameter, which can be set between 0-1, and the specific value can be preset. It can be understood that the first pre-labeled training sample obtained in the embodiment of the present invention does not need to be reviewed and confirmed by experts to save workload.
[0135] For example, taking the preset confidence as 0.7 and the preset hyperparameter as 0.5 as an example, it is assumed that the embodiment of the present invention is a three-category emotion classification problem, that is, the number m of classification prediction categories output by the classification prediction model is 3, which are specifically divided into positive, negative, and neutral. For the pre-labeled samples, the classification prediction model predicts that the corresponding probability distribution data is (0.1, 0.6, 0.3), that is, the probability data of positive is 0.1, the probability data of negative is 0.6, and the probability data of neutral is 0.3. It can be understood that the order of the probability distribution data corresponding to the output classification prediction categories is corresponding in the training and prediction processes of the model. Since in the probability distribution data, the largest probability data, that is, the confidence σ, is 0.6, which is lower than the preset confidence of 0.7, this pre-labeled sample is abandoned. The probability distribution data corresponding to another pre-labeled sample is (0.02, 0.9, 0.08), that is, the probability data of positive is 0.02, the probability data of negative is 0.9, and the probability data of neutral is 0.08. Then the pseudo label (that is, determining the first pre-labeled training sample) is the classification prediction category corresponding to the largest probability data 0.9, that is, negative. It can be understood that the positive, negative, and neutral data of the embodiment of the present invention are immediately labeled. Through the above method, the first pre-labeled training sample is screened out from the pre-labeled samples according to the probability distribution data corresponding to the pre-labeled samples. The training weight value corresponding to the first pre-labeled training sample is calculated to be 0.5*0.9=0.45. Finally, based on the loss function, the classification prediction model is trained using the first pre-labeled training sample, the training weight value corresponding to the first pre-labeled training sample, and the training set. Specifically, the corresponding training weight value 0.45 is multiplied before the loss function corresponding to the first pre-labeled training sample.
[0136] In order to give a lower training weight value to the first pre-labeled training sample with low confidence, the embodiment of the present invention uses w to distinguish the labeled sample from the first pre-labeled training sample corresponding to the pseudo label. For example, if w is 0.5, the training weight of the first pre-labeled training sample is always lower than 0.5, which specifically stipulates the upper limit of the pseudo label weight. Specifically, the preset confidence of this embodiment is 0.7, and the preset hyperparameter is 0.5.
[0137] The embodiment of the present invention utilizes the first pre-labeled training sample by adopting a dynamic weight method to make full use of the pre-labeled sample. When iteratively updating the classification prediction model, the labeled samples and the pre-labeled samples are used at the same time. For the pre-labeled sample, the classification prediction model is used to give a pseudo label, that is, the first pre-labeled training sample with high confidence is selected, and based on the confidence corresponding to the first pre-labeled training sample, a dynamic weight is added to the loss function corresponding to the first pre-labeled training sample.
[0138] It is understandable that the embodiments of the present invention can be applied to text classification and other related classifications, such as news classification, sentiment analysis, text review, etc., which can effectively save manual annotation. In addition, the embodiments of the present invention can also be used in combination with other existing screening methods, and the scores of each screening method are weighted and re-ranked to select pre-annotated training samples with high comprehensive value.
[0139] It is understandable that, in the related art, although there is a way to train the model using unlabeled samples, it does not distinguish between the real labels manually labeled and the pseudo labels predicted by the untrained classification prediction model. The first pre-labeled training samples obtained by the embodiment of the present invention are more valuable and targeted, and are representative. The embodiment of the present invention also uses a classification prediction model to automatically label the first pre-labeled training samples, which can effectively reduce the labeling cost.
[0140] In addition, the second aspect of the present invention further provides a sample processing device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor.
[0141] The processor and the memory may be connected via a bus or other means.
[0142] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0143] The non-transient software program and instructions required to implement the sample processing method of the first embodiment are stored in the memory. When executed by the processor, the sample processing method in the above embodiment is executed, for example, the sample processing method described above is executed. Figure 1 Steps S100 to S500 of the method, Figure 2 Steps S101 to S102, Figure 3 Step S201 of the method, Figure 4 Step S301 of the method, Figure 5 Step S110 of the method, Figure 6 Steps S210 to S230 of the method, Figure 7 Step S310 of the method, Figure 8 Steps S410 to S430 of the method, Fig. 9 Steps S501 to S503 of the method, Fig.10 Steps S510 to S530 of the method, Fig.11 Method steps S600 to S700 in.
[0144] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0145] In addition, an embodiment of the present invention further provides a computer-readable storage medium, which stores computer-executable instructions. The computer-executable instructions are executed by a processor or a controller, for example, by a processor in the above-mentioned device embodiment, so that the above-mentioned processor can execute the sample processing method in the above-mentioned embodiment, for example, execute the above-mentioned Figure 1 Steps S100 to S500 of the method, Figure 2 Steps S101 to S102, Figure 3 Step S201 of the method, Figure 4 Step S301 of the method, Figure 5 Step S110 of the method, Figure 6 Steps S210 to S230 of the method, Figure 7 Step S310 of the method, Figure 8 Steps S410 to S430 of the method, Fig. 9 Steps S501 to S503 of the method, Fig.10 Steps S510 to S530 of the method, Fig.11 Method steps S600 to S700 in.
[0146] It will be appreciated by those skilled in the art that all or some of the steps and systems in the disclosed method above may be implemented as software, firmware, hardware and appropriate combinations thereof. Some physical components or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or a non-transitory medium) and a communication medium (or a temporary medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that may be used to store desired information and may be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0147] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above-mentioned implementation mode. Technical personnel familiar with the field can also make various equivalent deformations or substitutions without violating the spirit of the present invention. These equivalent deformations or substitutions are all included in the scope defined by the claims of the present invention.
Claims
1. A sample processing method, comprising: Perform data perturbation processing on the preset unlabeled samples to obtain perturbed samples; The unlabeled samples are text data; Determine the disturbance sample and the unlabeled sample as unlabeled target samples; Inputting the disturbance sample and the unlabeled sample into a classification prediction model to obtain disturbance probability distribution data and unlabeled probability distribution data; Calculating stability data according to a first stability algorithm, the disturbance probability distribution data, and the unlabeled probability distribution data; Obtaining pre-labeled training samples according to the stability data; The classification prediction model is trained using the pre-labeled training samples and a preset training set until the classification prediction model meets a preset stop training condition.
2. The method according to claim 1, characterized in that The disturbance probability distribution data is probability distribution data corresponding to the classification prediction of the disturbance sample, and the unlabeled probability distribution data is probability distribution data corresponding to the classification prediction of the unlabeled sample.
3. The method according to claim 1, characterized in that The calculation formula of the first stability algorithm is: ; Among them, the is a first stability parameter, wherein the first stability parameter is used to characterize the stability data, is the number of the disturbance samples, is the number of classification prediction categories output by the classification prediction model, is the disturbance probability distribution data, is the unlabeled probability distribution data, , .
4. A sample processing method, comprising: When the current training round for training the classification prediction model is greater than a preset round threshold, a preset unlabeled sample corresponding to the current training round is determined as an unlabeled target sample of the current training round, and a preset unlabeled sample corresponding to a next training round after the current training round is determined as an unlabeled target sample of the next training round, wherein the unlabeled sample is text data; For the current training round, the corresponding unlabeled target sample is input into the classification prediction model to obtain the current round probability distribution data of the current training round; For the next training round, the corresponding unlabeled target samples are input into the classification prediction model to obtain the next round probability distribution data of the next training round; By analogy, the classification prediction model is trained for multiple rounds to obtain multiple current round probability distribution data and multiple next round probability distribution data; Calculating stability data according to a second stability algorithm, the current round probability distribution data, and the next round probability distribution data; Obtaining pre-labeled training samples according to the stability data; The classification prediction model is trained using the pre-labeled training samples and a preset training set until the classification prediction model meets a preset stop training condition.
5. The method according to claim 4, characterized in that The calculation formula of the second stability algorithm is: ; Among them, the is a second stability parameter, wherein the second stability parameter is used to characterize the stability data, The total number of training rounds for training the classification prediction model, is the current training round, the is the preset round threshold, is the number of classification prediction categories output by the classification prediction model, is the probability distribution data of the current round, is the probability distribution data for the next round, .
6. The method according to any one of claims 1 to 5, characterized in that: The step of obtaining a pre-labeled training sample according to the stability data includes: sorting the unlabeled samples according to the stability data; Screening the sorted unlabeled samples to obtain training samples to be labeled; The to-be-labeled training samples are labeled according to the classification prediction model to obtain pre-labeled training samples.
7. The method according to any one of claims 1 to 5, characterized in that: The method of training the classification prediction model by using the pre-labeled training samples and the preset training set also includes: Inputting the preset pre-labeled samples into the classification prediction model to obtain probability distribution data corresponding to the pre-labeled samples; Obtaining a first pre-labeled training sample according to the probability distribution data corresponding to the pre-labeled sample, wherein the confidence corresponding to the first pre-labeled training sample is greater than or equal to a preset confidence; The classification prediction model is trained using the first pre-labeled training sample, the training weight value corresponding to the first pre-labeled training sample, and the training set, wherein the training weight value is the product of the confidence corresponding to the first pre-labeled training sample and a preset hyperparameter.
8. The method according to any one of claims 1 to 5, characterized in that: The step of training the classification prediction model by using the pre-labeled training samples and the preset training set until the classification prediction model meets the preset stop training condition includes: Using the pre-labeled training samples and the preset training set to train the classification prediction model to obtain a candidate prediction model; Inputting a preset test set into the candidate prediction model to obtain test data; When the test data meets the expected test results, it is determined that the classification prediction model meets the preset stop training condition.
9. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Obtain target samples to be labeled; The target samples to be labeled are labeled according to the classification prediction model that meets the training stop condition.
10. A sample processing device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the sample processing method according to any one of claims 1 to 9 when executing the computer program.
11. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the sample processing method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Model training method and device based on semi-supervised learning and electronic equipment
CN111476256A
Training sample construction method and device, electronic equipment and storage medium
CN113590764A