A method, system, device and storage medium for improving the quality of a classification learning data set

By obtaining the error transfer probability matrix and error rate of image classification data sets, filtering and correcting error labels, the data set error problem caused by non-professional annotation is solved, and data set quality improvement and network performance improvement are achieved.

CN113919439BActive Publication Date: 2025-07-18NANJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111233079.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-22
Publication Date
2025-07-18
Estimated Expiration
2041-10-22

AI Technical Summary

Technical Problem

In the image classification, the data set error labels caused by non-professional annotator labeling errors exist, which affects the performance of the classifier and the high cost of expert labeling, resulting in a decrease in the data set scale and insufficient network training adequacy.

Method used

The error transfer probability matrix of the tag is obtained through the network output of the anchor sample, combined with the error rate and weight of the tag, the error label sample is selected, and the error transfer probability matrix is used to correct the error label, update the data set, and iterate multiple times until the proportion of clean labels in the data set increases.

Benefits of technology

While ensuring the data volume, the data set error rate is significantly reduced, network generalization performance is improved, the processing capability of error tag samples is improved, and data processing costs are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113919439B_ABST
    Figure CN113919439B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, device and storage medium for improving the quality of a classification learning data set, belonging to the technical field of image classification. The method includes: updating the data set by using a pre-designed updating method; outputting the data set in response to detecting that the proportion of clean labels in the data set does not increase; updating the data set again by using the pre-designed updating method in response to detecting that the proportion of clean labels in the data set increases; the pre-designed updating method includes: obtaining an error transfer probability matrix of labels through the network output of anchor samples; obtaining the error rate and weight of labels according to the error transfer probability matrix of labels, and obtaining the weighted average error rate of the data set according to the error rate and weight of labels; sorting data samples according to the probability of label annotation errors, screening out mislabeled samples in combination with the weighted average error rate of the data set, and correcting the labels of the mislabeled samples by using the error transfer probability matrix of labels to update the data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method, system, device and storage medium for improving the quality of a classification learning data set, and belongs to the technical field of image classification. Background Art

[0002] Neural networks have made great progress in the field of image processing technologies such as image classification and recognition, object detection, etc., which benefits from their ability to discover complex structures in high-dimensional data; usually, a large number of labeled data sets are required to process these complex problems. However, the acquisition of a large number of accurate and reliable labeled data sets is often very expensive and time-consuming; in recent years, crowdsourcing has gradually become the main solution for obtaining a large number of labeled data sets, which distributes data samples to a large number of non-professional annotators on the network for annotation. However, the abilities and preferences of each annotator are different, and the completed labels will have errors. For example, for some biomedical images, their sample annotation often requires professional knowledge, and non-professional annotators are very likely to mislabel the samples, thus generating a data set with incorrect labels. The existence of incorrect labels will deteriorate the performance of the trained classifier.

[0003] The prior art uses the memory characteristics of deep networks to screen and remove incorrect data in the data set to improve the quality of the data set, but the determination of the error level is a challenge. Usually, a small part of the data set is annotated by experts to estimate the error level of the entire data set. The introduction of experts increases the cost of data processing; at the same time, the incorrect label data selected is removed, reducing the scale of the data set and losing the guarantee of the sufficiency of network training. Summary of the Invention

[0004] The purpose of the present invention is to provide a method, system, device and storage medium for improving the quality of a classification learning data set, which, while ensuring the data volume, minimizes the error level of the data set, improves the generalization performance of the network, and enhances the ability of the classification network to handle samples with incorrect labels.

[0005] To achieve the above purpose, the present invention is implemented by adopting the following technical solutions:

[0006] In a first aspect, the present invention provides a method for improving the quality of a classification learning data set, including:

[0007] Updating the data set by using a pre-designed updating method; outputting the data set in response to detecting that the proportion of clean labels in the data set does not increase; and updating the data set again by using the pre-designed updating method in response to detecting that the proportion of clean labels in the data set increases;

[0008] The pre-designed updating method includes:

[0009] Obtain the error transition probability matrix of the labels from the network output of the anchor samples;

[0010] Obtain the error rate and weight of the labels according to the error transition probability matrix of the labels, and obtain the weighted average error rate of the dataset according to the error rate and weight of the labels;

[0011] Sort the data samples according to the probability of label annotation error, filter out the mislabeled samples in combination with the weighted average error rate of the dataset, and correct the labels of the mislabeled samples by combining the error transition probability matrix of the labels with the weight of the labels, and update the dataset.

[0012] Combined with the first aspect, further, the pre-designed update method further includes the step of obtaining anchor samples:

[0013] Directly train the network with the dataset containing mislabels to obtain the conditional probabilities of the data samples corresponding to various labels, and select the anchor samples of various labels according to the conditional probabilities.

[0014] Combined with the first aspect, further, the error transition probability matrix of the labels is obtained by the following method:

[0015]

[0016] where P is the error transition probability matrix of the labels, represents the probability that label i is mislabeled as label j, c is the total number of label categories, and i, j take values from 1,..., c.

[0017] Combined with the first aspect, further, the error rate and weight of the labels are obtained by the following method:

[0018] The error rate of the label is The weight of the label is where p ii represents the probability that label i is correctly labeled, c is the total number of label categories, m i is the observed number of i-class labels in the dataset, n j is the actual number of j-class labels in the dataset.

[0019] Combined with the first aspect, further, the weighted average error rate of the dataset is obtained by the following method:

[0020]

[0021] where wanr is the weighted average error rate of the dataset, p jj represents the probability that label j is correctly labeled, n j is the actual number of j-class labels in the dataset, m i is the observed number of i-class labels in the dataset, and c is the total number of label categories.

[0022] In combination with the first aspect, further, the labels of the mislabeled sample are corrected as follows:

[0023]

[0024] Where Y n-l is the corrected label, P is the error transition probability matrix of the label, is the network output of the mislabeled sample, X n-l is the mislabeled sample, is the set of mislabels.

[0025] In combination with the first aspect, further, it also includes the step of detecting whether the proportion of clean labels in the dataset increases:

[0026] If the trace of the error transition probability matrix of the label becomes larger, the proportion of clean labels in the dataset increases, otherwise the proportion of clean labels in the dataset does not increase.

[0027] In the second aspect, the present invention also provides a system for improving the quality of a classification learning dataset, including:

[0028] An update module: used to update the dataset by using a pre-designed update method;

[0029] An output module: used to output the dataset in response to detecting that the proportion of clean labels in the dataset does not increase;

[0030] A re-update module: used to re-update the dataset by using a pre-designed update method in response to detecting that the proportion of clean labels in the dataset increases.

[0031] In the third aspect, the present invention also provides a device for improving the quality of a classification learning dataset, including a processor and a storage medium;

[0032] The storage medium is used to store instructions;

[0033] The processor is used to operate according to the instructions to execute the steps of the method according to any one of the first aspects.

[0034] In the fourth aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method according to any one of the first aspects are implemented.

[0035] Compared with the prior art, the beneficial effects achieved by the present invention are:

[0036] A method, system, device, and storage medium for improving the quality of a classification learning data set provided by the present invention obtain an error transfer probability matrix of labels through the network output of anchor samples, and then obtain the error rate of the labels. Combining the sorting of the error probabilities of label annotations and the weighted average error rate of the data set, error label samples are screened out. The labels of the error label samples are corrected by using the error transfer probability matrix in combination with the weights of the labels, greatly reducing the error rate of the data set, improving the generalization performance of the network, and enhancing the ability of the classification network to handle error label samples. The error transfer probability matrix of the labels is obtained through the network output of the anchor samples, and then the error rate of the labels is obtained without using additional expert annotations, reducing the cost of data processing. At the same time, the screened error label samples are not removed, but the labels of the error label samples are corrected by using the error transfer probability matrix. Without reducing the data volume, the error level of the data set is reduced and the data set is updated. In response to detecting an increase in the proportion of clean labels in the data set, the data set is updated again using a pre-designed update method, and the update is performed multiple times until the proportion of clean labels in the data set does not increase, minimizing the error level of the data set. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 FIG. is a flowchart of a method for improving the quality of a classification learning data set provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the protection scope of the present invention.

[0039] Embodiment 1

[0040] As Figure 1 shown, the present invention provides a method for improving the quality of a classification learning data set, including the following steps:

[0041] S1. Update the data set using a pre-designed update method;

[0042] S2. Output the data set in response to detecting that the proportion of clean labels in the data set does not increase;

[0043] S3. Update the data set again using a pre-designed update method in response to detecting that the proportion of clean labels in the data set increases;

[0044] Among them, the pre-designed update method includes:

[0045] Obtain an error transfer probability matrix of labels through the network output of anchor samples;

[0046] Obtain the error rate and weight of the label according to the error transfer probability matrix of the label, and obtain the weighted average error rate of the dataset according to the error rate and weight of the label;

[0047] Sort the data samples according to the probability of label annotation error, screen out the mislabeled samples in combination with the weighted average error rate of the dataset, correct the labels of the mislabeled samples by using the error transfer probability matrix of the label in combination with the weight of the label, and update the dataset.

[0048] As Figure 1 shown, the complete operation flow chart of the present invention is shown. The dataset is continuously iteratively updated, and finally the error rate of the dataset label will be significantly lower than that of the original dataset.

[0049] In one data update, it can be roughly divided into four steps; in the first step, directly train the network using the dataset with mislabeled labels, output the probability that the data sample is labeled as various labels, and use the data sample with the maximum probability labeled as a certain label as the anchor sample of that label. Calculate the error transfer probability matrix of the label based on the network output of the anchor sample; in the second step, observe the error rate of each label according to the error transfer probability matrix of the label, and calculate the weighted average error rate of the dataset in combination with the actual weight of each label in the entire dataset; in the third step, use the characteristic that the deep network preferentially remembers the true label samples and lags in remembering the mislabeled samples during training, and the loss value distribution of the mislabeled samples is slightly larger than that of the correct label samples after stabilization. Sum the loss values of the data samples in each iterative training, sort the data samples in descending order according to the size of the loss sum, and the order of this sorting reflects the probability of mislabeling of each data sample. Then, screen out the data with a high probability of mislabeling according to the weighted average error rate; in the fourth step, for the screened mislabeled samples, correct their mislabeling to a certain extent according to the error transfer probability matrix of the label, so as to obtain a new dataset with a lower error rate without reducing the data volume.

[0050] The operation flow of the present invention is specifically as follows:

[0051] Step 1-1: The labeled dataset used to train the classifier is labeled by non-professionals, and the labels of some samples may be mislabeled; let the input sample set be and its corresponding noisy label set be The total number of label categories is c. Use the dataset with mislabeled labels to fully train the classifier network until the training accuracy tends to be stable. After stabilization, the output of the sample through the network is where the element represents the probability that the sample x is labeled as , and according to select the sample with the maximum probability labeled as label i in the dataset as the anchor sample of label i:

[0052] Step 1-2: Obtain the anchor samples of the labels from Step 1-1 where x i represents the anchor sample of label i, and based on the network output of the anchor sample calculate the error transition probability matrix of the quasi-calculated label where (i, j take values from 1 to c) represents the probability that label i is mislabeled as label j.

[0053] Step 2: In Step 1-2, the error transition probability matrix P of the labels is calculated. Due to the existence of mislabels in the dataset, the observed label distribution does not represent the actual distribution of the sample labels. Let the observed distribution of the sample labels be where m i is the total number of label i in the label set ; Assume the actual distribution of the sample labels is where n i is the actual number of label i in the dataset; and the error transition probability matrix of the labels satisfy the following equation:

[0054]

[0055] Solving the equation gives the actual distribution of the sample labels whose weight in the dataset is and the individual error rates of each type of sample are The weighted sum gives the weighted average error rate as This value can more accurately reflect the error level of the dataset.

[0056] Step 3-1: For the classification network in Step 1-1, use the dataset to train it for multiple rounds. Let the total number of training rounds be k, and the number of iterations in each round be t; During the training process, record the loss of each sample in each iteration where l i,j is the loss value of sample i in the jth iteration. Sum the losses of each iteration where Sort the loss sums in descending order to get B = rank(L). The order of this sequence reflects the probability size of the sample being labeled as a mislabel.

[0057] Step 3-2: In 3-1, obtain the probability size ranking R of the sample being labeled as a mislabel. Combine with the weighted average error rate wanr of the dataset obtained in Step 2, and select the samples included in the top-wanr of the ranking B as the mislabel samples Xn-l = {x n-l}。

[0058] Step 4: For the filtered error label samples X n-l , its corresponding error label is With the help of the error transition probability matrix P of the labels in Step 2 and the weights of various labels in the overall data set obtained correct the error label as follows: where Y n-l is the corrected label. After completion of the correction, merge it with the data remaining in Step 3-2 to form a new data set, and the error rate of this new data set will be lower than that of the original data set.

[0059] Step 5: Iteratively update the updated data set according to the above steps to further reduce the error level of the data set. Observe the change of the trace tr(P) of the error transition probability matrix P of the labels during the iteration. If it no longer increases significantly, it indicates that the update can no longer effectively reduce the error level of the data set, and stop the update; output the network and verify the performance on the test set, and this network will be able to complete the classification task more accurately.

[0060] Example Two

[0061] The embodiment of the present invention also provides a system for improving the quality of a classification learning data set, including:

[0062] Update module: used to update the data set by using a pre-designed update method;

[0063] Output module: used to output the data set in response to detecting that the proportion of clean labels in the data set does not increase;

[0064] Re-update module: used to update the data set again by using a pre-designed update method in response to detecting that the proportion of clean labels in the data set increases.

[0065] Example Three

[0066] The embodiment of the present invention also provides a device for improving the quality of a classification learning data set, including a processor and a storage medium;

[0067] The storage medium is used to store instructions;

[0068] The processor is used to operate according to the instructions to execute the steps of the following method:

[0069] S1. Update the data set by using a pre-designed update method;

[0070] S2. Output the data set in response to detecting that the proportion of clean labels in the data set does not increase;

[0071] S3. In response to detecting an increase in the proportion of clean labels in the dataset, update the dataset again using a pre-designed update method;

[0072] Among them, the pre-designed update method includes:

[0073] Obtain the error transition probability matrix of the labels through the network output of the anchor samples;

[0074] Obtain the error rate and weight of the labels according to the error transition probability matrix of the labels, and obtain the weighted average error rate of the dataset according to the error rate and weight of the labels;

[0075] Sort the data samples according to the probability of label annotation error, screen out the mislabeled samples in combination with the weighted average error rate of the dataset, and use the error transition probability matrix of the labels in combination with the weight of the labels to correct the labels of the mislabeled samples and update the dataset.

[0076] Example 4

[0077] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the following method are implemented:

[0078] S1. Update the dataset using a pre-designed update method;

[0079] S2. In response to detecting that the proportion of clean labels in the dataset does not increase, output the dataset;

[0080] S3. In response to detecting an increase in the proportion of clean labels in the dataset, update the dataset again using a pre-designed update method;

[0081] Among them, the pre-designed update method includes:

[0082] Obtain the error transition probability matrix of the labels through the network output of the anchor samples;

[0083] Obtain the error rate and weight of the labels according to the error transition probability matrix of the labels, and obtain the weighted average error rate of the dataset according to the error rate and weight of the labels;

[0084] Sort the data samples according to the probability of label annotation error, screen out the mislabeled samples in combination with the weighted average error rate of the dataset, and use the error transition probability matrix of the labels in combination with the weight of the labels to correct the labels of the mislabeled samples and update the dataset.

[0085] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0086] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.

[0087] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.

[0088] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.

[0089] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A method for improving the quality of a classification learning data set, characterized in that, Including: Updating the dataset by the following method: Directly training a network using the dataset with incorrect labels to obtain the conditional probabilities of each type of label corresponding to the data samples, and selecting the anchor samples of each type of label according to the conditional probabilities; Based on the anchor samples, obtaining the error transition probability matrix of the labels through the following formula: Among them, P is the error transition probability matrix of the labels, which represents the probability that label i is mislabeled as label j. c is the total number of label categories, and i, j take values from 1 to c. x i represents the anchor sample of label i; The error rate and weight of the label are obtained by the following method: The error rate of the label is The weight of the label is where p ii represents the probability that the label i is correctly annotated, c is the total number of label categories, m i is the observed number of the i-th type of label in the dataset, n j is the actual number of the j-th type of label in the dataset; Obtaining the weighted average error rate of the dataset according to the error rate and weight of the labels; Sorting the data samples according to the probability of incorrect label annotation, screening out the incorrect label samples in combination with the weighted average error rate of the dataset, and correcting the labels of the incorrect label samples by combining the error transition probability matrix of the labels with the weight of the labels to update the dataset; Correcting the labels of the incorrect label samples, including: Among them, Y n-l is the corrected label, P is the error transition probability matrix of the label, is the network output of the mislabeled sample, X n-l is the mislabeled sample, is the mislabeled set; If the proportion of clean labels in the dataset does not increase, outputting the dataset; otherwise, updating the dataset again; The dataset is a biomedical image dataset.

2. The method for improving the quality of a classification learning data set according to claim 1, characterized in that The weighted average error rate of the dataset is obtained by the following method: Among them, wanr is the weighted average error rate of the dataset, p jj represents the probability that the label j is correctly annotated, n j is the actual number of class j labels in the dataset, m i is the observed number of class i labels in the dataset, and c is the total number of label categories.

3. The method for improving the quality of a classification learning data set according to claim 1, wherein, It also includes the step of detecting whether the proportion of clean labels in the dataset increases: If the trace of the error transition probability matrix of the labels becomes larger, the proportion of clean labels in the dataset increases; otherwise, the proportion of clean labels in the dataset does not increase.

4. A system for improving the quality of a classification learning data set, characterized in that, Including: Update module: used to update the dataset by the following method: Directly training a network using the dataset with incorrect labels to obtain the conditional probabilities of each type of label corresponding to the data samples, and selecting the anchor samples of each type of label according to the conditional probabilities; Based on the anchor samples, obtaining the error transition probability matrix of the labels through the following formula: where P is the error transfer probability matrix of the labels, which represents the probability that label i is mislabeled as label j, c is the total number of label categories, i, j take values from 1, …, c, and x i represents the anchor sample of label i; Obtain the error rate and weight of the label according to the error transfer probability matrix of the label, including: the error rate of the label is The weight of the label is where p ii represents the probability that label i is correctly labeled, c is the total number of label categories, m i is the number of observations of the i-th class label in the dataset, n j is the actual number of the j-th class label in the dataset; Obtaining the weighted average error rate of the dataset according to the error rate and weight of the labels; Sorting the data samples according to the probability of incorrect label annotation, screening out the incorrect label samples in combination with the weighted average error rate of the dataset, and correcting the labels of the incorrect label samples by combining the error transition probability matrix of the labels with the weight of the labels to update the dataset; Correcting the labels of the incorrect label samples, including: Among them, Y n-l is the corrected label, P is the error transition probability matrix of the label, is the network output of the mislabeled sample, X n-l is the mislabeled sample, is the mislabeled set; Output module: used to output the dataset in response to detecting that the proportion of clean labels in the dataset does not increase; Re-update module: used to update the dataset again in response to detecting that the proportion of clean labels in the dataset increases; The dataset is a biomedical image dataset.

5. An apparatus for improving the quality of a classification learning data set, characterized in that, Including a processor and a storage medium; The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Denoising method and device, computer equipment, storage medium and model training method

    CN110929733A