Method, apparatus, computer device, and storage medium for mislabeled data recognition

By enhancing and training the data set of mislabeled data, using discarded layers and multiple prediction methods, the loss value is calculated and sorted to filter mislabeled data, the problem of low filtering accuracy in the existing technology is solved, and higher screening accuracy and model generalization are achieved.

CN113536069BActive Publication Date: 2025-05-30SHENZHEN ZHUIYI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110891701.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-04
Publication Date
2025-05-30
Estimated Expiration
2041-08-04

AI Technical Summary

Technical Problem

The prior art has low accuracy in the selection of false annotation data, especially when the average value of the sample loss value of the correct annotation data and the wrong annotation data is close to or the same, it is difficult to effectively distinguish.

Method used

By enhancing the data set containing mislabeled data, a second data set is generated, and the second data set is input to the neural network model containing the discarded layer for training, a trained prediction model is obtained. Then, by predicting and recording the loss values ​​multiple times, the first loss average, the second loss average, and the loss variance value of each subdata are calculated, and the data set is sorted according to these values, thereby filtering out the mislabeled data.

Benefits of technology

The accuracy of mislabeled data filtering is improved, and the loss value difference between correct data and mislabeled data is amplified through the use of data enhancement and discarding layers, the overfitting of mislabeled data is reduced, and the generalization and robustness of the model is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113536069B_ABST
    Figure CN113536069B_ABST
Patent Text Reader

Abstract

The present application relates to a method, apparatus, computer device, and storage medium for mislabeled data recognition. The method includes: augmenting a first data set to obtain a second data set; inputting the second data set into a neural network model with a dropout layer for model training to obtain a prediction model, and obtaining the first loss average value of each sub-data in the second data set at each training epoch during the training process; performing multiple predictions on the first data set through the prediction model, and at each prediction, discarding some neurons in the prediction model through the dropout layer for prediction; obtaining the second loss average value and the loss variance value of each sub-data according to the second loss value of each sub-data at each prediction; sorting each sub-data in the first data set according to the first loss average value, the second loss average value, and the loss variance value to screen mislabeled data. The solution of the present application can improve the accuracy of mislabeled data screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and particularly relates to a method, apparatus, computer device, and storage medium for identifying mislabeled data. Background Art

[0002] Data is the core of machine learning, and the quality of data will greatly affect the performance of the model. However, both manually labeled data sets and data sets collected from the Internet contain mislabeled data, and cleaning mislabeled data is a time-consuming and laborious project.

[0003] Currently, the main method for screening mislabeled data generally uses the average value of the sample loss values during the model training process as the screening feature. However, sometimes the average values of the sample loss values of correctly labeled data and mislabeled data are relatively close or the same. Therefore, it is not easy to distinguish between correctly and mislabeled data, and it is easy to reduce the accuracy of mislabeled data screening. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a method, apparatus, computer device, and storage medium for identifying mislabeled data that can improve the accuracy of mislabeled data screening.

[0005] A method for identifying mislabeled data, the method comprising:

[0006] Enhancing a first data set containing mislabeled data to obtain a second data set;

[0007] Inputting the second data set into a neural network model with a dropout layer, training the neural network model using the second data set to obtain a trained prediction model including the dropout layer, and obtaining the first loss average value of each sub-data in the second data set during each training epoch in the training process;

[0008] Performing multiple predictions on the first data set through the prediction model. During each prediction, discarding some neurons in the prediction model through the dropout layer, and using the prediction model after discarding the neurons for prediction;

[0009] Obtaining the second loss average value and the loss variance value of each sub-data according to the second loss value of each sub-data in the first data set during each prediction;

[0010] Sorting each sub-data in the first data set according to the first loss average value, the second loss average value, and the loss variance value, and screening mislabeled data from the first data set according to the sorting result.

[0011] A mislabeled data identification apparatus, the apparatus comprising:

[0012] A data processing module for enhancing a first data set containing mislabeled data to obtain a second data set;

[0013] A training statistics module for inputting the second data set into a neural network model with a dropout layer, training the neural network model using the second data set to obtain a trained prediction model including the dropout layer, and obtaining an average first loss value of each sub-data according to the first loss value of each sub-data in the second data set at each training epoch during the training process;

[0014] A prediction statistics module for making multiple predictions on the first data set through the prediction model. At each prediction, a part of the neurons in the prediction model are discarded through the dropout layer, and the prediction model after discarding the neurons is used for prediction; according to the second loss value of each sub-data in the first data set at each prediction, an average second loss value and a loss variance value of each sub-data are obtained;

[0015] A sorting and screening module for sorting each sub-data in the first data set according to the average first loss value, the average second loss value, and the loss variance value, and screening mislabeled data from the first data set according to the sorting result.

[0016] In one embodiment, the prediction statistics module is further configured to randomly discard a part of the neurons in the prediction model through the dropout layer at each prediction; or, at each prediction, discard a part of the neurons in the prediction model through the dropout layer according to a preset dropout rule.

[0017] In one embodiment, the prediction statistics module is further configured to average the second loss values in the multiple prediction results of each sub-data in the first data set to obtain the average second loss value of each sub-data; according to the second loss value of each sub-data in the multiple prediction results of the first data set and the average second loss value of the sub-data, the loss variance value of each sub-data is obtained.

[0018] In one embodiment, the training statistics module includes:

[0019] An accuracy monitoring module for statistically calculating the accuracy of the neural network model at each training epoch during the process of training the neural network model using the second data set; comparing the accuracy of the current training epoch with the accuracy of the previous training epochs respectively; the previous training epochs are at least one; if the accuracy of the current training epoch is less than the accuracy of the previous training epochs, stop training to obtain a trained prediction model including the dropout layer.

[0020] In one embodiment, the sorting and screening module is further configured to initially sort each sub-data of the first data set according to the first loss average value; if there is sub-data with the same first loss average value in the first data set, then for the sub-data with the same first loss average value, perform a secondary sorting according to the second loss average value corresponding to the sub-data; if there is sub-data with the same second loss average value in the first data set after the secondary sorting, then for the sub-data with the same second loss average value, perform an advanced sorting according to the loss variance value corresponding to the sub-data.

[0021] In one embodiment, the sorting and screening module is further configured to sort each sub-data in the first data set according to the first loss average value, the second loss average value, and the loss variance value respectively, to obtain the ranking positions of the first loss average value, the second loss average value, and the loss variance value of each sub-data; perform a weighted summation operation on the ranking positions of the first loss average value, the second loss average value, and the loss variance value of each sub-data in the first data set; and sort according to the weighted summation operation results of each sub-data in the first data set.

[0022] In one embodiment, the sorting and screening module is further configured to determine the sub-data ranked in the top preset positions according to the sorting result; and determine the sub-data ranked in the top preset positions as mislabeled data.

[0023] A computer device includes a memory and a processor, the memory stores a computer program, and the processor executes the steps of the above mislabeled data recognition method.

[0024] A computer-readable storage medium stores a computer program thereon, and the computer program is executed by a processor to perform the steps of the above mislabeled data recognition method.

[0025] The above mislabeled data identification, device, computer equipment and storage medium enhance the first data set containing mislabeled data to obtain a second data set. By enhancing the first data set, the generalization and robustness of the training model can be improved, and the difference in loss values between correct data and mislabeled data can be amplified. Input the second data set into a neural network model with a dropout layer. During training, the neural network model reduces the complex co-adaptation relationship between neurons through the dropout layer. Use the second data set to train the neural network model to obtain a trained prediction model including the dropout layer, and obtain the average first loss value of each sub-data according to the first loss value of each sub-data in the second data set at each training epoch during the training process. Perform multiple predictions on the first data set through the prediction model. During each prediction, discard some neurons in the prediction model through the dropout layer, and use the prediction model after discarding the neurons for prediction. Obtain the average second loss value and loss variance value of each sub-data according to the second loss value of each sub-data in the first data set during each prediction. Sort each sub-data in the first data set according to the average first loss value, the average second loss value and the loss variance value, and screen mislabeled data from the first data set according to the sorting result. The screening conditions are diverse, avoiding misjudgment due to being limited by a single condition. When training the model, enable the dropout layer to prevent overfitting to mislabeled data, and enhance the first data set, thereby amplifying the loss value gap between correct and mislabeled data, so as to improve the accuracy of mislabeled data screening. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is an application environment diagram of the mislabeled data identification method in an embodiment;

[0027] Figure 2 It is a schematic flowchart of the mislabeled data identification method in an embodiment;

[0028] Figure 3 It is a schematic flowchart of the model training step in an embodiment;

[0029] Figure 4 It is a structural block diagram of the mislabeled data identification device in an embodiment;

[0030] Figure 5 It is a structural block diagram of the prediction statistics module in an embodiment;

[0031] Figure 6 It is an internal structural diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0033] The mislabeled data recognition method provided by the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal 110 communicates with the server 120 through a network. Among them, the terminal 110 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices, and the server 120 can be implemented by an independent server or a server cluster composed of multiple servers.

[0034] The server 120 can enhance the first data set containing mislabeled data to obtain a second data set; the server 120 can input the second data set into a neural network model containing a dropout layer, and use the second data set to train the neural network model to obtain a trained prediction model including the dropout layer, and obtain the average first loss value of each sub-data according to the first loss value of each sub-data in the second data set at each training epoch during the training process; the server 120 can perform multiple predictions on the first data set based on the prediction model. Each time a prediction is made, a part of the neurons in the prediction model are discarded through the dropout layer, and the prediction model after discarding the neurons is used for prediction; according to the second loss value of each sub-data in the first data set at each prediction, the average second loss value and the loss variance value of each sub-data are obtained; according to the average first loss value, the average second loss value, and the loss variance value, the sub-data in the first data set are sorted, and the server 120 can screen the mislabeled data from the first data set according to the sorting result.

[0035] In one embodiment, the first data set containing mislabeled data can be sent from the terminal 110 to the server 120. The server 120 can execute the mislabeled data screening method in each embodiment of the present application, screen the mislabeled data from the first data set, and then send the mislabeled data to the terminal 110.

[0036] In one embodiment, as Figure 2 shown, a mislabeled data recognition method is provided. In this embodiment, an example is given in which this method is applied to a server. It can be understood that this method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0037] S202. Enhance the first data set containing mislabeled data to obtain a second data set.

[0038] Among them, mislabeled data refers to data with incorrect labels. The methods in the embodiments of the present application aim to identify mislabeled data from the first dataset.

[0039] Dataset augmentation refers to the process of transforming the data in the dataset to generate new data. It can be understood that dataset augmentation is mainly to reduce the overfitting phenomenon of the neural network model. By transforming the data in the dataset and then training the neural network model, a neural network model with stronger generalization ability can be obtained.

[0040] In one embodiment, there are multiple augmentation methods. For example, at least one of the following processes can be performed on the data in the dataset: rotation, cropping, translation transformation, and noise perturbation, etc., to achieve the data augmentation effect.

[0041] In one embodiment, the augmentation method can include at least one of online augmentation and offline augmentation. Offline augmentation is to transform all the data in the dataset at one time before starting training. Online augmentation is to perform augmentation on part of the data in the dataset in each training epoch.

[0042] Specifically, the server can perform augmentation processing on the data in the first dataset containing mislabeled data to obtain a second dataset, and use the second dataset obtained after the augmentation processing as the training data for training the prediction model and execute step S204.

[0043] In one embodiment, the server can adopt different augmentation methods for the first dataset according to the size of the data volume of the first dataset.

[0044] Specifically, when the data volume of the first dataset is greater than a preset quantity threshold, that is, when the data volume of the first dataset is relatively large, online augmentation can be performed on the first dataset. That is, in each training epoch during the training process of the prediction model, part of the data in the first dataset can be augmented (i.e., online augmentation) to obtain the corresponding second dataset as the training data for that training epoch.

[0045] In one embodiment, the server can divide the first dataset into multiple portions of data, and perform augmentation on each portion of data in each training epoch during the training process of the prediction model to obtain the corresponding second dataset. It can be understood that by augmenting multiple portions of data respectively, multiple second datasets can be obtained, and subsequently, step S204 can be executed for each second dataset to train the prediction model. For example, if the first dataset includes 50,000 data, the server can evenly divide these 50,000 data into 100 portions, with 500 data in each portion, and perform random flipping or random cropping or adding noise (Gaussian, salt and pepper) to each of these 500 data in each training epoch to obtain the second dataset.

[0046] When the data volume of the first data set is less than or equal to a preset quantity threshold, that is, when the data volume of the first data set is relatively small, offline enhancement can be performed on the first data set, that is, before starting training, overall enhancement is performed on each data in the first data set to obtain a second data set. For example, if the first data set includes 500 data, after performing enhancement processing such as random flipping, random cropping, or adding noise (Gaussian, salt and pepper) to each sub-data among the 500 data before training, a second data set is obtained.

[0047] S204, input the second data set into a neural network model containing a dropout layer, use the second data set to train the neural network model to obtain a trained prediction model including the dropout layer, and obtain the average first loss value of each sub-data according to the first loss value of each sub-data in the second data set at each training epoch during the training process.

[0048] Among them, the dropout layer is used to temporarily turn off neurons in the neural network model with a certain probability, which can prevent overfitting of the network to a certain extent.

[0049] Specifically, the server can input the second data set into a neural network model containing a dropout layer, use the second data set to train the neural network model to obtain a trained prediction model including the dropout layer. It can be understood that if the neural network model is trained in multiple training epochs, then in each training epoch, the corresponding second data set can be used to train the neural network model. Then, in each training epoch, the first loss value of each sub-data in the corresponding second data set can be obtained. For each sub-data in the second data set, there are corresponding different first loss values in different training epochs. The server can perform an average calculation on the multiple first loss values corresponding to the same sub-data in different training epochs to obtain the average first loss value of each sub-data.

[0050] It can be understood that if offline enhancement processing is performed on the first data set in step 102, then in each training epoch, the second data set after offline enhancement processing can be used for model training. If online enhancement processing is performed on the first data set, that is, the second data set obtained by performing online enhancement processing on some data in the first data set in each training epoch, then the second data set after online enhancement processing in each training epoch is used for model training.

[0051] In one embodiment, during the process of training a neural network model using a second data set, some neurons can be randomly and temporarily turned off according to a specified dropout probability through a dropout layer to train the neural network model. In this way, the training process does not particularly rely on the features extracted from a certain neuron, preventing overfitting. For example, if the dropout probability is 0.2, the dropout layer can randomly and temporarily turn off some neurons with a dropout probability of 0.2.

[0052] It can be understood that by discarding some neurons through the dropout layer of the neural network model, the network parameters can be updated during the training period, reducing the complex co-adaptation relationship between neurons, preventing overfitting to the data set, thereby preventing the model from overfitting to mislabeled data, amplifying the loss difference between the correct data and the mislabeled data during the training process, and thus improving the accuracy of model training.

[0053] In one embodiment, the server can perform iterative training on the neural network model using the second data set. During the training process, the accuracy of the neural network model can be detected, and it can be determined whether to stop training based on the growth of the accuracy. If the growth of the accuracy meets the preset training stop condition, the training can be stopped to obtain a trained prediction model including a dropout layer.

[0054] In one embodiment, the preset training stop condition may include: the growth rate of the accuracy is less than a preset speed threshold or there is no growth. That is, when the growth rate of the accuracy is less than the preset speed threshold or there is no growth, the training can be stopped.

[0055] It can be understood that during the period of training the neural network model, mislabeled data is more difficult to learn compared to correctly labeled data. Therefore, in the early stage of training the neural network model, the accuracy of the second data set will increase rapidly, while it will increase slowly in the later stage, and the model begins to fit the mislabeled data. When the accuracy increases slowly or there is no increase in the accuracy, stopping the training can prevent the neural network model from fitting the mislabeled data, thereby reducing the impact of the mislabeled data on the neural network model during the training process, and thus widening the loss value between the mislabeled data and the correctly labeled data.

[0056] S206, perform multiple predictions on the first data set based on the prediction model. In each prediction, discard some neurons in the prediction model through the dropout layer, and use the prediction model after discarding the neurons to make predictions; according to the second loss value of each sub-data in the first data set in each prediction, obtain the second loss average value and the loss variance value of each sub-data.

[0057] Among them, variance is a measure of the degree of dispersion when measuring a random variable or a set of data in probability theory and statistical variance. The loss variance value is the value obtained by calculating the variance of the loss values of each sub-data in the first data set. The second loss average value is the value obtained by averaging the loss values of each sub-data in the first data set. It can be understood that by discarding some neurons through the dropout layer of the neural network model, the network parameters can be updated, and the uncertainty of the prediction of the mislabeled data by the neural network model can be amplified. Therefore, the second loss value of the mislabeled data will be larger than that of the correctly labeled data, and the change in the second loss value is also large. Perform multiple predictions, and count the magnitude and change range of the loss values of each sub-data in multiple prediction periods.

[0058] Specifically, the server predicts the first data set through the prediction model. When predicting, some neurons are discarded through the dropout layer of the prediction model, and multiple predictions are performed. The second loss value of the sub-data in each prediction period is recorded. Then, the average calculation and variance calculation are respectively performed on the second loss values of each sub-data in the first data set to obtain the second loss average value and the loss variance value of each sub-data, so as to judge the average magnitude and change range of the second loss values obtained by the sub-data in multiple predictions. It can be understood that the loss variance value can be used to characterize the change range.

[0059] S208. Sort the sub-data in the first data set according to the first loss average value, the second loss average value, and the loss variance value, and screen out the mislabeled data from the first data set according to the sorting result.

[0060] Specifically, the server can sort the sub-data in the first data set according to the first average value, the second loss average value, and the loss variance value. The server can determine the sub-data ranked in the top preset positions or the sub-data ranked in the bottom preset positions as mislabeled data according to the final sorting result.

[0061] In one embodiment, if the final sorting result is in descending order and the loss values of the sub-data ranked in the front are relatively large, the server can select the sub-data ranked in the top preset positions as mislabeled data.

[0062] In another embodiment, if the final sorting result is in ascending order and the loss values of the sub-data ranked in the back are relatively large, the server can select the sub-data ranked in the bottom preset positions as mislabeled data.

[0063] The above mislabeled data identification, device, computer equipment, and storage medium enhance the first data set containing mislabeled data to obtain a second data set. By enhancing the first data set, the generalization and robustness of the training model can be improved, and the difference in loss values between correct data and mislabeled data can be amplified. Input the second data set into a neural network model with a dropout layer. During training, the neural network model reduces the complex co-adaptation relationship between neurons through the dropout layer. Use the second data set to train the neural network model to obtain a trained prediction model including the dropout layer, and obtain the first loss average value of each sub-data in the second data set at each training epoch during the training process. Based on the first data set, perform multiple predictions through the prediction model. At each prediction, discard some neurons in the prediction model through the dropout layer, and use the prediction model after discarding the neurons to make predictions. According to the second loss value of each sub-data in the first data set at each prediction, obtain the second loss average value and loss variance value of each sub-data. Sort the sub-data in the first data set according to the first loss average value, second loss average value, and loss variance value, and screen the mislabeled data from the first data set according to the sorting result. The screening conditions are diversified to avoid misjudgment due to being limited to a single condition. When training the model, enable the dropout layer to prevent overfitting to mislabeled data, and enhance the first data set, thereby amplifying the loss value gap between correct and mislabeled data, and improving the accuracy of mislabeled data screening.

[0064] In one embodiment, at each prediction, discarding some neurons in the prediction model through the dropout layer includes: at each prediction, randomly discard some neurons in the prediction model through the dropout layer; or, at each prediction, discard some neurons in the prediction model through the dropout layer according to a preset dropout rule.

[0065] Among them, the dropout layer is used to temporarily turn off the neurons in the prediction model with a certain probability. It can be understood that the prediction model makes multiple predictions for the same labeled data, and the dropout layer temporarily turns off some neurons at each prediction, so that the results of multiple predictions of the prediction model for the same labeled data may be different. Compared with correctly labeled data, for mislabeled data, the correct information required by the neurons contained in the mislabeled data is less than that of correctly labeled data, so mislabeled data is more likely to show instability in the results of multiple predictions.

[0066] In one embodiment, when making each prediction, the server can randomly and temporarily turn off some neurons through a dropout layer according to a specified dropout probability. For example, if the dropout probability is 0.2, the dropout layer can randomly and temporarily turn off some neurons with a dropout probability of 0.2. After multiple predictions like this, compared with correctly labeled data, the loss values output by the mislabeled data for multiple predictions show instability and a larger variation range.

[0067] In another embodiment, when making each prediction, the server can discard some neurons in the prediction model through the dropout layer according to a preset dropout rule. For example, for neurons of different categories, different dropout probabilities are specified to randomly and temporarily turn off a part of them; after multiple predictions like this, compared with correctly labeled data, the loss values output by the mislabeled data for multiple predictions show instability and a larger variation range.

[0068] In this embodiment, by discarding some neurons through the dropout layer of the prediction model, among the results of multiple predictions on the same labeled data, the loss values of the mislabeled data are more likely to show instability, so as to distinguish between correctly labeled data and mislabeled data according to the loss values obtained from the predictions, and improve the accuracy of screening mislabeled data.

[0069] In one embodiment, obtaining the second loss average value and the loss variance value of each sub - data according to the second loss value of each sub - data in the first data set during each prediction in S206 includes: averaging the second loss values of each sub - data in the first data set in multiple prediction results to obtain the second loss average value of each sub - data; and obtaining the loss variance value of each sub - data according to the second loss value of each sub - data in the first data set in multiple prediction results and the second loss average value of the sub - data.

[0070] Specifically, the server can record the second loss value of each sub - data in the first data set in each prediction result. After multiple predictions are completed, the server can sum up the second loss values and then calculate the average value according to the number of predictions to obtain the second loss average value. It can be understood that the second loss average value can represent the magnitude of the sub - data in multiple prediction results.

[0071] It can be understood that the loss variance value is calculated based on the second loss value of each sub - data in the first data set in multiple prediction results and the second loss average value of the sub - data, and can represent the variation range of the loss values of the sub - data in multiple prediction results.

[0072] In one embodiment, the server records the second loss value of each sub-data under each prediction and records the total number of predictions; sums up the second loss values of each sub-data to obtain the total second loss value, and divides the total second loss value by the total number of predictions to obtain the average second loss value; the server then calculates the loss variance value of each sub-data based on the second loss value and the average second loss value of each sub-data in multiple prediction results. Thus, through the calculated average second loss value and loss variance value, the server can know the magnitude and the variation range of the loss values of each sub-data of the first sub-data in multiple predictions, thereby providing more accurate information for mislabeling screening.

[0073] In this embodiment, by averaging the second loss values of each sub-data of the first data set in multiple prediction results, the average second loss value of each sub-data is obtained; based on the second loss value of each sub-data of the first data set in multiple prediction results and the average second loss value of the sub-data, the loss variance value of each sub-data is obtained; in this way, the magnitude of the loss value and the variation range of the loss value of the sub-data in the prediction process are obtained, providing more accurate information for screening mislabeled data.

[0074] In one embodiment, as Figure 3 shown, training the neural network model with the second data set in step S204 to obtain a trained prediction model including a dropout layer includes the following steps:

[0075] S302. During the process of training the neural network model with the second data set, the accuracy rate of the neural network model in each training epoch is statistically recorded.

[0076] Specifically, each training epoch is identified to distinguish training epochs. The server can, during the process of training the neural network model with the second data set, correspondingly record the accuracy rate of each training epoch and the identifier of the training epoch, and each training epoch identifier corresponds to an accuracy rate information. For example, the server can use numbers to represent training epochs, 1 represents the first training epoch, 2 represents the second training epoch, 3 represents the third training epoch, then the data recorded from the first training epoch to the third training epoch can be {{1, 40%}, {2, 50%}, {3, 60%}}.

[0077] In other embodiments, the server can also use other methods to correspondingly record the accuracy rate and the training epoch, for example, establishing a mapping correspondence relationship between the training epoch and the accuracy rate.

[0078] S304. Compare the accuracy rate of the current training epoch with the accuracy rates of the previous training epochs respectively; if the accuracy rate of the current training epoch is less than the accuracy rates of the previous training epochs, stop the training and obtain the trained prediction model including the dropout layer.

[0079] Among them, the previous training epoch refers to at least one consecutive epoch before the current training epoch. In one embodiment, the previous training epoch can be the nearest training epoch adjacent to the current training epoch before the current training epoch, or multiple consecutive training epochs adjacent to the current training epoch before the current training epoch. For example, the first training epoch is labeled as 1, the second training epoch is labeled as 2, and so on. The label of the nth training epoch is n. If the label of the current training epoch is n, the label of the previous training epoch can be n - 1, n - 2, n - 3 adjacent to the label of the current training epoch, or a training epoch n - 1 adjacent to the current training epoch.

[0080] Specifically, after each training epoch ends, compare the accuracy rate of the training epoch where it is located (i.e., the current training epoch) with the accuracy rate of the previous training epoch. If both are less than the accuracy rate of the previous training epoch, stop the training.

[0081] It can be understood that by recording the accuracy rate of each training time and labeling each training epoch, calculations related to the accuracy rate are performed after each training epoch ends, so as to know that the accuracy rate of the neural network model has increased slowly or not increased, and further know that the neural network model starts to fit mislabeled data. At this time, the server can stop the training, prevent the neural network model from fitting mislabeled data, reduce the impact of mislabeled data on the neural network model, and thus widen the loss value between the mislabeled data and the correctly labeled data.

[0082] For example, the label of the current training epoch is 500 and the accuracy rate is 80%. The recorded training epoch labels and corresponding accuracy rate information of the previous training epochs are {{499, 83%}, {498, 83%}, {497, 84%}}. Determine that the accuracy rate of 80% in the training epoch with the label of 500 is less than the accuracy rates corresponding to 499, 498, and 497, and then stop the training to prevent the neural network model from fitting mislabeled data.

[0083] In this embodiment, during the training of the neural network model, the accuracy rate of each training epoch is statistically calculated. When it is determined that the accuracy rate of the current training epoch is less than the accuracy rates of the previous training epochs, the training is stopped to prevent the model from fitting mislabeled data, thereby reducing the impact of mislabeled data on the model during the training process, widening the difference between the correctly labeled data and the mislabeled data, and providing more accurate information for the screening of mislabeled data.

[0084] In one embodiment, sorting the sub - data in the first data set according to the first loss average value, the second loss average value, and the loss variance value includes: initially sorting the sub - data in the first data set according to the first loss average value; if there are sub - data with the same first loss average value in the first data set, then for the sub - data with the same first loss average value, performing a secondary sorting according to the second loss average value corresponding to the sub - data; if there are sub - data with the same second loss average value in the first data set after the secondary sorting, then for the sub - data with the same second loss average value, performing an advanced sorting according to the loss variance value corresponding to the sub - data.

[0085] It can be understood that sorting the sub - data in the first data set using the first loss average value, the second loss average value, and the loss variance value belongs to multi - condition sorting; in multi - condition sorting, sorting can be performed by prioritizing the conditions. For example, taking the first condition as the main one and the second and third conditions as the secondary ones; first sort according to the first condition, and for the data with the same first condition, further sort according to the second condition. If the first condition and the second condition are the same, further sort according to the third condition.

[0086] It can be understood that by prioritizing the three conditions of the first loss average value, the second loss average value, and the loss variance value to sort the sub - data in the first data set, ordered sub - data can be obtained, which is convenient for finding sub - data with relatively large loss values and relatively large variation ranges of loss values.

[0087] In one embodiment, the server sorts the sub - data in the first data set in ascending order according to the first loss average value; for the sub - data with the same first loss average value in the first data set, performing a secondary ascending sorting according to the second loss average value corresponding to the sub - data; if there are sub - data with the same second loss average value in the first data set after the secondary sorting, then for the sub - data with the same second loss average value, performing an ascending sorting according to the loss variance value corresponding to the sub - data, obtaining ordered data, which is convenient for screening and also provides more diverse information for mis - annotation screening.

[0088] In this embodiment, initially sorting the sub - data in the first data set according to the first loss average value; if there are sub - data with the same first loss average value in the first data set, then for the sub - data with the same first loss average value, performing a secondary sorting according to the second loss average value corresponding to the sub - data; if there are sub - data with the same second loss average value in the second data set after the secondary sorting, then for the sub - data with the same second loss average value, performing an advanced sorting according to the loss variance value corresponding to the sub - data. In this way, the screening of mis - annotated data no longer only relies on the information of the first loss average value, but also adds the second loss average value and the loss variance value, improving the accuracy of mis - annotation screening.

[0089] In one embodiment, sorting the sub - data in the first data set according to the first loss average value, the second loss average value, and the loss variance value includes: respectively sorting the sub - data in the first data set according to the first loss average value, the second loss average value, and the loss variance value to obtain the ranking positions of the first loss average value, the ranking positions of the second loss average value, and the ranking positions of the loss variance value for each sub - data; performing a weighted summation operation on the ranking positions of the first loss average value, the ranking positions of the second loss average value, and the ranking positions of the loss variance value for each sub - data in the first data set; and sorting according to the results of the weighted summation operation for each sub - data in the first data set.

[0090] Among them, weighted summation means multiplying and then adding and summing according to the importance coefficients (i.e., weights) of each data. The weighted summation value corresponding to each sub - data in the first data set is obtained by setting different weights for the ranking positions of the first loss average value, the ranking positions of the second loss average value, and the ranking positions of the loss variance value, multiplying them, and then adding and summing.

[0091] It can be understood that in the multi - condition sorting of the sub - data in the first data set, the weighted algorithm is adopted in this embodiment. Different weights are set for different conditions respectively. After performing the weighted summation operation, a new weighted summation value is obtained, which is used as the only condition for sorting the sub - data in the first data set. This unique condition summarizes the situations of the first loss average value, the second loss average value, and the loss variance value. Sorting according to this unique condition can obtain ordered sub - data, which is convenient for finding sub - data with relatively large loss values and relatively large variation ranges of loss values.

[0092] In one embodiment, the server respectively sorts the sub - data in the first data set in ascending order according to the first loss average value, the second loss average value, and the loss variance value to obtain the ranking positions of the first loss average value, the ranking positions of the second loss average value, and the ranking positions of the loss variance value for each sub - data; performs a weighted summation operation on the ranking positions of the first loss average value, the ranking positions of the second loss average value, and the ranking positions of the loss variance value for each sub - data, where the weight of the ranking position of the first loss average value is set to 0.6, the weight of the ranking position of the second loss average value is set to 0.3, and the weight of the ranking position of the loss variance value is set to 0.1; and then sorts the first data set in descending order according to the results of the weighted summation operation, obtaining ordered data, which is convenient for screening and also provides more diverse information for mis - annotation screening.

[0093] In this embodiment, the sub - data of the second data set are initially sorted according to the first loss average value; if there is sub - data with the same first loss average value in the second data set, then for the sub - data with the same first loss average value, a secondary sorting is performed according to the second loss average value corresponding to the sub - data; if there is sub - data with the same second loss average value in the second data set after the secondary sorting, then for the sub - data with the same second loss average value, a further sorting is performed according to the loss variance value corresponding to the sub - data. In this way, the screening of mis - labeled data no longer depends only on the information of the first loss average value, but also adds the second loss average value and the loss variance value, improving the accuracy of mis - labeled data screening.

[0094] In one embodiment, screening mis - labeled data from the first data set according to the sorting result includes: determining, according to the sorting result, the sub - data ranked in the top preset positions; and determining the sub - data ranked in the top preset positions as mis - labeled data.

[0095] Specifically, if the sub - data are sorted in descending order, then the sub - data with larger loss values and larger loss value change ranges will be ranked in the front positions, and the server can determine the sub - data ranked in the top preset positions as mis - labeled data.

[0096] In one embodiment, the server sorts the sub - data of the first data set in descending order according to the first loss average value; for the sub - data with the same first loss average value in the first data set, a secondary descending sorting is performed according to the second loss average value corresponding to the sub - data; if there is sub - data with the same second loss average value in the first data set after the secondary sorting, then for the sub - data with the same second loss average value, a descending sorting is performed according to the loss variance value corresponding to the sub - data. In this way, the sub - data with larger loss values and larger loss value change ranges will be ranked in the front positions, and multiple sub - data in the front positions are determined as mis - labeled data.

[0097] In this embodiment, through sorting, the sub - data with larger loss values and larger loss value change ranges are ranked in the front positions, facilitating the screening of mis - labeled data and improving the screening efficiency of mis - labeled data.

[0098] It should be understood that although the steps in the flowcharts in some embodiments of the present application are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the sequence indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict sequence limit, and these steps can be executed in other sequences. Moreover, at least a part of the steps in the flowchart may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution sequence of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0099] In one embodiment, as Figure 4 shown, a mislabeled data recognition device 400 is provided, including: a data processing module 402, a training statistics module 404, a prediction statistics module 406, and a sorting and screening module 408, where:

[0100] The data processing module 402 is configured to enhance a first data set containing mislabeled data to obtain a second data set.

[0101] The training statistics module 404 is configured to input the second data set into a neural network model with a dropout layer, train the neural network model using the second data set to obtain a trained prediction model including the dropout layer, and obtain the first loss average value of each sub-data according to the first loss value of each sub-data in the second data set at each training epoch during the training process.

[0102] The prediction statistics module 406 is configured to perform multiple predictions on the first data set through the prediction model. At each prediction, a part of the neurons in the prediction model are discarded through the dropout layer, and the prediction model after discarding the neurons is used for prediction; according to the second loss value of each sub-data in the first data set at each prediction, the second loss average value and the loss variance value of each sub-data are obtained.

[0103] The sorting and screening module 408 is configured to sort each sub-data in the first data set according to the first loss average value, the second loss average value, and the loss variance value, and screen out the mislabeled data from the first data set according to the sorting result.

[0104] In one embodiment, the prediction statistics module 406 is further configured to randomly discard a part of the neurons in the prediction model through the dropout layer at each prediction; or, at each prediction, a part of the neurons in the prediction model are discarded through the dropout layer according to a preset dropout rule.

[0105] In one embodiment, the prediction statistics module 406 is further configured to average the second loss values of each sub-data in the first data set in multiple prediction results to obtain the second loss average value of each sub-data; and obtain the loss variance value of each sub-data according to the second loss value and the second loss average value of each sub-data in the first data set in multiple prediction results.

[0106] In one embodiment, as Figure 5 shown, the training statistics module 404 includes: an accuracy monitoring module 404a and a loss value statistics module 404b; where:

[0107] The accuracy monitoring module 404a is configured to, during the process of training the neural network model using the second data set, count the accuracy of the neural network model in each training epoch; compare the accuracy of the current training epoch with the accuracy of the previous training epoch respectively; the previous training epoch is at least one; if the accuracy of the current training epoch is less than the accuracy of the previous training epoch, stop training to obtain a trained prediction model including a dropout layer.

[0108] The loss value statistics module 404b is configured to obtain the first loss average value of each sub-data according to the first loss value of each sub-data in the second data set in each training epoch during the training process.

[0109] In one embodiment, the sorting and screening module 408 is further configured to initially sort each sub-data in the first data set according to the first loss average value; if there are sub-data with the same first loss average value in the first data set, then for the sub-data with the same first loss average value, perform a secondary sorting according to the second loss average value corresponding to the sub-data; if there are sub-data with the same second loss average value in the first data set after the secondary sorting, then for the sub-data with the same second loss average value, perform an advanced sorting according to the loss variance value corresponding to the sub-data.

[0110] In one embodiment, the sorting and screening module 408 is further configured to sort each sub-data in the first data set according to the first loss average value, the second loss average value, and the loss variance value respectively, to obtain the ranking position of the first loss average value, the ranking position of the second loss average value, and the ranking position of the loss variance value of each sub-data; perform a weighted summation operation on the ranking position of the first loss average value, the ranking position of the second loss average value, and the ranking position of the loss variance value of each sub-data in the first data set; and sort according to the weighted summation operation result of each sub-data in the first data set.

[0111] In one embodiment, the sorting and screening module 408 is further configured to determine the sub-data ranked in the top preset positions according to the sorting result; and determine the sub-data ranked in the top preset positions as mislabeled data.

[0112] The above mislabeled data recognition device enhances the first data set containing mislabeled data to obtain a second data set. By enhancing the first data set, the generalization and robustness of the training model can be improved, and the difference in loss values between correct data and mislabeled data can be amplified. The second data set is input into a neural network model with a dropout layer. During training, the neural network model reduces the complex co-adaptation relationship between neurons through the dropout layer. The neural network model is trained using the second data set to obtain a trained prediction model including the dropout layer. Based on the first loss value of each sub-data in the second data set at each training epoch during the training process, the first loss average value of each sub-data is obtained. The first data set is predicted multiple times by the prediction model. Each time a prediction is made, a part of the neurons in the prediction model are discarded through the dropout layer, and the prediction model after discarding the neurons is used for prediction. Based on the second loss value of each sub-data in the first data set each time a prediction is made, the second loss average value and the loss variance value of each sub-data are obtained. According to the first loss average value, the second loss average value, and the loss variance value, the sub-data in the first data set are sorted, and mislabeled data is screened from the first data set according to the sorting result. The screening conditions are diverse, avoiding misjudgment due to being limited by a single condition. When training the model, the dropout layer is enabled to prevent overfitting to mislabeled data, and the first data set is enhanced, thereby amplifying the difference in loss values between correctly and wrongly labeled data, and thus improving the accuracy of mislabeled data screening.

[0113] For the specific limitations of the above mislabeled data recognition device, reference can be made to the limitations of the above mislabeled data recognition method in the previous text, which will not be elaborated here. Each module in the above mislabeled data recognition device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.

[0114] In one embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a punctuation mark annotation method.

[0115] Those skilled in the art can understand that Figure 6 The structure shown in Figure 6 is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0116] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0117] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0118] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application may include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0119] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0120] The above embodiments merely illustrate several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several variations and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A mislabeled data recognition method, applied to a server, characterized in that, the server includes a processor for providing computing and control capabilities, a network interface for communicating with external terminals through a network connection, and an internal memory for providing an environment for the operation of computer programs. The method includes: the network interface receives a first data set containing mislabeled data sent by a terminal. The processor enhances the first data set in an offline enhancement or online enhancement manner according to the data volume size of the first data set to obtain a second data set. The enhancement includes at least one of rotation, cropping, translation transformation, and noise perturbation. Among them, the first data set includes a data set collected from the Internet. The offline enhancement method includes: when the data volume of the first data set is less than or equal to a preset quantity threshold, enhancing each data in the first data set as a whole before starting training. The online enhancement method includes: when the data volume of the first data set is greater than the preset quantity threshold, the processor divides the first data set into multiple portions of data and enhances each portion of data separately during each training epoch in the training process of the prediction model; the processor inputs the second data set into a neural network model with a dropout layer, trains the neural network model using the second data set to obtain a trained prediction model including the dropout layer, and obtains the first loss average value of each sub-data according to the first loss value of each sub-data in the second data set during each training epoch in the training process; the processor makes multiple predictions based on the first data set through the prediction model. During each prediction, the processor discards some neurons in the prediction model through the dropout layer and uses the prediction model after discarding the neurons for prediction; the processor obtains the second loss average value and the loss variance value of each sub-data according to the second loss value of each sub-data in the first data set during each prediction; the processor sorts each sub-data in the first data set according to the first loss average value, the second loss average value, and the loss variance value, and screens out mislabeled data from the first data set according to the sorting result; the network interface sends the mislabeled data to the terminal.

2. The method according to claim 1, characterized in that, during each prediction, discarding some neurons in the prediction model through the dropout layer includes: during each prediction, randomly discarding some neurons in the prediction model through the dropout layer; or, during each prediction, discarding some neurons in the prediction model through the dropout layer according to a preset dropout rule.

3. The method according to claim 1, characterized in that, obtaining the second loss average value and the loss variance value of each sub-data according to the second loss value of each sub-data in the first data set during each prediction includes: averaging the second loss values in the multiple prediction results of each sub-data in the first data set to obtain the second loss average value of each sub-data; Obtain the loss variance value of each sub - data according to the second loss value of each sub - data in the multiple prediction results of the first data set and the second average loss of the sub - data.

4. The method according to claim 1, wherein, the training of the neural network model using the second data set to obtain a trained prediction model including the dropout layer includes: During the process of training the neural network model using the second data set, count the accuracy of the neural network model in each training epoch; Compare the accuracy of the current training epoch with the accuracy of the prior training epochs respectively; the prior training epochs are at least one; If the accuracy of the current training epoch is less than the accuracy of the prior training epochs, stop training to obtain a trained prediction model including the dropout layer.

5. The method according to claim 1, wherein, the sorting of each sub - data in the first data set according to the first average loss, the second average loss and the loss variance value includes: Perform an initial sorting of each sub - data in the first data set according to the first average loss; If there are sub - data with the same first average loss in the first data set, then for the sub - data with the same first average loss, perform a secondary sorting according to the second average loss corresponding to the sub - data; If there are sub - data with the same second average loss in the first data set after the secondary sorting, then for the sub - data with the same second average loss, perform an advanced sorting according to the loss variance value corresponding to the sub - data.

6. The method according to claim 1, wherein, the sorting of each sub - data in the first data set according to the first average loss, the second average loss and the loss variance value includes: Sort each sub - data in the first data set according to the first average loss, the second average loss and the loss variance value respectively to obtain the ranking position of the first average loss, the ranking position of the second average loss and the ranking position of the loss variance value for each sub - data; Perform a weighted summation operation on the ranking position of the first average loss, the ranking position of the second average loss and the ranking position of the loss variance value for each sub - data in the first data set; Sort according to the weighted summation operation results of each sub - data in the first data set.

7. The method according to any one of claims 1 to 6, wherein, the screening of mis - labeled data from the first data set according to the sorting result includes: According to the sorting result, determine the sub - data ranked in the top preset positions; Determine the sub - data ranked in the top preset positions as mis - labeled data.

8. A mis - labeled data identification device, applied to a server, wherein, the server includes a processor for providing computing and control capabilities, a network interface for communicating with external terminals through a network connection, and an internal memory for providing an environment for the operation of computer programs. The device includes: A data processing module, configured to receive a first data set containing mislabeled data sent by a terminal, and enhance the first data set by means of offline enhancement or online enhancement according to the data volume of the first data set, so as to obtain a second data set; the enhancement includes at least one of rotation, cropping, translation transformation, and noise perturbation; wherein, the first data set includes a data set collected from the Internet; the offline enhancement method includes: when the data volume of the first data set is less than or equal to a preset quantity threshold, globally enhancing each data in the first data set before starting training; the online enhancement method includes: when the data volume of the first data set is greater than the preset quantity threshold, the processor divides the first data set into multiple portions of data, and enhances each portion of data separately in each training epoch during the training process of the prediction model; A training statistics module, configured to input the second data set into a neural network model with a dropout layer, train the neural network model by using the second data set to obtain a trained prediction model including the dropout layer, and obtain the first loss average value of each sub-data according to the first loss value of each sub-data in the second data set in each training epoch during the training process; A prediction statistics module, configured to perform multiple predictions on the first data set through the prediction model. In each prediction, a part of neurons in the prediction model are discarded through the dropout layer, and the prediction model after discarding the neurons is used for prediction; according to the second loss value of each sub-data in the first data set in each prediction, obtain the second loss average value and the loss variance value of each sub-data; A sorting and screening module, configured to sort each sub-data in the first data set according to the first loss average value, the second loss average value, and the loss variance value, screen mislabeled data from the first data set according to the sorting result; and send the mislabeled data to the terminal.

9. A computer device, including a memory and a processor, where the memory stores a computer program, characterized in that, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • A method and apparatus for generating a model

    CN109376267A

  • Model training method device, equipment and storage medium

    CN111428008A