Data processing method and device, electronic equipment, storage medium and program product
By analyzing the combined probability distribution matrix of the number of forgotten times and confidence of the sample data, and identifying and eliminating the noise sample data, the problem of difficulty in accurately identifying noise sample data in the prior art is solved, and the quality and prediction performance of the model training data are improved.
Patent Information
- Application Number
- CN202510119072.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to accurately identify and eliminate noise sample data, resulting in model training errors and affect prediction accuracy and generalization capabilities.
By analyzing the number of forgotten times of sample data during model training, and determining the candidate noise sample data based on the confidence combined probability distribution matrix, the sample data whose prediction labels are inconsistent with the annotation labels are eliminated.
The recognition efficiency and accuracy of noise sample data are improved, and the quality of model training data is ensured, thereby improving the prediction accuracy and generalization ability of the model.
Smart Images

Figure CN120011815A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to data processing technology, and in particular to a data processing method, device, electronic device, storage medium and program product. Background Art
[0002] In the process of machine learning model training, accurate sample data can ensure that the model learns the correct features and patterns during the training process, improving the accuracy and generalization ability of the model prediction. Since these sample data come from various channels, it is inevitable to encounter some incorrectly labeled sample data, that is, noisy sample data. If noisy sample data is used for model training or parameter fine-tuning, it may cause the model to learn incorrect knowledge, resulting in incorrect and unexpected outputs.
[0003] In the prior art, noise samples are usually identified based on business understanding and manual rules, as well as the K-nearest neighbor algorithm. However, the method of identifying noise samples based on business understanding and manual rules has weak generalization ability, mainly relies on manual experience, and it is difficult to summarize general filtering rules, which is not universal. The method of identifying noise samples based on the K-nearest neighbor algorithm has a large amount of calculation and is easily affected by noise interference and dimensionality disaster, resulting in inaccurate classification.
[0004] Therefore, how to improve the recognition efficiency and accuracy of noise sample data has become a technical problem that needs to be solved urgently. Summary of the invention
[0005] The present application provides a data processing method, device, electronic device, storage medium and program product to solve the problem that it is difficult to accurately remove noise sample data from sample data in existing methods.
[0006] In one aspect, the present application provides a data processing method, the method comprising:
[0007] Obtaining statistical data of sample data in a sample data set; the statistical data comprising: the number of forgetting times when the target model is trained using the sample data;
[0008] Based on the statistical data of the sample data, N candidate sample data are determined; the value of N is related to the confidence joint probability distribution matrix corresponding to the sample data set;
[0009] The first sample data whose predicted label is inconsistent with the marked label among the N candidate sample data is removed from the sample data set to remove the noise sample data.
[0010] Optionally, the determining N candidate sample data based on the statistical data of the sample data comprises:
[0011] Sorting the sample data in descending order of the number of forgetting times of the sample data;
[0012] The first N sample data in the sorting are used as the candidate sample data.
[0013] Optionally, the statistical data further includes: a predicted probability of the sample data corresponding to each annotated label; and the method further includes:
[0014] Based on the predicted probability of each labeled label corresponding to the sample data, obtaining a confidence joint probability distribution matrix corresponding to the sample data set;
[0015] The value of N is determined based on the confidence joint probability distribution matrix and the number of samples in the sample data set.
[0016] Optionally, determining the value of N based on the confidence joint probability distribution matrix and the number of samples in the sample data set includes:
[0017] Obtaining the prediction error probability corresponding to the sample data set according to the sum of the error probability elements corresponding to each labeled label in the confidence joint probability distribution matrix;
[0018] The value of N is determined according to the product of the prediction error probability and the number of samples.
[0019] Optionally, obtaining a confidence joint probability distribution matrix corresponding to the sample data set based on the predicted probability of each annotated label corresponding to the sample data set includes:
[0020] Based on the predicted probability of each labeled label corresponding to the sample data, obtaining a confidence joint count matrix corresponding to the sample data set;
[0021] Based on the confidence joint count matrix, the confidence joint probability distribution matrix is obtained.
[0022] Optionally, obtaining a confidence joint count matrix corresponding to the sample data set based on the predicted probability of each annotated label corresponding to the sample data set includes:
[0023] Based on the predicted probability of each labeled label corresponding to the sample data, obtain the predicted label corresponding to the sample data and the probability corresponding to each predicted label;
[0024] Based on the predicted label corresponding to the sample data and the probability corresponding to each predicted label, obtaining a confidence threshold corresponding to each predicted label;
[0025] Based on the confidence threshold corresponding to each predicted label, the predicted label corresponding to the sample data, and the probability corresponding to each predicted label, a confidence joint count matrix corresponding to the sample data set is obtained.
[0026] Optionally, the acquiring, based on the predicted label corresponding to the sample data and the probability corresponding to each predicted label, a confidence threshold corresponding to each predicted label includes:
[0027] Based on the average value of the probability corresponding to each predicted label corresponding to the sample data, a confidence threshold corresponding to each predicted label is determined.
[0028] Optionally, acquiring a confidence joint count matrix corresponding to the sample data set based on a confidence threshold corresponding to each predicted label, a predicted label corresponding to the sample data, and a probability corresponding to each predicted label includes:
[0029] Based on the confidence threshold corresponding to each predicted label and the probability corresponding to each predicted label, obtain the number of sample data with correct predictions and the number of sample data with incorrect predictions corresponding to each predicted label;
[0030] Based on the number of correctly predicted sample data corresponding to each predicted label and the number of incorrectly predicted sample data, a confidence joint counting matrix corresponding to the sample data set is constructed.
[0031] Optionally, before acquiring the confidence joint probability distribution matrix based on the confidence joint count matrix, the method further includes:
[0032] The confidence joint count matrix is modified based on the number of sample data corresponding to each annotated label in the sample data set.
[0033] Optionally, the modifying the confidence joint counting matrix based on the number of sample data corresponding to each annotated label in the sample data set includes:
[0034] According to the number of sample data with correct predicted labels corresponding to each labeled label and the number of sample data with incorrect predicted labels, obtain the proportion of sample data with correct predicted labels corresponding to each labeled label and the proportion of sample data with incorrect predicted labels;
[0035] Obtaining a corrected number of sample data for which each labeled label corresponds to a correct predicted label according to the proportion of sample data for which each labeled label corresponds to a correct predicted label, and the number of sample data for which each labeled label corresponds to a correct predicted label in the sample data set;
[0036] Obtaining the corrected number of sample data with incorrect predicted labels corresponding to each labeled label according to the proportion of sample data with incorrect predicted labels corresponding to each labeled label and the number of sample data corresponding to each labeled label in the sample data set;
[0037] The confidence joint counting matrix is updated based on the correct number of sample data with correct predicted labels corresponding to each labeled label and the correct number of sample data with incorrect predicted labels.
[0038] Optionally, the method further comprises:
[0039] The second sample data among the last M sample data in the sorting is removed from the sample data set; the second sample data is sample data whose predicted label is consistent with the marked label.
[0040] Optionally, the forgetting times of the last M sample data are all less than or equal to a preset forgetting times.
[0041] Optionally, the obtaining statistical data of the sample data in the sample data set includes:
[0042] Using the sample data set, training the target model in a cross-validation manner;
[0043] Based on the training process data of the sample data, statistical data of the sample data are obtained.
[0044] Optionally, the using the sample data set to train the target model in a cross-validation manner includes:
[0045] Dividing the sample data set to obtain K equal sample data subsets;
[0046] The target model is trained using K equal subsets of sample data.
[0047] Optionally, obtaining statistical data of the sample data based on the training process data of the sample data includes:
[0048] Based on the prediction result corresponding to when the sample data is used as the prediction sample data, obtaining the prediction probability and the prediction label of the sample data;
[0049] Based on the prediction result corresponding to when the sample data is used as training sample data, the number of forgetting times of the sample data is obtained.
[0050] Optionally, the method further comprises:
[0051] The target model is retrained using the sample data set from which the noise sample data and the second sample data are removed.
[0052] In a second aspect, the present application provides a data processing device, the device comprising:
[0053] An acquisition module is used to acquire statistical data of sample data in a sample data set; the statistical data includes: the number of forgetting times when the target model is trained using the sample data;
[0054] A processing module is used to determine N candidate sample data based on the statistical data of the sample data; remove the first sample data whose predicted label is inconsistent with the marked label among the N candidate sample data from the sample data set to remove noise sample data; the value of N is related to the confidence joint probability distribution matrix corresponding to the sample data set.
[0055] In a third aspect, the present application provides an electronic device, the electronic device comprising: a processor, and a memory communicatively connected to the processor;
[0056] The memory stores computer-executable instructions;
[0057] The processor executes the computer-executable instructions stored in the memory to implement the method as described in any one of the first aspects.
[0058] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method as described in any one of the first aspects.
[0059] In a fifth aspect, the present application provides a computer program product, including a computer program, which implements any method in the first aspect when executed by a processor.
[0060] Sample data is forgotten during model training, which means that the sample data is correctly learned by the model in the Tth iteration, but is incorrectly learned by the model in the T+1th iteration. The distribution of the number of times normal sample data is forgotten during model training is quite different from that of the number of times noise sample data is forgotten. Most normal sample data are no longer forgotten after completing one learning training, and a small number of normal sample data are no longer forgotten after being forgotten several times. However, the expected number of times noise sample data is forgotten is relatively large. Therefore, the more sample data is forgotten, the more likely it is to be noise sample data.
[0061] Therefore, the data processing method, device, electronic device, storage medium and program product provided by the present application can accurately identify noise sample data by analyzing the number of forgetting times of sample data. On this basis, the sample data set is cleaned, and the sample data that causes errors in the training of the target model is eliminated, so as to improve the accuracy and precision of the sample data set. Then, when the target model is subsequently trained using the noise-removed sample data, the target model can be prevented from learning the noise sample data, thereby improving the accuracy of the target model, that is, improving the quality of model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0063] Figure 1 This is a schematic diagram of the forgetting times distribution of normal sample data and noise sample data;
[0064] Figure 2 A flowchart of a data processing method provided in an embodiment of the present application;
[0065] Figure 3 A schematic diagram of a flow chart of a method for obtaining statistical data of sample data by cross-validation provided in an embodiment of the present application;
[0066] Figure 4 A schematic diagram of a method for training a target model using K equal portions of sample data subsets provided in an embodiment of the present application;
[0067] Figure 5 A flowchart of a method for cleaning a sample data set based on statistical data of the sample data provided in an embodiment of the present application;
[0068] Figure 6 A flowchart of a data processing method provided in an embodiment of the present application;
[0069] Figure 7 A schematic diagram of a target model training result after removing noise sample data provided in an embodiment of the present application;
[0070] Figure 8 A schematic diagram of a model training result after removing simple sample data provided in an embodiment of the present application;
[0071] Fig. 9 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application;
[0072] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0073] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0074] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0075] In the process of machine learning model training, the accuracy of sample data is crucial. As the input basis of the model, sample data directly affects the learning effect and optimization direction of the model. Accurate sample data can ensure that the model learns the correct features and patterns, thereby improving the accuracy and generalization ability of the model prediction.
[0076] However, since these sample data come from various channels, it is inevitable to encounter some incorrectly labeled sample data, that is, noisy sample data. If noisy sample data is used for model training or parameter fine-tuning, the model may learn incorrect knowledge, resulting in incorrect and unexpected outputs.
[0077] In the prior art, noise sample data is usually identified based on business understanding and manual rules, as well as the K nearest neighbor algorithm. The identification of noise sample data based on business understanding and manual rules relies on experienced data analysts to discover the rules of noise sample data through data analysis and formulate filtering rules to remove noise sample data. However, it is difficult to summarize general filtering rules through business understanding and manual rules for artificially introduced mislabeled sample data and randomly set biased noise data samples. Therefore, this method has weak generalization ability and is not universal.
[0078] Based on the K-nearest neighbor algorithm to identify noise sample data, it is usually assumed that similar samples have the same classification label. The suspected noise sample data is classified through the sample similarity model and the K-nearest neighbor algorithm to identify the noise sample data. This method has a large amount of calculation. When predicting the classification of each sample data set, it is necessary to re-perform global operations. For sample data sets with large sample data capacity, the amount of calculation is huge and the model training efficiency is low. It is also easily disturbed by noise sample data. If the selected similar samples contain noise sample data, it is easy to be disturbed by the noise sample data, resulting in inaccurate classification. In addition, this method needs to consider the influence of all dimensions when calculating the similarity between two samples in a high-dimensional feature space, making the calculation of the distance between features very complicated. As the feature dimension increases, the complexity of calculating the distance between features by this method will also increase significantly. Therefore, this method is also prone to dimensionality disaster.
[0079] Therefore, both of the above two methods have the problem of being difficult to accurately remove noise sample data from sample data.
[0080] The inventors found that in the model training iteration, the model has the phenomenon of "learning and forgetting, forgetting to learn". Specifically, a certain sample data is correctly learned and accurately predicted by the model in the Tth iteration. However, in the T+1 iteration, the feature learning of the sample data is reversed, resulting in an incorrect prediction.
[0081] Taking the image processing model as an example, the inventor selected 10,000 images as a sample data set, randomly reversed the true labels of 5% of the normal sample data, and used them as noise sample data to conduct the following experiment:
[0082] The above sample data set is used to train the model, and finally the forgetting frequency distribution of normal sample data and noise sample data is obtained. Figure 1 This is a schematic diagram of the forgetting times distribution of normal sample data and noise sample data. Figure 1 As shown in the distribution diagram, the X-axis represents the number of forgetting times, and the Y-axis represents the proportion of sample data. Figure 1 The samples shown are also called sample data, and the embodiments of the present application are all described with reference to the sample data.
[0083] like Figure 1 As shown in the figure, the distribution of the number of forgetting times of normal sample data is quite different from that of the number of forgetting times of noise sample data. The distribution of normal samples is a long-tail distribution. Most samples are no longer forgotten after completing one learning training, and a small number of samples are no longer forgotten after forgetting 2-5 times. The distribution of the number of forgetting times of noise sample data is basically irregular, and the overall expected value of the number of forgetting times is too large.
[0084] Therefore, based on the conclusion of the above experiment, it can be inferred that the greater the number of forgetting times, the more likely it is to be noise sample data. In other words, if the sample data frequently shows the phenomenon of "learning and forgetting, forgetting to learn", it indicates that the sample data is more likely to be noise sample data.
[0085] In view of this, the present application provides a data processing method, which determines the candidate sample data suspected to be noise sample data in the sample data set based on the number of times the sample data in the sample data set is forgotten during the iterative training of the target model, and then determines the candidate sample data suspected to be noise sample data in the sample data set based on the confidence joint probability distribution matrix corresponding to the sample data set. And remove the noise sample data in the candidate sample data. This method identifies the sample data that causes errors in the training of the target model by analyzing the number of times the sample is forgotten. On this basis, the sample data set is cleaned to remove the sample data that causes errors in the training of the target model. Thereby improving the accuracy and precision of the sample data set, and thus improving the quality of model training.
[0086] The data processing method provided in the present application can be applied to a model training platform, or a model training platform for cleaning training data, or an electronic device, etc. The following is an explanation using a model training platform as an example.
[0087] In one embodiment, the model training platform can be deployed entirely in a cloud environment. For example, if the computing resources included in the cloud environment are servers running virtual machines, the model training platform can be independently deployed on a server or virtual machine in the cloud environment, or the model training platform can be distributedly deployed on multiple servers in the cloud environment, or distributedly deployed on multiple virtual machines in the cloud environment, or distributedly deployed on servers and virtual machines in the cloud environment.
[0088] The model training platform provided in the embodiment of the present application can also be deployed in different environments in a distributed manner. The model training platform provided in the present application can be logically divided into multiple parts, each of which has different functions. The various parts in the model training platform can be deployed in any two or three of the terminal computing device (located on the user side), the edge environment, and the cloud environment. The various parts of the model training platform deployed in different environments or devices collaborate to realize the function of data processing. It should be understood that the embodiment of the present application does not restrictively divide which parts of the model training platform are deployed in what environment. In actual application, it can be adaptively deployed according to the computing power of the terminal computing device, the resource occupancy of the edge environment and the cloud environment, or the specific application requirements.
[0089] The model training platform can also be deployed separately on a computing device in any environment (for example, deployed separately on an edge server in an edge environment).
[0090] The technical solutions of the embodiments of the present application are described in detail below in conjunction with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0091] Figure 2 A flow chart of a data processing method provided in an embodiment of the present application. Figure 2 As shown, the method comprises the following steps:
[0092] S201, obtaining statistical data of sample data in a sample data set; the statistical data includes: the number of forgetting times when the sample data is used to train a target model.
[0093] The sample data set is a data set used in the target model training, and the sample data set includes multiple sample data and the annotation labels corresponding to the sample data. The statistical data is the data obtained based on the prediction results during the target model training process. The annotation labels corresponding to the sample data are the labels of the sample data categories.
[0094] The above statistical data can be the data recorded by the model training platform during the training process of the target model, or it can be the data collected based on the training results of each iteration cycle after the training is completed, or it can be the data collected based on the training results after the training results of each iteration cycle are synchronized to other devices.
[0095] The number of forgetting times included in the statistical data refers to the total number of forgetting times during the training process of the target model. For example, when a sample data is correctly learned and accurately predicted by the model in the Tth iteration, but is incorrectly learned and incorrectly predicted by the model in the T+1th iteration, it is recorded as one forgetting of the sample data.
[0096] It should be understood that the embodiments of the present application do not limit whether the statistical data also includes other content.
[0097] Optionally, in some embodiments, the above operation of training the target model using the sample data set may also be performed by other model training platforms or systems or devices, which is not limited in this application.
[0098] This embodiment does not limit the above-mentioned method of using the sample data set to train the target model. For example, the sample data set can be used to train the target model in a cross-validation manner.
[0099] The cross-validation method is a method for the generalization ability of the target model. By dividing the sample data set into multiple subsets, each subset has the opportunity to be used as a training set and a validation set, thereby reducing the risk of overfitting and improving the accuracy of the target model performance estimation. The cross-validation method can be, for example, any cross-validation method such as K-fold cross-validation, simple cross-validation, and leave-one-out method, which is not limited in the embodiments of the present application. The above K refers to the number of subsets into which the sample data set is divided.
[0100] S202. Determine N candidate sample data based on the statistical data of the sample data; the value of N is related to the confidence joint probability distribution matrix corresponding to the sample data set.
[0101] The candidate sample data refers to sample data that is screened out based on statistical data and is suspected to be noise sample data.
[0102] The confidence joint probability distribution matrix is used to show the joint probability distribution between the predicted label and the annotated label of the sample data. This joint distribution reflects the distribution between the noise sample data and the normal sample data in the sample data set. The elements in the matrix represent the confidence probability of the predicted label corresponding to the annotated label of the sample data. Therefore, through the confidence joint probability distribution matrix, the sample data suspected of being noise samples in the sample data set can be accurately judged, and the sample data set can be cleaned on this basis to ensure the accuracy of the sample data set.
[0103] Exemplarily, the model training platform can obtain the confidence joint probability distribution matrix corresponding to the sample data set based on the predicted probability of each labeled label corresponding to the sample data, and then determine the value of N based on the confidence joint probability distribution matrix and the number of samples in the sample data set.
[0104] For example, N sample data in the sample data set may be randomly selected as candidate sample data.
[0105] For another example, the sample data may be sorted in descending order of the number of forgetting times of the sample data, and the first N sample data in the sorting are used as candidate sample data.
[0106] S203: Remove the first sample data whose predicted label is inconsistent with the labeled label among the N candidate sample data from the sample data set to remove noise sample data.
[0107] In this implementation, the above statistical data may carry a predicted label, or may carry content that can calculate or derive the predicted label, such as the predicted probability of each labeled label. The predicted probability of the sample data includes the predicted probability corresponding to each labeled label. This probability is used to determine the final predicted label of the sample data. The predicted label may be the labeled label with the highest predicted probability among the predicted probabilities corresponding to each labeled label.
[0108] Among them, the sample data whose predicted labels are consistent with the labeled labels in the first N candidate sample data may be sample data that is difficult to train, while the sample data whose predicted labels are inconsistent with the labeled labels are usually noise sample data. Therefore, by predicting labels and labeling labels, the noise sample data can be accurately screened out and removed from the sample data with a high number of forgetting times, while the sample data that is difficult to train continues to be retained in the sample data set. Based on the above method, not only can the accuracy of data cleaning be improved, but also the accuracy of subsequent model training can be guaranteed.
[0109] In addition to noise sample data, there can also be simple sample data in the sample data set. Many sample data have the same type of dominant features. After completing the first or first few iterations of training, these sample data will not be forgotten. This type of sample data is simple sample data. Simple sample data is usually large in number, and sample data of the same type has no benefit for the training of the target model. In the process of training the target model, removing simple sample data from the sample data set can improve the efficiency of target model training without reducing the accuracy of the target model.
[0110] After the simple sample data completes the first or first few iterations of training, forgetting training will not occur. Therefore, in some embodiments, the simple sample data can be further removed from the sample data set in combination with the number of forgetting times to further reduce the dimension of the data set and remove sample data that has no gain effect on the target model. While ensuring the accuracy of the target model, the target model training cycle can be greatly shortened and the efficiency of target model training can be improved.
[0111] For example, the first forgetting times threshold can be used to screen the simple sample data to be removed. That is, when the forgetting times of the sample data are less than or equal to the first forgetting times threshold, it means that the sample data is simple sample data in the target model training process, and the sample data is removed from the sample data set.
[0112] For another example, the sample data may be sorted in descending order of the number of forgetting times of the sample data. M sample data in the sorting are removed from the sample data set, or the second sample data among the last M sample data are removed from the sample data set. The second sample data is the sample data whose predicted label is consistent with the marked label.
[0113] It should be understood that the first forgetting times threshold may be a preset forgetting times threshold, or may be a forgetting times threshold determined based on the forgetting times of each sample data in the sample data set. For example, the forgetting times threshold may be obtained by performing statistical analysis on the forgetting times of the sample data set and calculating any statistical value such as the median or average of the forgetting times.
[0114] The above M value can be directly given a fixed value by comprehensively considering the size of the sample data set, or it can be dynamically adjusted by combining the training requirements of the target model.
[0115] Optionally, in some embodiments, the method of the embodiments of the present application can be used to clean only simple sample data during the data cleaning process, or other sample data that requires the number of forgetting times as a judgment condition can be cleaned, etc., without limitation.
[0116] The method provided in the present application is based on the number of times sample data in a sample data set is forgotten during the iterative training of a target model, and the confidence joint probability distribution matrix corresponding to the sample data set, to determine candidate sample data suspected to be noise sample data, and to remove the noise sample data from the candidate sample data. The cleaning of the sample data set is achieved. The method identifies the sample data that causes errors in the training of the target model by analyzing the number of times the sample is forgotten. On this basis, the sample data set is cleaned to remove the sample data that causes errors in the training of the target model, thereby improving the accuracy of the sample data set and further improving the quality of model training.
[0117] The following example uses the model training platform to train the target model using the K-fold cross-validation method and obtain statistical data.
[0118] Figure 3 A flow chart of a method for obtaining sample data statistical data by cross-validation is provided in an embodiment of the present application. Figure 3 As shown, the method comprises the following steps:
[0119] S301. Divide the sample data set to obtain K equal sample data subsets.
[0120] Each sample data subset is part of the sample data in the sample data set.
[0121] Exemplarily, the sample data set is evenly divided into K sample data subsets of equal size, or as equal and exclusive as possible, to ensure that each subset can represent the characteristics of the entire sample data set.
[0122] The size of K can be determined by comprehensively considering the size of the sample data set, and the embodiment of the present application does not limit the size of K.
[0123] S302: Use K equal sample data subsets to train the target model.
[0124] For example, Figure 4 A schematic diagram of a method for training a target model using K equal portions of sample data subsets provided in an embodiment of the present application. Figure 4 As shown in the figure, K rounds of cross validation will be performed during the training of the target model. When performing K rounds of cross validation, each cross validation uses a sample data subset as a validation set, and the remaining K-1 sample data subsets as training sets to iteratively train the target model. In each round of cross validation, the target model will undergo multiple iterations of training. In each iteration, the training set will be used to train the target model first, and then the validation set will be used to verify the performance of the target model. This continues until the iterative training of this cross validation is completed.
[0125] S303: Obtain statistical data based on the prediction results corresponding to the sample data when the sample data is used as the prediction sample data.
[0126] The above-mentioned prediction sample data is the sample data for result prediction after the target model is trained. In the embodiment of the present application, the sample data is a subset of the sample data in the training set. The statistical data may include, for example, the prediction probability, prediction label and number of forgetting of the sample data.
[0127] For example, based on the prediction result corresponding to the sample data when the sample data is used as the predicted sample data, the predicted label and predicted probability of the sample data are obtained.
[0128] The predicted probability of the sample data includes the predicted probability corresponding to each annotated label. This probability is used to determine the final predicted label of the sample data. The predicted label can be the label with the highest predicted probability among the predicted probabilities corresponding to each annotated label.
[0129] For another example, based on the corresponding prediction result when the sample data is used as training sample data, the number of forgetting times of the sample data is obtained.
[0130] Exemplarily, in each iteration, the target model predicts the sample data in the training set based on the current training state.
[0131] For example, after the validation set verifies the performance of the target model, when the training set sample data subset is used as the prediction sample data to predict the result, the target model will obtain the prediction result of the sample data of the data subset and record the prediction result of the current iterative training. After the iterative training of this cross-validation is completed, the number of forgetting times of each sample data can be calculated based on these prediction results.
[0132] For example, when a sample data is correctly learned and accurately predicted by the model in the Tth iteration, but is incorrectly learned and incorrectly predicted by the target model in the T-1th iteration, it is recorded as a forgetting of the sample data. During the entire iterative training process, the total number of times the predicted label of the sample data changes is the number of forgetting times obtained by the cross-validation of this round. Since K rounds of cross-validation will be performed, each sample data subset will obtain K-1 forgetting times. The K-1 forgetting times of each sample data subset are added and averaged to obtain the final prediction result of the target model.
[0133] The above method divides the sample data set into K sample data subsets of equal size or as equal as possible, and performs K-1 predictions on each sample subset. The K-1 prediction results of each sample data subset are averaged to obtain the number of forgetting times of the target model during the training process.
[0134] This method can prevent information loss and information bias during the training of the target model. It ensures that the number of forgetting times of each data subset during the training process can be taken into account, providing a more reliable prediction of the target model performance. It ensures that each sample data subset has the opportunity to participate in the prediction of the target model as a set.
[0135] The following example takes the target model as a binary classification model. Assume that the sample data set used to train the target model has 500 sample data, 300 of which are labeled a, and 200 are labeled b. A feasible implementation method is provided. The target model predicts the sample data of the data subset, and the prediction results may include the content shown in Table 1:
[0136] Table 1
[0137]
[0138]
[0139] The method provided in the present application trains the sample data set by cross-validation and obtains the number of forgetting, prediction probability and prediction category of the sample data. During the training process, the method divides the sample data set into K equal sample data subsets. Each iterative training selects a different sample data subset as a training set and a validation set. Each sample data has the opportunity to be used for training and validation, so that the accuracy of the sample data set can be more comprehensively evaluated.
[0140] The following is an example of how to determine N candidate sample data based on the statistical data of the sample data to eliminate the noise sample data. Figure 5A flow chart of a method for cleaning a sample data set based on statistical data of the sample data provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, the method includes:
[0141] S501 , sorting the sample data in descending order of the number of forgetting times of the sample data.
[0142] Exemplarily, the number of forgetting times of the sample data is obtained as described above, and the sample data are sorted in order from high to low according to the number of forgetting times of each sample data in the sample data set during the target model training process.
[0143] S502: Remove the first sample data whose predicted label is inconsistent with the marked label among the first N sample data from the sample data set.
[0144] Among them, the sample data whose predicted labels are consistent with the labeled labels in the first N sample data may be sample data that is difficult to train, while the sample data whose predicted labels are inconsistent with the labeled labels are usually noise sample data. Therefore, by predicting the labels and labeling the labels, the noise sample data can be accurately screened out and removed from the sample data with a high number of forgetting times, while the sample data that is difficult to train continues to be retained in the sample data set. Based on the above method, not only can the accuracy of data cleaning be improved, but also the accuracy of subsequent model training can be guaranteed.
[0145] As shown above, the N value can be directly given a fixed value by comprehensively considering the size of the sample data set, or the N value can be dynamically adjusted by combining the training requirements of the target model.
[0146] Exemplarily, the model training platform can obtain the confidence joint probability distribution matrix corresponding to the sample data set based on the predicted probability of each labeled label corresponding to the sample data, and then determine the value of N based on the confidence joint probability distribution matrix and the number of samples in the sample data set.
[0147] In the above method, we counted the number of times each sample data is forgotten through target model training. Sample data with a high number of forgetting times are suspected to be noise sample data, or they may be normal sample data with unclear features. In order to correctly distinguish these two types of sample data, the confidence joint probability distribution matrix is used to further judge the noise sample data.
[0148] The confidence joint probability distribution matrix is used to show the joint probability distribution between the predicted label and the annotated label of the sample data. This joint distribution reflects the distribution between the noise sample data and the normal sample data in the sample data set. The elements in the matrix represent the confidence probability of the predicted label corresponding to the annotated label of the sample data. Therefore, through the confidence joint probability distribution matrix, the sample data suspected of being noise samples in the sample data set can be accurately judged, and the sample data set can be cleaned on this basis to ensure the accuracy of the sample data set.
[0149] For example, the prediction error probability corresponding to the sample data set can be obtained based on the sum of the error probability elements corresponding to each labeled label in the confidence joint probability distribution matrix. The value of N is determined based on the product of the prediction error probability and the number of samples.
[0150] The sum of the error probability elements corresponding to each of the above-mentioned annotation labels is the sum of the confidence joint probabilities of the sample data whose annotation label i is inconsistent with the predicted label j. That is, the sum of the elements i≠j in the confidence joint probability distribution matrix.
[0151] In the following, i represents the marked label and j represents the predicted label, which will not be described in detail. For example, in a binary classification model, i is a or b, and j is a or b. The following formulas and tables are all based on the binary classification model as an example, which will not be described in detail.
[0152] For example, the value of N can be determined using the following formula (1).
[0153] N=n·∑ i≠j q i,j (1)
[0154] Where n is the number of sample data in the sample data set, ∑ i≠j q i,j is the sum of the elements i≠j in the confidence joint probability distribution matrix.
[0155] An example is given in conjunction with Table 1, as shown in the following formula (2).
[0156] N = 500*(q a,b +q b,a )=500*(0.09+0.11)=100 (2)
[0157] For example, the confidence joint probability distribution matrix can be calculated by the confidence joint count matrix corresponding to the sample data set.
[0158] Exemplarily, for example, the following formulas (3) and (4) may be used to obtain a confidence joint probability distribution matrix.
[0159]
[0160] Among them, Q i,j is the confidence joint probability distribution matrix, q i,j Represents the elements in the confidence joint probability distribution matrix. i,j is the value of each element in the confidence joint count matrix, ∑ i,j∈[m] c i,j is the sum of all elements in the confidence joint count matrix, and [m] is the set of all annotated labels of the sample data in the sample dataset.
[0161] The above confidence joint count matrix is a matrix that shows the quantitative relationship between the target model's predicted labels and the labeled labels. The elements of this matrix represent the number of samples of the predicted labels corresponding to the labeled labels of each sample data. This matrix helps to identify samples suspected to be noise sample data in the sample data set.
[0162] Exemplarily, based on the predicted probability of each annotated label corresponding to the sample data, a confidence joint count matrix corresponding to the sample data set is obtained.
[0163] For example, the following formulas (5), (6) and (7) can be used to construct the confidence joint count matrix corresponding to the sample data set.
[0164]
[0165]
[0166] in, It represents the set of sample data in the sample data set whose predicted label corresponds to a probability greater than the confidence threshold. y represents the label i. Denotes the predicted label. X y=i Indicates that the labeled data set is i samples. represents the probability corresponding to the predicted label, Indicates the sample data whose probability corresponding to the predicted label j is greater than the confidence threshold, t j Represents the confidence threshold. It means that the predicted label j is the label with the highest predicted probability among the predicted probabilities of each label.
[0167] C i,j is the confidence joint count matrix, c i,j is the element in the confidence joint count matrix, It represents the number of sample data with correct prediction labels corresponding to each sample data in the sample data set after the confidence threshold screening, and the number of sample data with wrong prediction labels. The sample data with correct prediction labels corresponding to each sample data is i=j, and the sample data with wrong prediction labels is i≠j.
[0168] Continuing to refer to the above example, combined with Table 1, the confidence joint count matrix is shown in Table 2:
[0169] Table 2
[0170] <![CDATA[c i,j ]]> j=a j=b i=a 100 40 i=b 50 180
[0171] In the embodiments of the present application, the description of each matrix is illustrated in a table, and the content of each cell in the table is equivalent to the corresponding element in the matrix.
[0172] As shown in the table, the number of sample data with labeled i as a and predicted label j as a is 100; the number of sample data with labeled i as b and predicted label j as b is 180; the number of sample data with labeled i as a and predicted label j as b is 40; the number of sample data with labeled i as b and predicted label j as a is 50.
[0173] The above confidence threshold t j It is a preset critical value used to evaluate the credibility of the prediction results of the target model. It is usually associated with the predicted probability output of the target model and determines the acceptance criteria of the target model prediction. In the classification task, when the probability corresponding to the predicted label exceeds this threshold, the predicted label will be adopted as the final decision.
[0174] The embodiment of the present application does not limit the setting method of the confidence threshold. The confidence threshold can be given a fixed value according to the requirements of the target model, or it can be dynamically calculated by using the predicted probability during the training of the target model, thereby improving the accuracy of the confidence threshold.
[0175] The following is an explanation of how to obtain the confidence threshold based on the predicted probability of each labeled label based on sample data.
[0176] First, based on the predicted probability of each labeled label corresponding to the sample data, the predicted label corresponding to the sample data and the probability corresponding to each predicted label can be obtained.
[0177] As shown in Table 1, the predicted probability of label a of sample data 1 is 0.9, and the predicted probability of label b is 0.1. The predicted label of sample data 1 is a, and the probability corresponding to the predicted label is 0.9. The predicted probability of label a of sample data 2 is 0.4, and the predicted probability of label b is 0.6. The predicted label of sample data 2 is b, and the probability corresponding to the predicted label is 0.6. The predicted labels and the corresponding probabilities of the predicted labels are obtained for other sample data in the same way as above, which will not be described in detail.
[0178] Secondly, based on the predicted labels corresponding to the sample data and the probability corresponding to each predicted label, the confidence threshold corresponding to each predicted label is obtained.
[0179] Exemplarily, the confidence threshold corresponding to each predicted label is determined based on the average value of the probabilities corresponding to the predicted labels corresponding to the sample data.
[0180] For example, the confidence threshold can be obtained using the following formula (8).
[0181]
[0182] in, is the number of all sample data in the sample data set whose predicted labels are a or b, p j Represents the probability corresponding to the predicted label of each sample data in the sample data set. Represents the set of sample data in the sample data set whose sample data prediction labels are a or b respectively. It means summing the probability values corresponding to the sample data predicted label being a or b.
[0183] It should be understood that the above formula is only an exemplary way of calculating the confidence threshold, and appropriate deformation can be made on the basis of the above formula, for example, adding a coefficient and / or constant, etc., and there is no limitation on this.
[0184] Continuing with the above example, the probability corresponding to each sample data prediction label in the sample data set can be shown in Table 3:
[0185] Table 3
[0186]
[0187] Continuing with the above example, the sum of the probability values corresponding to the predicted labels of each sample data in the sample data set can be shown in Table 4:
[0188] Table 4
[0189]
[0190] For example, there are 240 sample data with the predicted label a and 260 sample data with the predicted label b in the sample data set. The calculation results of the confidence thresholds corresponding to the predicted labels can be shown in Table 5:
[0191] Table 5
[0192]
[0193] For example, there are 240 sample data with the predicted label a and 260 sample data with the predicted label b in the sample data set. The calculation results of the confidence thresholds corresponding to the predicted labels can be shown in Table 5:
[0194] After obtaining the confidence threshold, Table 5 obtains the confidence joint count matrix corresponding to the sample data set based on the confidence threshold corresponding to each predicted label, the predicted label corresponding to the sample data, and the probability corresponding to each predicted label.
[0195] Exemplarily, based on the confidence threshold corresponding to each predicted label and the probability corresponding to each predicted label, the number of sample data predicted correctly and the number of sample data predicted incorrectly corresponding to each predicted label are obtained. Based on the number of sample data predicted correctly and the number of sample data predicted incorrectly corresponding to each predicted label, a confidence joint count matrix corresponding to the sample data set is constructed.
[0196] The sample data with the correct predicted label is the sample data with the same predicted label as the labeled label. The sample data with the wrong predicted label is the sample data with the wrong predicted label as the sample data.
[0197] Since the number of sample data in the confidence joint counting matrix is affected by the confidence threshold, the number of sample data in the confidence joint counting matrix may be different from the number of sample data in the sample data set.
[0198] When there is incomplete consistency, the count matrix can be used directly for subsequent processing, or the confidence joint count matrix can be corrected based on the number of sample data corresponding to each labeled label in the sample data set, and then the corrected confidence joint count matrix can be used to calculate the confidence joint probability distribution matrix. In other words, when there is incomplete consistency, it can be determined whether the deviation between the two is within the tolerance range. If it is not within the tolerance range, this method can be used for correction.
[0199] Exemplarily, based on the number of sample data with correct predicted labels corresponding to each labeled label and the number of sample data with incorrect predicted labels, the proportion of sample data with correct predicted labels corresponding to each labeled label and the proportion of sample data with incorrect predicted labels are obtained.
[0200] Secondly, according to the proportion of sample data with correct predicted labels corresponding to each labeled label and the number of sample data corresponding to each labeled label in the sample data set, the number of corrected sample data corresponding to the correct predicted labels for each labeled label is obtained. According to the proportion of sample data with incorrect predicted labels corresponding to each labeled label and the number of sample data corresponding to each labeled label in the sample data set, the number of corrected sample data corresponding to the incorrect predicted labels for each labeled label is obtained.
[0201] Finally, the confidence joint count matrix is updated based on the number of corrected sample data with correct predicted labels corresponding to each annotated label, as well as the number of corrected sample data with incorrect predicted labels.
[0202] The above process can be expressed, for example, by the following formulas (9) and (10):
[0203]
[0204] in, is the modified confidence joint count matrix, [m] is the set of all labeled labels of sample data in the sample data set, k is the predicted label, ∑ k∈[m] c i,k is the sum of all sample data labeled a or b in the confidence joint count matrix, c i,j is the element in the confidence joint count matrix. |X y=i | is the number of all sample data with the same label in the sample dataset.
[0205] Continuing with the above example, the modified confidence joint count matrix result can be shown in Table 6:
[0206] Table 6
[0207]
[0208] Table 6 is illustrated in conjunction with Table 2. As shown in Table 6, the number of sample data corrected by labeling i as a and predicting label j as a is 214; the number of sample data corrected by labeling i as b and predicting label j as b is 157; the number of sample data corrected by labeling i as a and predicting label j as b is 86; the number of sample data corrected by labeling i as b and predicting label j as a is 43.
[0209] Optionally, a confidence joint probability distribution matrix may be obtained based on the modified confidence joint count matrix, thereby improving the accuracy of the confidence probability matrix.
[0210] For example, the following formulas (11) and (12) may be used to obtain the confidence joint probability distribution matrix.
[0211]
[0212] in, is the element in the modified confidence joint count matrix, ∑ i,j联合概率 is the sum of all elements in the corrected confidence joint count matrix.
[0213] Continuing with the above example, the confidence joint probability distribution matrix result can be shown in Table 7:
[0214] Table 7
[0215] <![CDATA[q i,j ]]> j=a j=b i=a 0.43=214 / 500 0.11=57 / 500 i=b 0.09=43 / 500 0.31=157 / 500
[0216] Combined with Table 6, Table 7 is illustrated. As shown in Table 7, when the labeled label i is a, the confidence joint probability of predicting the label j is a is 0.43; when the labeled label i is b, the confidence joint probability of predicting the label j is b is 0.31; when the labeled label i is a, the confidence joint probability of predicting the label j is b is 0.11; when the labeled label i is b, the confidence joint probability of predicting the label j is a is 0.09.
[0217] The method provided by the present application sorts the sample data according to the number of forgetting times of each sample data in the sample data set during the target model training process from high to low. The first sample data whose predicted label is inconsistent with the labeled label among the first N sample data is removed from the sample data set.
[0218] By sorting the sample data, the method can preliminarily identify the sample data suspected to be noise sample data in the sample data set. These samples suspected to be noise sample data can be processed preferentially, thereby improving the efficiency and pertinence of data cleaning. The first sample data whose predicted label is inconsistent with the labeled label in the first N sample data is removed from the sample data set. Sample data that contributes little to the training of the target model or causes errors in the training of the target model can be removed. This improves the accuracy of the sample data set and shortens the training cycle of the target model.
[0219] The following is a specific example to illustrate how to implement data processing. Figure 6 A flowchart of a data processing method provided in an embodiment of the present application is shown in FIG. Figure 6 As shown, the method may include the following steps:
[0220] S601. Divide the sample data set to obtain K equal sample data subsets.
[0221] S602: Use K equal portions of sample data subsets to train the target model. Based on the prediction results corresponding to when the sample data is used as the prediction sample data, obtain the statistical results of the sample data.
[0222] The statistical results may include, for example, prediction probability, prediction label, and number of forgetting times.
[0223] S603: Sort the sample data in descending order of the number of forgetting times of the sample data.
[0224] S604: Determine a confidence threshold corresponding to each predicted label based on an average value of probabilities corresponding to the predicted labels corresponding to the sample data.
[0225] S605. Based on the confidence threshold corresponding to each predicted label, the predicted label corresponding to the sample data, and the predicted probability corresponding to each predicted label, obtain a confidence joint count matrix corresponding to the sample data set.
[0226] S606: Calibrate the confidence joint counting matrix based on the number of sample data corresponding to each annotated label in the sample data set.
[0227] S607: Update the confidence joint counting matrix based on the correct number of sample data with correct predicted labels corresponding to each labeled label and the correct number of sample data with incorrect predicted labels.
[0228] S608. Obtain a confidence joint probability distribution matrix based on the calibrated confidence joint count matrix.
[0229] S609: Obtain the prediction error probability corresponding to the sample data set according to the sum of the error probability elements corresponding to each labeled label in the confidence joint probability distribution matrix, and determine the value of N according to the product of the prediction error probability and the number of samples.
[0230] S610: Remove the first sample data whose predicted labels are inconsistent with the marked labels in the first N sample data and the sample data whose predicted labels are consistent with the marked labels in the last M sample data from the sample data set.
[0231] The value of M can be set according to the training situation of the target model.
[0232] S611. Retrain the target model using the cleaned sample data set.
[0233] The method provided in the present application trains the sample data set by adopting a cross-validation method, so that each sample data has the opportunity to be used for training and validation, thereby being able to more comprehensively evaluate the accuracy of the sample data set.
[0234] The sample data are sorted in the order of the number of times each sample data in the sample data set is forgotten during the training of the target model. The confidence joint probability distribution matrix is calculated by confidence learning, and the number of samples to be removed is determined by the confidence joint probability distribution matrix and the number of samples in the sample data set. Based on the number of samples to be removed and the sorting of the sample data, the noise sample data and simple sample data are removed from the data set. This achieves dimensionality reduction of the sample data set, greatly shortening the training cycle of the target model while ensuring the accuracy of the sample data set.
[0235] The cleaned sample data set is used to retrain the target model to ensure the accuracy and stability of the target model.
[0236] The following will compare and analyze the impact on the training results of the target model after removing noise sample data in the sample data set using the method provided by this application and the method provided by the K nearest neighbor algorithm.
[0237] One possible implementation method, taking the image processing model as an example, is to first select 10,000 images as the validation sample data set for training, randomly reverse the true labels of 5% of the sample data, and artificially inject noise. Then select 10,000 images from them as the test sample data set for training.
[0238] Secondly, the method provided in the embodiment of the present application and the K nearest neighbor method in the prior art are used to identify noise sample data from the test sample data set.
[0239] Finally, after completing the noise sample identification, 1% to 10% of the noise sample data are selected for elimination. Then, the validation sample data set is used to train and predict the screened sample data to obtain the training results of the target model after eliminating the noise sample data.
[0240] It can be seen from the above results that by removing the noise sample data through the method of the embodiment of the present application, the accuracy of the target model can be significantly improved while achieving dimensionality reduction of the sample data set and shortening the training cycle of the target model.
[0241] Figure 7 A schematic diagram of a target model training result after removing noise sample data provided by an embodiment of the present application. Figure 7 As shown, the closer the proportion of the filtered noise sample data is to the proportion of the real noise sample data, the higher the accuracy of the target model. And compared with the K nearest neighbor method, the accuracy of the target model can be improved by up to 0.89% by filtering out the noise sample data using the method provided in the embodiment of the present application.
[0242] The following will describe in detail the method provided by this application and the impact on the training results of the target model after removing simple sample data from the sample data set.
[0243] One possible implementation method, taking the image processing model as an example, is to first select 10,000 images as the validation sample data set for training, randomly reverse the true labels of 5% of the sample data, and artificially inject noise. Then select 10,000 images as the test sample data set for training.
[0244] Secondly, the target model is trained on the test sample data set using the method provided in the embodiment of the present application.
[0245] Finally, after the target model training is completed, a portion of samples are randomly removed from the simple sample data with a forgetting time of 0. Then, the validation sample data set is used to train and predict the screened sample data to obtain the training results of the target model after removing the simple sample data.
[0246] Figure 8 A schematic diagram of a target model training result after removing simple sample data provided in an embodiment of the present application. Figure 8 As shown in the figure, removing simple sample data of a certain range of classes has almost no effect on the final accuracy of the target model. When 5% of the simple sample data is removed, the accuracy of the target model drops by 0.001%, when 20% of the simple sample data is removed, the accuracy of the target model drops by 0.008%, and when 40% of the simple sample data is removed, the accuracy of the target model drops by only 0.024%.
[0247] From the above results, it can be seen that the simple sample data is removed from the data set by the method provided in the embodiment of the present application. Although the accuracy of the target model has shown a slightly downward trend, the impact of this decline on the accuracy of the target model is almost negligible. However, when the target model is trained on the sample data set after the simple sample data is removed, the training cycle of the target model can be greatly shortened, thereby improving the training efficiency while keeping the performance of the target model basically stable.
[0248] The above is the method of the embodiment of the present application. The device provided by the embodiment of the present application is described below.
[0249] Fig. 9 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application is shown in FIG. Fig. 9 As shown, the data processing device includes: an acquisition module 901 and a processing module 902.
[0250] The acquisition module 901 is used to acquire statistical data of sample data in the sample data set. The statistical data includes: the number of forgetting times when the sample data is used to train the target model.
[0251] Processing module 902 is used to determine N candidate sample data based on the statistical data of the sample data. The first sample data whose predicted label is inconsistent with the labeled label in the N candidate sample data is removed from the sample data set to remove the noise sample data. The value of N is related to the confidence joint probability distribution matrix corresponding to the sample data set.
[0252] Optionally, the processing module 902 is specifically configured to sort the sample data in descending order of the number of forgetting times of the sample data, and use the first N sample data in the sorting as candidate sample data.
[0253] Optionally, the statistical data also includes: the predicted probability of the sample data corresponding to each labeled label; the processing module 902 is further used to obtain the confidence joint probability distribution matrix corresponding to the sample data set based on the predicted probability of the sample data corresponding to each labeled label. Based on the confidence joint probability distribution matrix and the number of samples in the sample data set, the value of N is determined.
[0254] In a possible implementation, processing module 902 is specifically configured to obtain the prediction error probability corresponding to the sample data set according to the sum of the error probability elements corresponding to each labeled label in the confidence joint probability distribution matrix, and determine the value of N according to the product of the prediction error probability and the number of samples.
[0255] In a possible implementation, the processing module 902 is specifically configured to obtain a confidence joint count matrix corresponding to the sample data set based on the predicted probability of each labeled label corresponding to the sample data set, and obtain a confidence joint probability distribution matrix based on the confidence joint count matrix.
[0256] For example, processing module 902 is specifically used to obtain the predicted label corresponding to the sample data and the probability corresponding to each predicted label based on the predicted probability of each labeled label corresponding to the sample data. Based on the predicted label corresponding to the sample data and the probability corresponding to each predicted label, obtain the confidence threshold corresponding to each predicted label. Based on the confidence threshold corresponding to each predicted label, the predicted label corresponding to the sample data, and the probability corresponding to each predicted label, obtain the confidence joint count matrix corresponding to the sample data set.
[0257] Exemplarily, the processing module 902 is specifically configured to determine a confidence threshold corresponding to each predicted label based on an average value of probabilities corresponding to each predicted label corresponding to the sample data.
[0258] Exemplarily, the processing module 902 is specifically used to obtain the number of sample data with correct predictions corresponding to each prediction label and the number of sample data with incorrect predictions based on the confidence threshold corresponding to each prediction label and the probability corresponding to each prediction label. Based on the number of sample data with correct predictions corresponding to each prediction label and the number of sample data with incorrect predictions, a confidence joint count matrix corresponding to the sample data set is constructed.
[0259] In a possible implementation, the processing module 902 is further used to correct the confidence joint count matrix based on the number of sample data corresponding to each labeled label in the sample data set before obtaining the confidence joint probability distribution matrix based on the confidence joint count matrix.
[0260] For example, processing module 902 is specifically used to obtain the proportion of sample data with correct prediction labels corresponding to each labeled label and the proportion of sample data with incorrect prediction labels according to the number of sample data with correct prediction labels corresponding to each labeled label and the number of sample data with incorrect prediction labels. According to the proportion of sample data with correct prediction labels corresponding to each labeled label and the number of sample data corresponding to each labeled label in the sample data set, obtain the correct number of sample data with correct prediction labels corresponding to each labeled label. According to the proportion of sample data with incorrect prediction labels corresponding to each labeled label and the number of sample data corresponding to each labeled label in the sample data set, obtain the correct number of sample data with incorrect prediction labels corresponding to each labeled label. Based on the correct number of sample data with correct prediction labels corresponding to each labeled label and the correct number of sample data with incorrect prediction labels, update the confidence joint counting matrix.
[0261] Optionally, the processing module 902 is specifically configured to remove the second sample data among the last M sample data in the sorting from the sample data set. The second sample data is sample data whose predicted label is consistent with the marked label.
[0262] For example, the number of forgetting times of the last M sample data is less than or equal to the preset number of forgetting times.
[0263] Optionally, the acquisition module 901 is specifically configured to use the sample data set to train the target model in a cross-validation manner, and obtain statistical data of the sample data based on the training process data of the sample data.
[0264] In a possible implementation, the acquisition module 901 is specifically configured to divide the sample data set into K equal sample data subsets, and use the K equal sample data subsets to train the target model.
[0265] In a possible implementation, the acquisition module 901 is further configured to acquire the prediction probability and prediction label of the sample data based on the prediction result corresponding to the sample data when the sample data is used as the prediction sample data, and to acquire the number of forgetting times of the sample data based on the prediction result corresponding to the sample data when the sample data is used as the training sample data.
[0266] Optionally, the processing module 902 is further configured to retrain the target model using the sample data set from which the noise sample data and the second sample data are removed.
[0267] The data processing device provided in the embodiment of the present application can be used to execute the aforementioned data processing method. Its implementation principle, process and beneficial effects can be found in the aforementioned embodiment and will not be repeated here.
[0268] Fig.10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Fig.10 As shown, the electronic device 1000 may include: a memory 1001 and a processor 1002. Optionally, the electronic device may further include a communication interface 1003, wherein the memory 1001 and the processor 1002 communicate; illustratively, the memory 1001, the processor 1002 and the communication interface 1003 may communicate via a communication bus 1004, the memory 1001 is used to store a computer program, and the processor 1002 executes the computer program to implement the method of the above embodiment.
[0269] Optionally, the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc. The steps in the method embodiments disclosed in the present application may be directly implemented as being executed by a hardware processor, or may be implemented by a combination of hardware and software modules in the processor.
[0270] An embodiment of the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, a method in any of the above method embodiments is implemented.
[0271] An embodiment of the present application also provides a computer program product, including a computer program, which implements the method in any of the above method embodiments when the computer program is executed by a processor.
[0272] All or part of the steps of the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a readable memory. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned memory (storage medium) includes: read-only memory (ROM), RAM, flash memory, hard disk, solid state drive, magnetic tape, floppy disk, optical disc and any combination thereof.
[0273] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processing unit of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0274] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0275] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0276] The database management system provided in the embodiment of the present application can be used to execute the aforementioned data query method. Its implementation principle, process and beneficial effects can be found in the aforementioned embodiment and will not be repeated here.
[0277] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0278] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A data processing method, characterized in that: The method comprises: Obtaining statistical data of sample data in a sample data set; the statistical data comprising: the number of forgetting times when the target model is trained using the sample data; Based on the statistical data of the sample data, N candidate sample data are determined; the value of N is related to the confidence joint probability distribution matrix corresponding to the sample data set; The first sample data whose predicted label is inconsistent with the marked label among the N candidate sample data is removed from the sample data set to remove the noise sample data.
2. The method according to claim 1, characterized in that The determining N candidate sample data based on the statistical data of the sample data includes: Sorting the sample data in descending order of the number of forgetting times of the sample data; The first N sample data in the sorting are used as the candidate sample data.
3. The method according to claim 2, characterized in that The statistical data also includes: the predicted probability of the sample data corresponding to each labeled label; the method also includes: Based on the predicted probability of each labeled label corresponding to the sample data, obtaining a confidence joint probability distribution matrix corresponding to the sample data set; The value of N is determined based on the confidence joint probability distribution matrix and the number of samples in the sample data set.
4. The method according to claim 3, characterized in that The determining the value of N based on the confidence joint probability distribution matrix and the number of samples in the sample data set includes: Obtaining the prediction error probability corresponding to the sample data set according to the sum of the error probability elements corresponding to each labeled label in the confidence joint probability distribution matrix; The value of N is determined according to the product of the prediction error probability and the number of samples.
5. The method according to claim 3, characterized in that: The step of obtaining a confidence joint probability distribution matrix corresponding to the sample data set based on the predicted probability of each labeled label corresponding to the sample data set includes: Based on the predicted probability of each labeled label corresponding to the sample data, obtaining a confidence joint count matrix corresponding to the sample data set; Based on the confidence joint count matrix, the confidence joint probability distribution matrix is obtained.
6. The method according to claim 5, characterized in that The step of obtaining a confidence joint count matrix corresponding to the sample data set based on the predicted probability of each labeled label corresponding to the sample data set includes: Based on the predicted probability of each labeled label corresponding to the sample data, obtain the predicted label corresponding to the sample data and the probability corresponding to each predicted label; Based on the predicted label corresponding to the sample data and the probability corresponding to each predicted label, obtaining a confidence threshold corresponding to each predicted label; Based on the confidence threshold corresponding to each predicted label, the predicted label corresponding to the sample data, and the probability corresponding to each predicted label, a confidence joint count matrix corresponding to the sample data set is obtained.
7. The method according to claim 6, characterized in that The step of obtaining the confidence threshold corresponding to each prediction label based on the prediction label corresponding to the sample data and the probability corresponding to each prediction label includes: Based on the average value of the probability corresponding to each predicted label corresponding to the sample data, a confidence threshold corresponding to each predicted label is determined.
8. The method according to claim 6, characterized in that The step of obtaining a confidence joint count matrix corresponding to the sample data set based on the confidence threshold corresponding to each predicted label, the predicted label corresponding to the sample data set, and the probability corresponding to each predicted label includes: Based on the confidence threshold corresponding to each predicted label and the probability corresponding to each predicted label, obtain the number of sample data with correct predictions and the number of sample data with incorrect predictions corresponding to each predicted label; Based on the number of correctly predicted sample data corresponding to each predicted label and the number of incorrectly predicted sample data, a confidence joint counting matrix corresponding to the sample data set is constructed.
9. The method according to claim 5, characterized in that Before acquiring the confidence joint probability distribution matrix based on the confidence joint count matrix, the method further includes: The confidence joint count matrix is modified based on the number of sample data corresponding to each annotated label in the sample data set.
10. The method according to claim 9, characterized in that The step of modifying the confidence joint counting matrix based on the number of sample data corresponding to each annotated label in the sample data set includes: According to the number of sample data with correct predicted labels corresponding to each labeled label and the number of sample data with incorrect predicted labels, obtain the proportion of sample data with correct predicted labels corresponding to each labeled label and the proportion of sample data with incorrect predicted labels; Obtaining a corrected number of sample data for which each labeled label corresponds to a correct predicted label according to the proportion of sample data for which each labeled label corresponds to a correct predicted label, and the number of sample data for which each labeled label corresponds to a correct predicted label in the sample data set; Obtaining the corrected number of sample data with incorrect predicted labels corresponding to each labeled label according to the proportion of sample data with incorrect predicted labels corresponding to each labeled label and the number of sample data corresponding to each labeled label in the sample data set; The confidence joint counting matrix is updated based on the correct number of sample data with correct predicted labels corresponding to each labeled label and the correct number of sample data with incorrect predicted labels.
11. The method according to any one of claims 2 to 9, characterized in that: The method further comprises: The second sample data among the last M sample data in the sorting is removed from the sample data set; the second sample data is sample data whose predicted label is consistent with the marked label.
12. The method according to claim 11, characterized in that The number of forgetting times of the last M sample data is less than or equal to the preset number of forgetting times.
13. The method according to any one of claims 3 to 9, characterized in that: The obtaining of statistical data of the sample data in the sample data set includes: Using the sample data set, training the target model in a cross-validation manner; Based on the training process data of the sample data, statistical data of the sample data are obtained.
14. The method according to claim 13, characterized in that The using the sample data set to train the target model in a cross-validation manner includes: Dividing the sample data set to obtain K equal sample data subsets; The target model is trained using K equal subsets of sample data.
15. The method according to claim 14, characterized in that The training process data based on the sample data is used to obtain statistical data of the sample data, including: Based on the prediction result corresponding to when the sample data is used as the prediction sample data, obtaining the prediction probability and the prediction label of the sample data; Based on the prediction result corresponding to when the sample data is used as training sample data, the number of forgetting times of the sample data is obtained.
16. The method according to any one of claims 1 to 9, characterized in that: The method further comprises: The target model is retrained using the sample data set from which the noise sample data and the second sample data are removed.
17. A data processing device, characterized in that: The device comprises: An acquisition module is used to acquire statistical data of sample data in a sample data set; the statistical data includes: the number of forgetting times when the target model is trained using the sample data; A processing module is used to determine N candidate sample data based on the statistical data of the sample data; remove the first sample data whose predicted label is inconsistent with the marked label among the N candidate sample data from the sample data set to remove noise sample data; the value of N is related to the confidence joint probability distribution matrix corresponding to the sample data set.
18. An electronic device, characterized in that: The electronic device comprises: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 16.
19. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to any one of claims 1 to 16.
20. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 16 when being executed by a processor.