Deep learning efficient training method for error label data

By adopting greedy algorithms and adaptive update strategies in deep learning, the problem of difficult to take into account training efficiency and accuracy caused by wrong label data is solved, and efficient training in data lake scenarios is achieved, training efficiency is improved and model performance is maintained.

CN120258088APending Publication Date: 2025-07-04BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510351110.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In deep learning, there is a problem that wrong label data makes it difficult to balance training efficiency and model accuracy, especially in data lake scenarios, where existing methods are inefficient and complex in dealing with error label detection and subset selection due to data diversity and label uncertainty.

Method used

A low-complexity greedy algorithm is used to select the most profitable instance in each iteration, combine the adaptive update strategy, error label detection and subset selection are performed through iterative working methods, and the training process is optimized by gradient approximation and adaptive update.

Benefits of technology

It significantly improves training efficiency while maintaining the accuracy of the model, and can efficiently train deep learning models in the presence of wrong labels, improving training efficiency by more than 10 times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258088A_ABST
    Figure CN120258088A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning, and discloses a deep learning efficient training method for error tag data. In the whole training process, the Deem adopts an iterative working mode, and error tag cleaning and subset selection operations are executed at the same time in each iteration. According to the method, a low-complexity greedy algorithm is designed for solving, weights are distributed for subset instances through an approximation ratio mode to establish a mapping relation with a complete set, and the subset gradient can still well represent the complete set. In order to improve the efficiency, the Deem designs an adaptive subset updating strategy. And each iteration only pays attention to the affected part of the previous subset and data set change increment update, so that repeated calculation is avoided, and the training efficiency is improved. Wide experiments on six data sets and ten baseline methods show that the Deem can improve the training efficiency by 10 times on the premise of not sacrificing the accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and particularly to an efficient training method for deep learning oriented to mislabeled data. Background Art

[0002] Modern deep learning relies on a large amount of training data to train large models, consuming huge computing resources, and the training process is expensive and time-consuming. At the same time, in the data lake environment for storing and managing large-scale heterogeneous data, due to diverse data sources and uneven data quality, the labeled data in the data lake often has noise and mislabeling. In addition, the unstructured data characteristics of large-scale data lakes make data cleaning and annotation more difficult, further increasing the challenges in the deep learning training process. Since data lakes usually store massive information from different business systems, sensors, logs, and external data sources, label errors are more common in the data lake scenario, and mislabeled data will seriously affect the model performance. Generally speaking, dealing with this problem is divided into two steps: first, run an effective mislabel detection method to improve the quality of training data, and then run a subset selection method to speed up the training to achieve it. However, existing methods such as (KNN, HEARDING) perform poorly in actual scenarios when dealing with these problems because they only focus on accelerating training (subset selection) or eliminating the impact of mislabeled data (mislabel detection) in one aspect. There are already mature algorithms to support the above two problems respectively, and the operations are relatively simple. However, in the data lake scenario, data often exists in a multi-source, dynamic, and unstructured form, making the direct application of these algorithms difficult.

[0003] Therefore, the present invention aims to solve the problem that it is difficult to balance training efficiency and model accuracy in deep learning. At the same time, for these two goals, an effective method is proposed to learn a high-quality deep learning model using a large training data set containing mislabeled data (especially the data in the data lake).

[0004] This method mainly faces two challenges: First, due to the correlation between gradients and training losses, mislabel detection and subset selection may promote each other, and the computational process is relatively complex. Second, considering that the labels are modeled as probability distributions over all possible classes, the proposed method must address the problem of gradient approximation under the label uncertainty existing in the data lake. For the former, existing subset selection techniques mainly focus on efficiency issues rather than solving the effectiveness issues caused by mislabels. However, mislabeled data often changes the decision boundary, thereby degrading the model performance, and the diversity of data in the data lake exacerbates the complexity of this problem. Therefore, the present invention intelligently utilizes the training losses and gradients to achieve training efficiency and model quality by interleaving the mislabel detection process and the subset selection process. More specifically, Deem models the subset selection problem as the problem of selecting a subset of training instances that accurately approximates the gradients computed over the entire training set. This approximation in turn significantly accelerates the loss-based mislabel detection. For the latter, most existing methods are specific to certain models, such as k-means, convex loss models. Or they require a validation set as a reference, which has problems in practical applications. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides an efficient deep learning training method for mislabeled data. This method designs a low-complexity greedy algorithm that adds the instance with the highest gain in each iteration to construct the entire subset through a finite number of iterations. In each iteration, this method iteratively samples instances to select the instance with the maximum gain. The advantage of this processing is that the time complexity is relatively low, thereby theoretically ensuring the effectiveness and accuracy of the model. Considering that the cost of regularly re-running the subset selection algorithm in each iteration is very high. Therefore, the present invention designs an effective adaptive update strategy to selectively update the subset according to the change of gradients. More specifically, the present invention continuously monitors the gradient changes of all instances and identifies those instances that have significantly changed in the gradient space and updates the subset.

[0006] The present invention designs an efficient data training framework. Specifically, it relates to a method for achieving efficient deep learning training through subset selection and adaptive data instance augmentation in the presence of mislabels. The efficient deep learning training method for mislabeled data provided by the present invention includes the following steps:

[0007] (1) Mislabel detection. Based on the characteristic that the deep learning model first fits the correctly labeled instances and then the mislabeled instances during training, the mislabeled instances have a larger loss during training. Accordingly, the present invention reduces the impact of mislabels on training by iteratively removing the instances with large losses and continuously repeats and optimizes the mislabel detection in subsequent iterations.

[0008] After processing the mislabeled instance examples, the present invention evaluates the gradient approximation error e between the original dataset D and the subset C. If e is greater to a certain extent than the error e' caused by the last subset re-selection or the subset C is an empty set, then this indicates that local updates may not be sufficient to maintain the representativeness of the subset. In this case, the present invention triggers a subset re-selection process.

[0009] (2) Uncertainty-based subset selection method. The gradient descent method estimates the complete gradient based on the gradients of each instance. Therefore, a good subset is a small number of instances whose weighted sum of gradients is very close to the complete gradient. Then, the gradient descent algorithm can be applied to this subset, using the weights as the step size. In this way, the performance of the model generated by Deem can be comparable to that of the model trained using the complete training set, but with much higher efficiency.

[0010] (3) Optimizing the subset selection problem. The above preparations theoretically complete the process of subset selection. However, from the perspective of the specific implementation process, it is infeasible to find the optimal solution of the equation by enumerating the gradients of the entire parameter space and considering all possible subsets. Therefore, the present invention provides a series of optimization methods, including calculating and minimizing the upper bound of the gradient approximation error, and designing a greedy algorithm to effectively solve this problem with an approximation ratio.

[0011] (4) Adaptive subset update. The present invention designs an adaptive subset update algorithm so that each instance in the complete dataset can still be well represented by the instances in the subset.

[0012] Preferably, in the above data training method for a deep learning model, the step (1) mislabel detection mainly includes the following steps:

[0013] (1-1) Construct a loss function based on soft labels. Training with mislabels will degrade the model performance. To mitigate their impact on gradient approximation, the present invention converts the labels of the training instances into soft labels. Formally, each instance o i is associated with a set of probabilities P(l|x i ), where x i is the feature vector of the instance o i , and l is a value in the set of soft labels. Therefore, the loss function for training a deep learning model based on soft labels can be expressed as: where θ is the parameter selected by the model and f(x) is the loss function.

[0014] (1-2) Calculate the probability of the soft label. As mentioned above, the loss is effective in measuring the cleanliness of training instances. To avoid the bias of the loss in a single iteration, the present invention considers the average loss of the recent several epochs. In addition to the average loss, the standard deviation of the losses of these epochs should also be considered. The reason is that an instance with a high loss but also a high standard deviation indicates that it may be a clean instance that is difficult to fit. Therefore, the possibility of its being mislabeled should be reduced. Take the arithmetic mean of the normalized average loss function and the standard deviation as the evaluation metric h i (θ|x i ,l) value, and calculate The present invention identifies the instance with the lowest p(l|x i ) as a mislabel and removes it from D.

[0015] After completing the processing of the mislabeled data, the subset selection method based on uncertainty described in steps (2) and (3) can be used successively to process the data set, including the following steps:

[0016] (2-1) Define the subset selection problem based on the soft label. The established equation aims to minimize the difference between the subset and the complete set given the entire parameter space θ∈Θ. The present invention solves this problem by assigning a positive weight ω to each instance in the subset and calculating the gradient of the subset as the weighted sum of the gradients of these selected instances. Since the labels of instances in different classes vary greatly, there are significant differences in the gradients. Therefore, the present invention can select subsets for instances in different classes separately and then merge them. Without loss of generality, the discussion of the present invention will focus on the subset selection problem of one class. j to solve this problem and calculate the gradient of the subset as the weighted sum of the gradients of these selected instances. Since the labels of instances in different classes vary greatly, there are significant differences in the gradients. Therefore, the present invention can select subsets for instances in different classes separately and then merge them. Without loss of generality, the discussion of the present invention will focus on the subset selection problem of one class.

[0017] (2-2) Establish the mapping between the subset and the complete data set.

[0018] i) Before selecting the subset, first determine the number K of subsets that are expected to be selected.

[0019] ii) Establish the mapping γ(j)=i, representing that the element c in the subset j is derived from the i-th instance of the complete data set.

[0020] iii) Establish the mapping representing that the i-th instance in the complete data set is assigned to the element c in the subset j .

[0021] iv) Set the weight index w j representing the element c in the subset jThe number of instances in the corresponding complete dataset. In the present invention, instances with similar probabilities and gradients are assigned to the same subset. It is not difficult to prove that the gradient of the subset constructed in this way is similar to the gradient of the complete dataset when the weight is used as the step size.

[0022] Preferably, the specific operation steps for optimizing subset selection in step S3 are as follows:

[0023] (3-1) Minimize the upper bound of the gradient approximation error. Based on the methodology of the present invention, when each instance o in the complete dataset D i is assigned to the instance with the most similar gradient in subset C, the upper bound can be minimized. To further limit the gradient approximation error for all θ ∈ Θ, the present invention transforms the subset selection under the uncertainty problem into: s.t. |C| ≤ K Next, the present invention shows that the gradient difference between two trained instances (i, j) can be bounded by the last layer of the neural network. Specifically, considering the T-layer perceptron, the present invention defines as the Lipschitz continuous activation function of layer t. θ (t) is the weight matrix of layer t. To calculate the gradient of the loss function f with respect to the weights θ (t) of the last layer, the present invention uses backpropagation as follows:

[0024] i) Determine the expression forms of the input and output of the last layer of the neural network: is the input of the last layer T, and the output before activation of the last layer is The activated output is

[0025] ii) Solve the gradient with respect to θ according to the chain rule: By the chain rule: where is expressed as

[0026] iii) Determine the expression of the maximum distance between two gradients in the entire parameter space Θ in the above formula as the approximation error bound: The maximum distance between two gradients in the parameter space Θ can be expressed as:

[0027]

[0028] (3-2) Greedy algorithm for subset selection. The present invention first proves that the subset selection problem under uncertainty has the submodular property and is NP-hard. Therefore, using the submodular property, the present invention develops a greedy algorithm to achieve the approximation ratio. The present invention iteratively adds the instance o that most significantly reduces the gradient approximation error i to subset C until K instances are obtained. Once the subset is selected, weights are assigned to the instances in C. The greedy algorithm provides an optimal solution Approximation, where α is the sample ratio based on the submodular property. This greedy algorithm adds the instance with the highest gain in each iteration, thus requiring K iterations to construct the entire subset. In each iteration, it iteratively samples instances to select the instance with the maximum gain. The benefit calculation is to calculate the upper bound of the gradient approximation error, and its time complexity is O(N).

[0029] Preferably, when faced with the increase in computational consumption caused by multiple iterations, the adaptive subset selection method described in step (4) can be well solved, and its main steps are as follows:

[0030] (4-1) Adaptive deletion of mislabeled instances. As discussed in the first section, in each iteration, a small group of mislabeled instances D m will be deleted. Each o i ∈D m can be divided into two cases. If o i ∈C, then the instance represented by o i should be reassigned, and the weight of the subset should be updated, and then instances must be incrementally supplemented. Conversely, it can be directly deleted from D.

[0031] (4-2) Adaptive gradient change augmentation. As the training continues, each instance gradient will change. When the change is large, it indicates that the instance can no longer be well represented by the original instances in the subset, and it is necessary to reassign the instance. That is, when it represents that o i significantly deviates from it is necessary to reassign o i and determine the weight. After the above reallocation, check whether these instances in the original subset can still represent the instances assigned to them based on the current parameter θ'. When it indicates that c j is far from the instance currently represented by c j . Therefore, adjust c j by deleting it and select another one from D to supplement. Finally, the present invention also proves that applying the subset update algorithm will produce a good limit on the gradient approximation error.

[0032] Compared with the related technologies, the efficient deep learning training method for mislabeled data provided by the present invention has the following beneficial effects:

[0033] The data training method for deep learning models proposed by the present invention can accurately model data in the presence of different types of incorrect labels and can select a suitable subset to replace the original dataset for more efficient training. By using an adaptive and continuously updated strategy, the present invention can update the data training strategy according to the feedback of the model to guide the data training process, significantly improving the efficiency of the model in dealing with large datasets.

[0034] Other features and advantages of the present application will be described in the following specification, and in part, will become apparent from the specification or will be understood by implementing the present application. Other advantages of the present application can be achieved and obtained through the solutions described in the specification and the drawings. Brief Description of the Drawings

[0035] Figure 1 It is a schematic diagram of the processing flow framework of the Deem system described in the present invention;

[0036] Figure 2 It is a schematic diagram of the method for establishing the mapping between the subset and the complete dataset described in the present invention;

[0037] Figure 3 It is a flowchart of the efficient deep learning training method for error label data proposed by the present invention. Detailed Embodiments

[0038] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings. The embodiments described in the drawings are only used to explain the present invention and cannot be construed as a limitation of the present invention.

[0039] Combined with Figures 1 - 3 , the specific implementation process of the present invention is elaborated as follows:

[0040] The flowchart of the efficient deep learning training method for error label data proposed by the present invention is as Figure 3 shown, and the method includes the following steps:

[0041] Step (1): Error label detection. Based on the characteristic that the deep learning model first fits the correct label instances and then the incorrect label instances during training, the incorrect label instances have a large loss during training. Accordingly, the present invention reduces the impact of incorrect labels on training by iteratively removing the instances with large losses and continuously repeats and optimizes the error label detection in subsequent iterations.

[0042] Step (1-1): Construct a loss function based on soft labels. Training with incorrect labels will reduce the performance of the model. To reduce their impact on gradient approximation, we convert the labels of the training instances into soft labels. Formally, each instance o i is associated with a set of probabilities P(l|xi ) corresponds, where x i is an instance o is the eigenvector of i, and l is a value in the set of soft labels. Therefore, the loss function for training a deep learning model based on soft labels can be expressed as: where θ is the parameter selected by the model, and f(x) is the loss function.

[0043] Step (1-2): Calculate the probability of the soft label. As mentioned above, the loss is effective in measuring the cleanliness of training instances. To avoid the bias of the loss in a single iteration, we consider the average loss of the last few epochs. In addition to the average loss, the standard deviation of the losses in these epochs should also be considered. The reason is that an instance with a high loss but also a high standard deviation indicates that it may be a clean instance that is difficult to fit. Therefore, the possibility of its being mislabeled should be reduced. The present invention takes the arithmetic mean of the normalized average loss function and the standard deviation as the evaluation index h i (θ|x i ,l) value, that is and calculate We identify the instance with the lowest p(l|x i ) as a mislabel and remove it from D.

[0044] After processing the mislabeled instances, we evaluate the gradient approximation error e between the original dataset D and the subset C. If e is greater to a certain extent than the error e' caused by the last subset reselection or the subset C is an empty set, then this indicates that the local update may not be sufficient to maintain the representativeness of the subset. In this case, we trigger the subset reselection process.

[0045] Step (2): Uncertainty-based subset selection method. The Gradient descent method estimates the complete gradient based on the gradient of each instance. Therefore, a good subset is a small number of instances whose weighted sum of gradients is very close to the complete gradient. Then, the gradient descent algorithm can be applied to this subset, using the weights as the step size. In this way, the performance of the model generated by Deem can be comparable to that of the model trained using the complete training set, but with much higher efficiency. The equation established in the present invention takes into account the influence of uncertain labels, that is, the probability p index established in the first part is introduced to optimize the gradient descent method.

[0046] Step (2-1): Define the subset selection problem based on soft labels. The established objective of the equation is to minimize the difference between the subset and the complete set given the entire parameter space θ∈Θ. We assign a positive weight ω to each instance in the subset jTo solve this problem, the gradient of the subset is calculated as the weighted sum of the gradients of these selected instances. Since the labels of instances in different classes vary greatly, there are significant differences in the gradients. Therefore, we can select subsets for instances in different classes separately and then combine them. Without loss of generality, the discussion of the present invention will focus on the problem of subset selection for one class. Specifically, the equation is:

[0047] where C l represents the set of instances of a specific class l in the subset. Naturally, since the C l corresponding to different classes are disjoint, so C = ∪C l , K l is the number of instances in the corresponding class l, that is

[0048] Step (2-2): Establish the mapping between the subset and the complete dataset.

[0049] i) Before selecting the subset, first determine the number K of subsets you hope to select.

[0050] ii) Establish the mapping γ(j) = i, j ∈ [1, N], i ∈ [1, K], representing that the element c j in the subset is derived from the i-th instance in the complete dataset.

[0051] iii) Establish the mapping c j ∈ C, representing that the i-th instance in the complete dataset is assigned to the element c j in the subset.

[0052] iv) Set the weight index representing the number of instances in the complete dataset corresponding to the element c j in the subset. The present invention assigns instances with similar probabilities and gradients to the same subset. It is not difficult to prove that the gradient of the subset is similar to the gradient of the complete dataset after constructing the subset with the weight as the step size.

[0053] As Figure 2 shown, γ(1) = 1, γ(2) = 4, γ(3) = 7, that is, the 1st, 4th, and 7th instances in the complete dataset are taken as subsets 1, 2, and 3. respectively represent that the 3rd, 4th, and 5th instances in the original dataset are assigned to subset 2 because they have similar color gradients (similar colors in the figure). Therefore, the weights ω1 = 2, ω2 = 3, ω3 = 3. The subset constructed in this way is effective because

[0054] Step (3): Optimize the subset selection problem. The above preparations have theoretically completed the process of subset selection. However, from the perspective of the specific implementation process, it is infeasible to find the optimal solution of the equation by enumerating the gradients of the entire parameter space and considering all possible subsets. Therefore, the present invention provides a series of optimization methods, including calculating and minimizing the upper bound of the gradient approximation error, and designing a greedy algorithm to effectively solve this problem with an approximation ratio.

[0055] Step (3-1): Minimize the upper bound of the gradient approximation error. Based on the methodology of the present invention, when each instance o in the complete dataset D i is assigned to the instance with the most similar gradient in the subset C, the upper bound can be minimized. To further limit the gradient approximation error for all θ ∈ Θ, we can finally transform the subset selection under the uncertainty problem into:

[0056] Next, it is found that the gradient difference between two trained instances (i, j) can be restricted by the last layer of the neural network. Specifically, considering the T-layer perception, we define (t) as the Lipschitz continuous activation function of layer t. θ (t) is the weight matrix of layer t. To calculate the gradient of the loss function f with respect to the weights θ

[0057] of the last layer, we use backpropagation. For simplicity, we denote the p corresponding to the instances (i, j) as p1 and p2 respectively, and the operations are as follows: As the input of the last layer T, the output of the last layer before activation is The activated output is

[0058] ii) Gradient with respect to θ: By the chain rule: where is denoted as

[0059] iii) Upper bound of the approximation error: The maximum distance between two gradients in the parameter space Θ can be expressed as:

[0060] Step (3-2): Greedy algorithm for subset selection. The present invention first proves that the subset selection problem under uncertainty has the submodular property and is NP-hard. Therefore, using the submodular property, we develop a greedy algorithm to achieve an approximation ratio.

[0061] The present invention first initializes the subset C to be empty, and then iterates through each instance in D\C to find the instance o that results in the largest error reduction i , and adds it to the subset C. However, to improve efficiency, we can sample a subset D s from D\C s , and select the instance that benefits the most from D . Once the subset is selected, we assign weights to the instances in C. The greedy algorithm provides an approximation of the optimal solution

[0062] , where α is the sample ratio based on the submodular property. s This greedy algorithm adds the instance with the highest benefit in each iteration, thus requiring K iterations to construct the entire subset. In each iteration, it iteratively samples instances to select the instance with the maximum benefit. The benefit calculation is to calculate the upper bound of the gradient approximation error, and its time complexity is O(N). Overall, its time complexity is O(|D

[0063] |KN). Although we have a total of r labels, in fact, for instances that may be clean, we do not need to consider their soft labels. In addition, for instances that are likely to be mislabeled, we only need to consider the top k labels (e.g., k = 3) that are most likely to be the true labels, rather than considering all their possible labels. Therefore, the impact of r on the time complexity can be ignored.

[0064] Step (4): Adaptive subset update. As discussed in the first 3 sections of the present invention, two factors will lead to the update of the selected subset, namely (1) the removal of mislabeled instances and (2) the gradient change caused by iterative training. However, it is expensive to routinely recalculate the subset from scratch at each iteration. Therefore, we have developed an adaptive strategy to incrementally update the subset based on the previous subset and dataset changes, so that each instance in the complete dataset can still be well represented by the instances in the subset. m Step (4-1): Adaptive removal of mislabeled instances. In each iteration, a small group of mislabeled instances D i will be removed. Each o m ∈ D i can be divided into two cases. If o i ∈ C, then the instance represented by o should be reallocated, and the weights of the subset should be updated, and then instances must be incrementally supplemented. Conversely, it is simply deleted from D, and the corresponding weights are reduced. Then we must incrementally supplement instances. Use the greedy algorithm starting from D, and reassign the instances in D to the instances in the subset together with the next gradient change step.

[0065] Step (4-2): Adaptive gradient change augmentation. As the training continues, the gradient of each instance changes. When the change is large, it indicates that the instance can no longer be well represented by the original instances in the subset and reallocation is necessary. Specifically, when (where θ is the parameter of the model in the previous training and θ' is the parameter of the model in the current iteration), it represents that the current o i and have a gradient difference greater than and the maximum difference between the instances represented by in the previous iteration, that is, o i deviates significantly from and reallocation of o i is necessary and the weights are determined.

[0066] After the above reallocation, we check whether these instances in the original subset can still represent the instances assigned to them based on the current parameter θ'. When it indicates that c j is far from the instance currently represented by c j . Therefore, we adjust c by deleting it j and select another one from D to supplement it. Assuming the number of instances with large gradient changes in C is v, the total time complexity is O(v|D s |N), which is actually less than the complexity of selecting the subset from scratch, thus improving the efficiency. Finally, the present invention proves that applying the subset update algorithm will impose good restrictions on the gradient approximation error.

[0067] The overall process is as Figure 1 shown. The present invention deletes some of the identified mislabeled instances from the original dataset, selects or updates the subset according to the judgment conditions, and uses the processed representative subset for the training of deep learning, and realizes the deletion of adaptive error instances and gradient descent optimization in this process. The iteration stops until the model converges.

[0068] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. An efficient deep learning training method for mislabeled data, characterized in that, It includes the following steps: S1. Error label detection: Based on the loss difference between correct label instances and wrong label instances during the training process of the deep learning model, identify and remove wrong label instances; S2. Uncertainty-based subset selection: According to the gradient approximation principle, select subsets for each category in the complete dataset respectively, make the weighted sum of subset gradients close to the gradient of the complete dataset by assigning weights, and establish the mapping relationship between the subset and the complete dataset; S3. Optimize subset selection: By minimizing the upper bound of the gradient approximation error, combined with the greedy algorithm, iteratively select the instances that most significantly reduce the error to construct the target subset, and assign weights to approximate the gradient of the complete dataset; S4. Adaptive subset update: Dynamically adjust the subset to cope with the deletion of wrong labels and gradient changes, including supplementing new instances after deleting wrong label instances in the subset, and reassigning deviated instances according to the gradient difference to ensure the representativeness of the subset.

2. The efficient deep learning training method for error-labeled data according to claim 1, wherein The error label detection in step S1 mainly includes the following steps: S1.

1. Construct a loss function based on soft labels: Convert the labels of training instances into soft labels. The loss function for training a deep learning model based on soft labels can be expressed as: where θ are the parameters selected by the model, f(x) is the loss function, and each instance o i is associated with a set of probabilities P(l|x i ), where x i is the feature vector of instance o i and l is a value in the set of soft labels; S1.

2. Calculate the probability of the soft label: Take the arithmetic mean of the normalized average loss function and the standard deviation as the evaluation metric h i (θ|x i ,l) values, and calculate Identify the instance with the lowest p(l|x i ) as the wrong label and remove it from the original dataset D.

3. The efficient deep learning training method for error-labeled data according to claim 1, characterized in that The uncertainty-based subset selection method in step S2 includes the following steps: S2.

1. Define the subset selection problem based on soft labels: Assign a positive weight ω to each instance in the subset j , and calculate the gradient of the subset as the weighted sum of the gradients of these selected instances. Select subsets separately according to the instances in different classes and merge them; S2.

2. Establish the mapping between the subset and the complete dataset: S2.2.

1. Before selecting the subset, first determine the number K of subsets to be selected; S2.2.

2. Establish the mapping γ(j) = i, representing the element c in the subset j derived from the i-th instance of the complete data set; S2.2.

3. Establish mapping It represents that the i-th instance in the complete dataset is assigned to the element c in the subset j ; S2.2.

4. Set the weight index w j Represents the element c in the subset j The number of instances in the corresponding complete dataset 4. The efficient deep learning training method for error-labeled data according to claim 1, wherein The specific operation steps of optimizing subset selection in step S3 are as follows: S3.

1. Minimize the upper bound of the gradient approximation error: Convert the subset selection under the uncertainty problem to: s.t. |C| ≤ K to limit the gradient approximation error for all θ ∈ Θ; Next, the gradient difference between two training instances (i, j) is limited by the last layer of the neural network; S3.

2. Greedy Algorithm for Subset Selection: Exploiting the submodular property, develop a greedy algorithm to achieve an approximation ratio, i.e., iteratively add the instance o that most significantly reduces the gradient approximation error i to the subset C until K instances are obtained; after selecting the subset, assign weights to the instances in C; The greedy algorithm provides an approximation where α is a sample ratio based on the submodular property.

5. The efficient deep learning training method for error-labeled data according to claim 1, characterized in that, The specific steps of adaptive subset update in step S4 are: S4.1, Instance deletion with adaptive error marking: In each iteration, a small group of instances D of error markings m will be deleted; each o i ∈ D m can be divided into two cases: If o i ∈ C, the instance represented by o i should be reallocated, and the weights of the subsets updated, then instances are incrementally supplemented; otherwise it is directly deleted from D; S4.

2. Adaptive gradient change supplementation: When o j ∈ D, it means that o i significantly deviates from then re - allocate o i and determine the weight; After the above reallocation, check whether these instances in the original subset can still represent the instances assigned to them based on the current parameter θ'; when o i ∈ D, this indicates that c j is far from the instance currently represented by c j ; adjust c by deleting it j and select another one from D to supplement it.

6. The efficient deep learning training method for error-labeled data according to claim 4, wherein In step S3.1, the specific operation of restricting the gradient difference between two trained instances (i, j) by the last layer of the neural network is: Based on the T-layer perception, define as the Lipschitz continuous activation function of layer t, θ (t) is the weight matrix of layer t; using backpropagation, calculate the gradient of the loss function f with respect to the weights θ (t) of the last layer as follows: S3.1.

1. Determine the expression forms of the input and output of the last layer of the neural network: As the input to the last layer T, the output of the last layer before activation is The activated output is S3.1.

2. Solve the gradient with respect to θ according to the chain rule: By the chain rule: where is expressed as S3.1.

3. Determine the expression of the maximum distance between two gradients in the entire parameter space Θ in the above formula as the approximation error bound: The maximum distance between two gradients in the parameter space Θ can be expressed as: