Method and system for automatically screening and recommending annotation data

By using a confidence-regularized loss function and an attention mechanism to filter high-quality samples, this method solves the problem of filtering noise-dependent labeled data in existing technologies, achieving efficient sample filtering and improved model performance.

CN120929925AInactive Publication Date: 2025-11-11HANGZHOU FIVE DIMENSIONAL DATA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511112134.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies struggle to adaptively select high-quality samples from labeled data with instance-dependent noise without relying on prior knowledge of noise, and to make reasonable use of the effective information of damaged samples, resulting in poor model training performance.

Method used

The classification model is trained using a confidence regularization loss function. The sum of the cross-entropy loss term and the confidence regularization term is used as the regularization loss value to dynamically select high-quality samples. The selection results are optimized through adaptive threshold iteration, and the sample features from historical rounds are mined by combining an attention mechanism to reduce noise interference.

Benefits of technology

Accurately remove noisy labels, select high-quality samples, improve model training effect and generalization ability, reduce noise interference, and improve model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929925A_ABST
    Figure CN120929925A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing, and provides an annotation data automatic screening recommendation method and system. The method comprises the following steps: receiving an initial data set containing a plurality of labeled samples; training a classification model by adopting a confidence regularization loss function on the basis that the current selection mark is a to-be-reserved mark sample; calculating the sum of the cross entropy loss item and the corresponding confidence regularization item of each labeled sample in the initial data set by using the trained classification model, and taking the sum as a regularization loss value of the labeled sample; and comparing the regularization loss value of each labeled sample with a corresponding screening threshold value, if the regularization loss value is smaller than the screening threshold value, updating the screening mark of the labeled sample to be reserved, otherwise, updating the screening mark of the labeled sample to be excluded. According to the method, noise labels can be accurately stripped, high-quality samples can be screened out, noise interference is reduced, and the model training effect and generalization ability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and more specifically, to a method and system for automatically filtering and recommending labeled data. Background Technology

[0002] In the fields of artificial intelligence and machine learning, model performance is highly dependent on large-scale, high-quality labeled data. With the rapid development of deep learning technology, the demand for labeled data for tasks such as image recognition and natural language processing is growing exponentially. However, noise is inevitably introduced during manual labeling—this noise may stem from the subjective judgment bias, limited professional knowledge, fatigue, or negligence of the labelers, resulting in some samples having labels that do not match the true categories (i.e., damaged samples). Noisy labeled data severely impacts model training: on the one hand, noise weakens the true correlation between features and labels; on the other hand, it may introduce spurious correlation patterns, causing the model to overfit to noisy labels and ultimately reducing generalization ability. Therefore, selecting high-quality samples (i.e., clean samples) from massive amounts of labeled data has become a crucial step in improving model performance.

[0003] Traditional noise processing methods are mainly divided into two categories: one is loss correction-based methods, which correct the loss function by estimating the noise transition matrix. However, these methods usually assume that the label noise is independent of the sample features (i.e., instance-independent noise) and require a lot of prior information to estimate the noise rate, which has limited applicability in real-world scenarios. The other is sample selection-based methods, which select samples with small losses as clean samples by setting a fixed threshold. However, these methods are sensitive to noise distribution, and their effectiveness drops significantly, especially in scenarios where the noise depends on the sample features (i.e., instance-dependent noise).

[0004] In practical applications, annotation noise often exhibits significant instance dependence. For example, in image classification tasks, blurry images or samples at class boundaries are more prone to mislabeling, while clear images have a lower annotation error rate; in medical image diagnosis, the annotation error of complex cases is much higher than that of typical cases. Existing methods for dealing with this type of noise have obvious limitations: either they require pre-estimation of complex noise distributions, resulting in high computational costs; or they rely on manually set screening thresholds, making it difficult to adapt to the noise characteristics of different datasets, leading to unstable screening accuracy.

[0005] Furthermore, most existing technologies treat clean samples and damaged samples as opposing groups, discarding damaged samples directly after screening and ignoring the valuable feature information they contain. In reality, the features of damaged samples may still contain valuable information for model training, and completely discarding them would waste data resources.

[0006] Therefore, how to adaptively select high-quality samples from labeled data with instance-dependent noise without relying on prior knowledge of noise, and how to make reasonable use of the effective information of damaged samples, has become a technical problem that urgently needs to be solved in the field of machine learning. Summary of the Invention

[0007] In response, the present invention provides a method, system, electronic device, computer storage medium, and computer program product for automatically filtering and recommending labeled data, in order to solve at least one of the above-mentioned technical problems.

[0008] In a first aspect, the present invention provides an automatic filtering and recommendation method for labeled data, comprising the following steps: receiving an initial dataset containing multiple labeled samples, wherein the labeled samples contain feature data and corresponding labeling; training a classification model using a confidence regularization loss function based on the currently selected labeled samples to be retained; calculating the sum of the cross-entropy loss term and the corresponding confidence regularization term of each labeled sample in the initial dataset using the trained classification model, as the regularization loss value of the labeled sample; comparing the regularization loss value of each labeled sample with the corresponding filtering threshold, wherein if the regularization loss value is less than the filtering threshold, the filtering label of the labeled sample is updated to be retained, otherwise it is updated to be excluded.

[0009] In a second aspect, the present invention provides an automatic filtering and recommendation system for labeled data. The system includes a receiving unit, a training unit, a loss calculation unit, and a filtering and recommendation unit. The receiving unit receives an initial dataset containing multiple labeled samples, each labeled sample containing feature data and corresponding labeling. The training unit trains a classification model using a confidence regularization loss function based on labeled samples currently marked as to be retained. The loss calculation unit uses the trained classification model to calculate the sum of the cross-entropy loss term and the corresponding confidence regularization term for each labeled sample in the initial dataset, using this sum as the regularization loss value for that labeled sample. The filtering and recommendation unit compares the regularization loss value of each labeled sample with a corresponding filtering threshold. If the regularization loss value is less than the filtering threshold, the filtering label of that labeled sample is updated to "to be retained"; otherwise, it is updated to "to be excluded."

[0010] In a third aspect, the present invention provides an electronic device comprising: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the computer program, when executed by the processor, implements the method as described in any of the preceding claims.

[0011] In a fourth aspect, the present invention provides a computer storage medium storing a computer program that can be executed by a processor to implement the method as described in any of the preceding claims.

[0012] In a fifth aspect, the present invention provides a computer program product comprising a computer program executable by a processor to implement the method as described in any of the preceding claims.

[0013] This invention dynamically selects samples to be retained, trains the model using a confidence-regularized loss function, calculates the regularized loss value and compares it with an adaptive threshold, and iteratively optimizes the selection results. This approach accurately removes noisy labels, selects high-quality samples, reduces noise interference, and improves model training performance and generalization ability. Attached Figure Description

[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a flowchart illustrating an automatic filtering and recommendation method for labeled data disclosed in an embodiment of the present invention.

[0016] Figure 2 This is a schematic diagram of the structure of an automatic filtering and recommendation system for labeled data disclosed in an embodiment of the present invention. Detailed Implementation

[0017] The following specific embodiments illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0018] Furthermore, the technical features involved in the different embodiments of this application described below can be combined with each other as long as they do not conflict with each other.

[0019] like Figure 1 As shown, this embodiment of the invention discloses an automatic filtering and recommendation method for labeled data, including the following method steps: S10, receiving an initial dataset containing multiple labeled samples, wherein the labeled samples contain feature data and corresponding labeling tags.

[0020] Annotated samples are basic data units containing feature data and annotation labels. Feature data refers to the original attributes of the sample (such as pixel information of an image, word vectors of text, etc.), which are the input basis for subsequent model learning. Annotation labels are the category labels of the feature data (such as image categories such as "cat" and "dog", and text sentiment labels such as "positive" and "negative"), which are the target references for model training.

[0021] After receiving the initial dataset uploaded by the user, it is stored in the sample library. The labels of these labeled samples may contain noise (such as human annotation errors), and subsequent steps will distinguish between noisy labels and clean labels.

[0022] S20: Based on the currently selected labeled samples to be retained, a classification model is trained using a confidence regularization loss function.

[0023] This step optimizes the classification model using a dynamically updated subset of samples and a specially designed loss function. The currently selected labeled samples to be retained refer to those deemed high-quality in the previous iteration (initially all samples). Choosing these high-quality samples as the training set reduces the interference of noisy labels on the classification model, making it more closely resemble the distribution characteristics of clean data.

[0024] The confidence regularization loss function consists of two parts: the first is the cross-entropy loss, which measures the degree of matching between the model's prediction and the labeled data, and is the basic loss for traditional classification tasks; the second is the confidence regularization term, which is obtained by calculating the expectation of the distribution of labeled data in the current dataset. Its role is to enhance the certainty of the classification model's prediction of samples—that is, to encourage the classification model to output high-confidence prediction results and avoid the classification model from falling into a low-confidence ambiguity due to noisy labels.

[0025] Through this step, the classification model can continuously learn on dynamically updated sample subsets, gradually improving its ability to identify clean samples and providing a reliable model foundation for subsequent sample quality assessment.

[0026] S30: Calculate the sum of the cross-entropy loss term and the corresponding confidence regularization term for each labeled sample in the initial dataset using the trained classification model, and use this sum as the regularization loss value for that labeled sample.

[0027] This step calculates the regularization loss value for each sample using the trained classification model, which serves as a quantitative basis for sample selection.

[0028] Specifically, for each labeled sample in the initial dataset, its cross-entropy loss is first calculated, reflecting the prediction error of the classification model for that sample's label; then, its corresponding confidence regularization term is calculated, reflecting the overall confidence level of the classification model's prediction for that sample. The sum of the two is the regularization loss value for that sample, which integrates prediction error and confidence information.

[0029] In practical screening, clean samples typically have lower regularization loss values ​​(the classification model predicts accurately and with high confidence), while samples with noisy labels often have higher regularization loss values ​​(large prediction bias or low confidence) due to labeling errors. Therefore, the regularization loss value can effectively distinguish sample quality and provide a quantitative standard for subsequent screening.

[0030] S40. Compare the regularization loss value of each labeled sample with the corresponding screening threshold. If the regularization loss value is less than the screening threshold, update the screening label of the labeled sample to be retained; otherwise, update it to be excluded.

[0031] This step classifies sample quality by comparing the regularization loss value with the screening threshold.

[0032] The selection threshold is calculated individually for each sample, for example, based on the average regularized loss of the model for that sample across all possible labels, rather than a fixed threshold, thus adapting to the noise characteristics of different samples. The design logic is: for a sample with acceptable labeling quality, its regularized loss value should be lower than the average loss level of the model for that sample across all labels.

[0033] If the regularization loss value of a sample is less than its screening threshold, it indicates that the sample's labeling quality is high (the classification model predicts it accurately and the confidence level meets expectations), and its screening label is updated to "to be retained"; otherwise, it indicates that the sample may have noisy labels, and its label is updated to "to be excluded". Through this step, dynamic screening of samples is achieved, gradually weeding out low-quality samples.

[0034] Repeat steps S20-S40 until the iteration termination condition is met, and recommend the currently marked samples to be retained as high-quality samples.

[0035] Each round involves retraining the classification model based on the samples selected for retention in the previous round. The new classification model is then used to re-evaluate the regularization loss values ​​of all samples, and the selection labels are updated according to the selection threshold. As the iterations proceed, the classification model's ability to identify clean samples continuously improves, and the set of samples selected for retention gradually approaches the true high-quality samples.

[0036] Termination conditions include, for example, the set of samples to be retained stabilizing (e.g., the rate of change of samples in two consecutive rounds is below a preset threshold) or reaching the maximum number of iterations, in order to balance screening accuracy and computational efficiency. Finally, the samples marked as to be retained after convergence are the high-quality samples selected, which can be directly used for subsequent model training to improve model performance.

[0037] This invention dynamically selects samples to be retained, trains the model using a confidence-regularized loss function, calculates the regularized loss value and compares it with an adaptive threshold, and iteratively optimizes the selection results. This approach accurately removes noisy labels, selects high-quality samples, reduces noise interference, and improves model training performance and generalization ability.

[0038] As an example, the confidence regularization loss function is as follows: ;in, For cross-entropy loss term, For confidence regularization, The regularization coefficient is . This represents the expected distribution of noise labels; For the feature data of the nth labeled sample, This represents the prediction result of the classification model for the nth labeled sample. The label is the annotation label corresponding to the nth labeled sample; is a random variable for noise labels.

[0039] This embodiment combines traditional cross-entropy loss with confidence regularization term to improve the robustness of the classification model to noisy labels through dual constraints.

[0040] For the cross-entropy loss term: For the model to sample The predicted probability distribution (e.g., softmax output); Indicates the model predicts samples Belongs to label The probability of.

[0041] By mapping probability to loss through logarithmic operations, the loss approaches 0 when the predicted probability is close to 1, and approaches infinity when the predicted probability is far from 1.

[0042] For the confidence regularization term: by introducing the expectation of the noise label distribution, the model's prediction behavior for noise samples is constrained, avoiding overfitting of the model to incorrect labels.

[0043] in, This is the regularization coefficient, which controls the strength of the confidence constraint; Indicates the distribution of noise labels The expectation is that the impact of all possible noise labels on the model prediction is taken into account. Indicates when the sample The real label is Cross-entropy loss at that time.

[0044] If the classification model has high confidence in predicting noisy samples (i.e.) (Large), but the label is actually incorrect. For other possible tags Expectations will increase, leading to This reduces the confidence level, thus suppressing such high-confidence erroneous predictions; conversely, if the classification model has low confidence in predicting noisy samples, then... This will increase, prompting the classification model to adjust its parameters to improve the confidence level of its prediction for that sample.

[0045] This embodiment, by calculating the expected distribution of noise labels, can adaptively mitigate the impact of noise labels on model training without prior knowledge of the noise rate. Compared to traditional methods that rely on fixed thresholds or assumptions about instance-independent noise, the loss function in this embodiment directly models instance-dependent noise, making it more relevant to real-world labeling scenarios. Furthermore, compared to fixed threshold filtering methods, the confidence regularization term in this embodiment dynamically adjusts the classification model's prediction behavior for different samples, rather than simply discarding high-loss samples, thus retaining samples that may contain valuable information.

[0046] As an example, the filtering threshold is determined as follows: for each labeled sample in the initial dataset, its regularization loss value under all possible labels is calculated, the average of all regularization loss values ​​is calculated to obtain the expected regularization loss, and the expected regularization loss is used as the filtering threshold.

[0047] Traditional methods use a single threshold (such as the quantile of the loss distribution), which cannot distinguish differences in sample difficulty, easily leading to the misclassification of complex samples or the omission of simple samples. To address this technical problem, this embodiment calculates a dynamic threshold based on the characteristics of each sample individually, achieving accurate evaluation of annotation quality. Specifically, for each labeled sample… Its screening threshold The calculation steps are as follows: First, calculate the full-label regularization loss using the following formula: ;in, For the set of all possible labels, It is a sample Predicted as a label Cross-entropy loss, It is the confidence regularization term.

[0048] Then, calculate the expected regularization loss: This value indicates the model's performance on the samples. The average prediction loss across all possible labels reflects the inherent difficulty of the sample.

[0049] Therefore, the screening threshold was determined to be... That is, the screening threshold for each sample is equal to its expected regularization loss.

[0050] The dynamic thresholding scheme in this embodiment can adapt to noise characteristics. Specifically, for complex samples (such as blurred images or semantic boundary samples), its... The threshold is naturally higher, and thus more lenient to avoid misjudgments; for simple samples (such as clear images), the threshold is strict to ensure high-confidence screening. Furthermore, compared to threshold setting methods based on prior knowledge, this invention does not require pre-estimation of noise rate or noise distribution, but is dynamically calculated entirely based on model prediction results, making it suitable for real-world scenarios where noise characteristics are unknown.

[0051] As an example, the step of training a classification model using a confidence regularization loss function based on the currently selected labeled samples to be retained includes: determining the selection label change sequence and predicted confidence sequence of samples corresponding to each historical training round; the attention module calculates the weight of each sample to be retained in the current training round based on the selection label change sequence and the predicted confidence sequence; the confidence regularization loss function is weighted based on the weights to obtain a weighted loss function; the classification model is trained using the weighted loss function; and the sample selection labels and predicted confidence of the current training round are stored.

[0052] Training using only the confidence regularization loss function, which applies a uniform treatment to samples to be retained in the current round without considering the dynamic changes of samples in historical rounds, may lead to the following problems: it cannot distinguish between consistently high-quality samples and suspicious samples with frequently fluctuating selection status; it may overfit to noisy or boundary samples; and the training direction of the classification model depends solely on the current sample distribution, lacking guidance from historical selection information, making it prone to convergence instability due to local noise interference during iteration. To address these technical problems, this invention introduces an attention mechanism to mine sample features from historical rounds, thereby overcoming the aforementioned deficiencies.

[0053] First, calculate the change sequence of screening labels and the prediction confidence sequence of corresponding samples in each historical training round of the classification model.

[0054] (1) Filtering sequences with changing markers: Record samples The screening status in the past K rounds (1 for retention, 0 for exclusion). For example, if a sample is marked as [1,0,1] in the first three rounds, it indicates that its quality is unstable and it may be a boundary sample or a noise sample.

[0055] (2) Predicting confidence sequences: Record the model's performance on samples over the past K rounds. The predicted probability of the labeled data. If the confidence level fluctuates greatly (e.g., [0.9, 0.3, 0.8]), it indicates that the classification model's judgment of the sample is unstable and there may be labeling noise.

[0056] Next, the attention module maps the two sequences to sample weights using a neural network (such as an MLP). The formula is: ,in, For learnable weight matrix, This is a bias term.

[0057] Samples that are consistently retained and have high confidence (e.g., sequences [1,1,1] with all confidence levels > 0.8) are assigned high weights (e.g., ...). For samples with frequently changing screening status or low confidence (e.g., sequences containing multiple zeros or large confidence fluctuations), assign low weights (e.g., ...). ).

[0058] Next, based on the weights obtained above, a weighted loss function is constructed and the model is trained.

[0059] Weighted loss function: ;in, The confidence regularization loss function is... For the calculation of samples The weight.

[0060] use Optimize model parameters and update the model to And store the filter flags for the current round. and prediction confidence This is used for weight calculation in the next training round.

[0061] As an example, the attention module calculates the weights of each sample to be retained in the current training round based on the filter label change sequence and the prediction confidence sequence, including: determining an interference intensity coefficient based on the round number value of the current training round, generating an interference signal based on the interference intensity coefficient, and adding the interference signal to the filter label change sequence and the prediction confidence sequence; wherein, the interference intensity coefficient is negatively correlated with the round number value; the attention module calculates the weights of each sample to be retained in the current training round based on the filter label change sequence and the prediction confidence sequence with the interference signal added.

[0062] Based on the aforementioned embodiments, this invention further dynamically introduces interference signals that are negatively correlated with the training rounds when the attention module calculates sample weights. This solves the problem that traditional attention mechanisms rely too much on stable features and are difficult to adapt to changes in the training stage, thereby achieving adaptive training that combines early noise-resistant learning with later fine-tuning.

[0063] The interference intensity coefficient is defined as: , where is the current training round, This is the initial interference strength (e.g., set to 0.1). This is the attenuation coefficient (e.g., set to 0.05). The larger the number of rounds t, the... The smaller the value (exponential decay), the stronger the early-stage interference and the weaker the later-stage interference.

[0064] For each sample Generate Gaussian interference ( (This represents the noise variance, e.g., set to 0.01), and the intensity is adjusted according to the round: ;in, It consists of the original screening marker variation sequence and the predicted confidence sequence. It is a new characteristic sequence of fusion interference.

[0065] Features containing interference The input attention module calculates weights using a multilayer perceptron (MLP) and softmax: .

[0066] Based on the scheme in this embodiment, in the early stage of training, strong interference can be used to filter historical noise and quickly establish the classification model's ability to learn from high-quality samples; while in the later stage of training, weak interference can be used to retain the dynamic features of samples and improve the screening accuracy of complex / boundary samples.

[0067] like Figure 2 As shown in the figure, this embodiment of the invention also provides an automatic filtering and recommendation system 100 for labeled data. The system 100 includes a receiving unit 101, a training unit 102, a loss calculation unit 103, and a filtering and recommendation unit 104. The receiving unit 101 receives an initial dataset containing multiple labeled samples, wherein the labeled samples contain feature data and corresponding labeling. The training unit 102 trains a classification model using a confidence regularization loss function based on the currently selected labeled samples to be retained. The loss calculation unit 103 uses the trained classification model to calculate the sum of the cross-entropy loss term and the corresponding confidence regularization term of each labeled sample in the initial dataset, as the regularization loss value of the labeled sample. The filtering and recommendation unit 104 compares the regularization loss value of each labeled sample with the corresponding filtering threshold. If the regularization loss value is less than the filtering threshold, the filtering label of the labeled sample is updated to be retained; otherwise, it is updated to be excluded.

[0068] As an example, the confidence regularization loss function is as follows: ;in, For cross-entropy loss term, For confidence regularization, The regularization coefficient is . This represents the expected distribution of noise labels; For the feature data of the nth labeled sample, This represents the prediction result of the classification model for the nth labeled sample. The label is the annotation label corresponding to the nth labeled sample; is a random variable for noise labels.

[0069] As an example, the filtering threshold is determined as follows: for each labeled sample in the initial dataset, its regularization loss value under all possible labels is calculated, the average of all regularization loss values ​​is calculated to obtain the expected regularization loss, and the expected regularization loss is used as the filtering threshold.

[0070] As an example, the training unit 102 is used to: determine the screening label change sequence and predicted confidence sequence of samples corresponding to each historical training round; the attention module calculates the weight of each sample to be retained in the current training round based on the screening label change sequence and the predicted confidence sequence; weight the confidence regularization loss function based on the weight to obtain a weighted loss function; train the classification model using the weighted loss function; and store the sample screening label and predicted confidence of the current training round.

[0071] As an example, the training unit 102 is configured to: determine an interference intensity coefficient based on the round number value of the current training round, generate an interference signal based on the interference intensity coefficient, and add the interference signal to the filter label change sequence and the prediction confidence sequence; wherein the interference intensity coefficient is negatively correlated with the round number value; and the attention module calculates the weight of each sample to be retained in the current training round based on the filter label change sequence and the prediction confidence sequence with the interference signal added.

[0072] This invention also provides an electronic device comprising: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the computer program, when executed by the processor, implements the method as described in any of the foregoing embodiments.

[0073] This invention also provides a computer storage medium storing a computer program that can be executed by a processor to implement the methods described in any of the foregoing claims.

[0074] This invention also provides a computer program product comprising a computer program that can be executed by a processor to implement the method as described in any of the foregoing embodiments.

[0075] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0076] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for automatically filtering and recommending labeled data, characterized in that: The method includes the following steps: receiving an initial dataset containing multiple labeled samples, wherein the labeled samples contain feature data and corresponding labeling; training a classification model using a confidence regularization loss function based on the currently selected labeled samples to be retained; calculating the sum of the cross-entropy loss term and the corresponding confidence regularization term for each labeled sample in the initial dataset using the trained classification model, as the regularization loss value of that labeled sample; comparing the regularization loss value of each labeled sample with the corresponding selection threshold, and if the regularization loss value is less than the selection threshold, updating the selection label of that labeled sample to be retained, otherwise updating it to be excluded.

2. The method for automatically filtering and recommending labeled data according to claim 1, characterized in that: The confidence regularization loss function is as follows: ;in, For cross-entropy loss term, For confidence regularization, The regularization coefficient is . This represents the expected distribution of noise labels; For the feature data of the nth labeled sample, This represents the prediction result of the classification model for the nth labeled sample. The label is the annotation label corresponding to the nth labeled sample; is a random variable for noise labels.

3. The method for automatically filtering and recommending labeled data according to claim 2, characterized in that: The filtering threshold is determined as follows: for each labeled sample in the initial dataset, calculate its regularization loss value under all possible labels, calculate the average of all regularization loss values ​​to obtain the expected regularization loss, and use the expected regularization loss as the filtering threshold.

4. The method for automatically filtering and recommending labeled data according to claim 3, characterized in that: Based on the currently selected labeled samples to be retained, a classification model is trained using a confidence regularization loss function. This includes: determining the selection label change sequence and predicted confidence sequence of samples corresponding to each historical training round; the attention module calculates the weights of each sample to be retained in the current training round based on the selection label change sequence and the predicted confidence sequence; the confidence regularization loss function is weighted based on the weights to obtain a weighted loss function; the classification model is trained using the weighted loss function, and the sample selection labels and predicted confidences of the current training round are stored.

5. The method for automatically filtering and recommending labeled data according to claim 4, characterized in that: The attention module calculates the weights of each sample to be retained in the current training round based on the filter label change sequence and the prediction confidence sequence, including: determining an interference intensity coefficient based on the round number of the current training round, generating an interference signal based on the interference intensity coefficient, and adding the interference signal to the filter label change sequence and the prediction confidence sequence; wherein the interference intensity coefficient is negatively correlated with the round number; the attention module calculates the weights of each sample to be retained in the current training round based on the filter label change sequence and the prediction confidence sequence with the added interference signal.

6. An automatic filtering and recommendation system for labeled data, characterized in that, The system includes a receiving unit, a training unit, a loss calculation unit, and a filtering and recommendation unit. The receiving unit receives an initial dataset containing multiple labeled samples, each labeled sample containing feature data and corresponding labels. The training unit trains a classification model using a confidence regularization loss function based on the currently selected labeled samples to be retained. The loss calculation unit uses the trained classification model to calculate the sum of the cross-entropy loss term and the corresponding confidence regularization term for each labeled sample in the initial dataset, using this sum as the regularization loss value for that labeled sample. The filtering and recommendation unit compares the regularization loss value of each labeled sample with a corresponding filtering threshold. If the regularization loss value is less than the filtering threshold, the filtering label of that labeled sample is updated to "to be retained"; otherwise, it is updated to "to be excluded." 7. The automatic filtering and recommendation system for labeled data according to claim 6, characterized in that: The confidence regularization loss function is as follows: ;in, For cross-entropy loss term, For confidence regularization, The regularization coefficient is . This represents the expected distribution of noise labels; For the feature data of the nth labeled sample, This represents the prediction result of the classification model for the nth labeled sample. The label is the annotation label corresponding to the nth labeled sample; is a random variable for noise labels.

8. An electronic device, characterized in that: The electronic device includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the computer program, when executed by the processor, implements the method as described in any one of claims 1-5.

9. A computer storage medium, characterized in that: The computer storage medium stores a computer program that can be executed by a processor to implement the method as described in any one of claims 1-5.

10. A computer program product, characterized in that: The computer program product includes a computer program that can be executed by a processor to implement the method as described in any one of claims 1-5.