Sample labeling method and device, equipment and storage medium
Patent Information
- Application Number
- CN202510495504.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure CN120030354A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data annotation technology, and in particular to a sample annotation method, device, equipment and storage medium. Background Art
[0002] The low efficiency of manual data labeling has become a major bottleneck in the field of machine learning. In order to reduce labeling costs and improve efficiency, the industry has proposed a variety of methods, but each method has certain limitations.
[0003] For example, the method of labeling after unsupervised learning classification, that is, the samples are preliminarily classified through unsupervised learning and then labeled manually. However, the classification accuracy of unsupervised learning is limited, which may result in the workload of subsequent manual labeling not being significantly reduced; or the method of sample derivation after manual labeling, that is, manual labeling is first performed, and then the amount of labeled data is expanded by deriving the labeled samples. This method may reduce the quality of labeling because the accuracy of the derived samples is often difficult to guarantee; or the samples with larger model training images are labeled by screening (for example, the samples of the model recognition result model), but it is difficult to accurately judge the importance of the samples, and the screening process itself may increase the computing cost and labor cost. Therefore, the above methods all have the problems of high cost, reduced labeling accuracy or greater implementation difficulty to a certain extent.
[0004] Therefore, how to reduce the cost of manual labeling while ensuring labeling accuracy is a problem that needs to be solved urgently. Summary of the invention
[0005] The main purpose of this application is to provide a sample labeling method, device, equipment and storage medium, aiming to solve the technical problem of how to reduce the cost of manual labeling while ensuring the accuracy of labeling.
[0006] To achieve the above objectives, the present application proposes a sample annotation method, which includes: The target model is trained based on the labeled data set so that the target model predicts the labels of the to-be-labeled data set, and obtains pseudo labels of each to-be-labeled data set and the confidence of the pseudo labels; Based on the dataset to be labeled and the corresponding pseudo labels and confidences, an original pseudo label dataset is constructed, and the original pseudo label dataset is divided according to the confidences to obtain a plurality of original pseudo label sub-datasets with different confidence intervals; Evaluate the accuracy of the pseudo labels of each original pseudo label sub-dataset, and weight each pseudo label according to the evaluation result and the confidence level to obtain a weighted pseudo label data set; The weighted pseudo-label dataset and the label dataset are merged to train the target model according to the merged new label dataset.
[0007] In one embodiment, the step of evaluating the accuracy of the pseudo-labels of each original pseudo-label sub-dataset and weighting each pseudo-label according to the evaluation result and the confidence level to obtain a weighted pseudo-label data set includes: For any original pseudo-label sub-dataset, determine the accuracy of each pseudo-label in the original pseudo-label sub-dataset, and use the accuracy as an evaluation result of the accuracy of the pseudo-label in the original pseudo-label sub-dataset; Determining a weighted weight of a corresponding pseudo-label according to the evaluation result and each confidence level in the original pseudo-label sub-dataset, and weighting the pseudo-label based on the weighted weight; After traversing each original pseudo-label sub-dataset, a weighted pseudo-label data set is constructed based on each weighted pseudo-label.
[0008] In one embodiment, the step of determining the accuracy of each pseudo-label in the original pseudo-label sub-dataset includes: Sampling the original pseudo-label sub-dataset to obtain a sample data set; Determining the correct number of each pseudo-label in the sample data set, and calculating the pseudo-label accuracy of the sample data set based on the correct number and the number of samples in the sample data set; The pseudo-label accuracy is used as the accuracy of each pseudo-label in the original pseudo-label sub-dataset.
[0009] In one embodiment, the step of determining the weighted weight of the corresponding pseudo-label according to the evaluation result and each confidence level in the original pseudo-label sub-dataset includes: For any pseudo-label in the original pseudo-label sub-dataset, constructing data to be mapped based on the evaluation result and the confidence of the pseudo-label; The data to be mapped is mapped to a target weight, and the target weight is used as a weighted weight of the pseudo label.
[0010] In one embodiment, after the step of merging the weighted pseudo-label dataset and the label dataset to train the target model according to the merged new label dataset, the following steps are further performed: Perform performance evaluation on the trained target model and compare the performance evaluation results with the preset indicators; When the comparison result does not reach the preset index, the step of training the target model based on the labeled data set is returned to be executed based on the merged new labeled data set.
[0011] In one embodiment, the step of performing performance evaluation on the trained target model further includes: Obtaining a first labeling cost of the label data set and a second labeling cost of the weighted pseudo-label data set; Calculating a total annotation cost according to the first annotation cost and the second annotation cost, and determining an improvement in the annotation efficiency based on a ratio of the total annotation cost to the performance evaluation result; The number of samples of the to-be-annotated data set, the confidence partitioning threshold and the weighting strategy of the original pseudo-label sub-data set are adjusted according to the labeling efficiency improvement status, and after the adjustment, the step of training the target model based on the labeled data set is returned to be executed.
[0012] In one embodiment, the step of training the target model based on the label data set also includes: Obtaining an original dataset to be labeled, and extracting part of the data from the original dataset to be labeled for multiple independent labelings to obtain multiple sets of candidate label datasets; For any sample data in each group of candidate label data sets, compare the labels of the sample data to obtain the labeling evaluation result of the sample data; When the labeling evaluation result meets the standard, the sample data and the corresponding label are used as target label data; If the labeling evaluation result does not meet the standard, the sample data is re-labeled, and the sample data and the corresponding secondary label are used as target label data; After traversing each sample data, a label data set is constructed based on each target label data.
[0013] In addition, to achieve the above purpose, the present application also proposes a sample labeling device, which includes: A prediction module is used to train a target model based on a label data set so that the target model predicts the label of the to-be-labeled data set, and obtains a pseudo label of each to-be-labeled data set and a confidence level of the pseudo label; A partitioning module, configured to construct an original pseudo-label data set based on the data set to be labeled and the corresponding pseudo-labels and confidences, and to partition the original pseudo-label data set according to the confidences to obtain a plurality of original pseudo-label sub-data sets with different confidence intervals; A weighting module is used to evaluate the accuracy of the pseudo labels of each original pseudo label sub-dataset, and weight each pseudo label according to the evaluation result and the confidence level to obtain a weighted pseudo label data set; The training module is used to merge the weighted pseudo-label data set with the label data set to train the target model according to the merged new label data set.
[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes an electronic device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the sample labeling method described above.
[0015] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the sample labeling method described above are implemented.
[0016] One or more technical solutions proposed in the present invention have at least the following technical effects: The present invention firstly trains the target model based on the labeled data set, so as to train the target model with the labeled data set, so that the model can predict the unlabeled data, generate pseudo labels, and evaluate the confidence of each pseudo label, thereby providing a basis for subsequent semi-supervised learning; then, an original pseudo-label data set is constructed based on the data set to be labeled and the corresponding pseudo labels and confidences, and the original pseudo-label data set is divided according to the confidences to obtain a plurality of original pseudo-label sub-data sets with different confidence intervals, and an original pseudo-label data set is constructed by combining the unlabeled data set with the pseudo labels and confidences, and the data set is divided into a plurality of sub-data sets according to the confidences. A set of sub-datasets is prepared, each sub-dataset corresponds to a confidence interval, which is helpful for subsequent different processing and evaluation of pseudo-labels with different confidence levels; then, the pseudo-label accuracy of each original pseudo-label sub-dataset is evaluated, and each pseudo-label is weighted according to the evaluation result and the confidence level to obtain a weighted pseudo-label data set, so as to achieve the effect of increasing the weight of high-quality pseudo-labels and reducing the weight of low-quality pseudo-labels, thereby effectively improving the effect of semi-supervised learning; finally, the weighted pseudo-label data set and the label data set are merged to train the target model according to the merged new label data set to improve the performance of the model, accelerate the model training process, and improve the accuracy of the model.
[0017] In summary, the present invention integrates the data confirmed by manual sampling into the semi-supervised learning process, and introduces adaptive adjustment of the pseudo-label weights in the labeling process, so as to increase the weight of high-quality pseudo-labels and reduce the weight of low-quality pseudo-labels, so as to effectively improve the accuracy of pseudo-labels. The weighted pseudo-label dataset is then merged with the label dataset to further train the target model. While avoiding reliance on a large amount of manually labeled data, the accuracy of sample labeling is greatly improved, thereby reducing the cost of manual labeling and improving the accuracy of sample labeling. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 A schematic diagram of the process flow provided for Example 1 of the sample labeling method of this application; Figure 2 A flow chart of Example 2 of the sample labeling method of this application; Figure 3 A schematic diagram of a brief flow chart of a sample labeling method provided in Example 2 of the present application; Figure 4 A schematic diagram of the pseudo-label sampling and confirmation process of the sample annotation method provided in Example 2 of the present application; Figure 5 A schematic diagram of the task flow of annotators in the sample annotation method provided in Example 2 of the present application; Figure 6 A schematic diagram of a semi-supervised training cycle flow of a sample labeling method provided in Example 2 of the present application; Figure 7 This is a schematic diagram of the module structure of the sample labeling device according to an embodiment of the present application; Figure 8 Schematic diagram of the device structure of the hardware operating environment involved in the sample labeling method in the embodiment of the present application.
[0021] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0022] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0023] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0024] The main solution of the embodiment of the present application is: training a target model based on a label data set so that the target model predicts the label of the data set to be labeled, and obtains a pseudo-label of each data to be labeled and a confidence of the pseudo-label; constructing an original pseudo-label data set based on the data set to be labeled and the corresponding pseudo-label and confidence, and dividing the original pseudo-label data set according to the confidence to obtain a plurality of original pseudo-label sub-data sets with different confidence intervals; evaluating the pseudo-label accuracy of each original pseudo-label sub-data set, and weighting each pseudo-label according to the evaluation result and the confidence to obtain a weighted pseudo-label data set; merging the weighted pseudo-label data set and the label data set to train the target model according to the merged new label data set.
[0025] In order to reduce the cost of labeling and improve efficiency, the existing sample labeling methods all have certain limitations. For example, the method of labeling after unsupervised learning classification, that is, the samples are preliminarily classified by unsupervised learning and then manually labeled. However, the classification accuracy of unsupervised learning is limited, which may lead to the workload of subsequent manual labeling not being significantly reduced; or the method of sample derivation after manual labeling, that is, manual labeling is first performed, and then the amount of labeled data is expanded by deriving the labeled samples. This method may reduce the quality of labeling because the accuracy of the derived samples is often difficult to guarantee; or the samples with larger model training images are labeled by screening (for example, the samples of the model recognition result model), but it is difficult to accurately judge the importance of the samples, and the screening process itself may increase the computational cost and labor cost. Therefore, the above methods all have the problems of high cost, reduced labeling accuracy or greater implementation difficulty to a certain extent. Therefore, how to reduce the cost of manual labeling while ensuring the accuracy of labeling is a problem that needs to be solved urgently.
[0026] The present application provides a solution, which integrates the data confirmed by manual sampling into the semi-supervised learning process, and introduces adaptive adjustment of the pseudo-label weights in the labeling process, so as to increase the weight of high-quality pseudo-labels and reduce the weight of low-quality pseudo-labels, so as to effectively improve the accuracy of pseudo-labels. The weighted pseudo-label dataset is then merged with the label dataset to further train the target model. While avoiding reliance on a large amount of manually labeled data, it also greatly improves the accuracy of sample labeling, thereby reducing the cost of manual labeling while improving the accuracy of sample labeling.
[0027] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of realizing the above functions, a sample annotation system, etc. The following takes the sample annotation system as an example to illustrate this embodiment and the following embodiments.
[0028] Based on this, the present application embodiment provides a sample labeling method, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the sample labeling method of the present application.
[0029] In this embodiment, the sample labeling method includes steps S10 to S40: Step S10, training the target model based on the label data set so that the target model predicts the label of the to-be-labeled data set, and obtains a pseudo label of each to-be-labeled data set and a confidence level of the pseudo label; It should be noted that the labeled dataset refers to a dataset that has been manually labeled and contains correct labels; the target model refers to a machine learning model that needs to be trained to learn data features and make predictions; the unlabeled dataset refers to a dataset that has not yet been labeled and requires a model to predict labels.
[0030] It is understandable that in semi-supervised learning, usually only a small amount of labeled data is available, while a large amount of unlabeled data is not effectively utilized, which makes it difficult for the model to learn sufficient feature representations. Therefore, step S10 is performed to establish the basic feature representation of the model by training the target model with a labeled data set. This can avoid the problem that the model cannot learn effective feature representations in the absence of sufficient labeled data, thereby improving the initial performance of the model and enabling it to effectively predict unlabeled data.
[0031] For example, collect and preprocess data sets with artificial labels, i.e., labeled data sets, perform necessary data enhancement and normalization, select a machine learning model suitable for the task, such as convolutional neural network for image classification, set hyperparameters such as optimization algorithm (such as Adam), loss function (such as cross entropy loss), and learning rate, use the labeled data set to train the target model, and use the early stopping method to monitor the performance of the validation set to prevent overfitting. Then use the trained model to predict the unlabeled data set to obtain the pseudo label and its confidence for each data to be labeled.
[0032] Step S20, constructing an original pseudo-label dataset based on the dataset to be labeled and the corresponding pseudo-labels and confidences, and dividing the original pseudo-label dataset according to the confidences to obtain a plurality of original pseudo-label sub-datasets with different confidence intervals; It should be noted that pseudo-label refers to the label predicted by the model for unlabeled data; confidence refers to the model's confidence in the predicted label, which can usually be expressed as a probability value; the original pseudo-label dataset refers to the dataset containing all the data to be labeled, their pseudo-labels and confidence levels; the original pseudo-label sub-dataset refers to the pseudo-label dataset with different confidence intervals obtained according to the confidence level, such as the original pseudo-label sub-dataset with high confidence, the original pseudo-label sub-dataset with medium confidence, and the original pseudo-label sub-dataset with low confidence.
[0033] It is understandable that in self-training or pseudo-labeling methods, all pseudo-labels are usually used directly for training, but the quality of these pseudo-labels varies, which may lead to a decline in model performance. Therefore, step S20 is performed. By dividing the pseudo-label sub-dataset according to the confidence level, it is possible to avoid using all pseudo-labels indiscriminately for training, resulting in a decline in model performance. Therefore, high-quality pseudo-labels can be used more specifically to improve the effect of model training.
[0034] Exemplarily, the data to be annotated is combined with its corresponding pseudo-label and confidence to form an original pseudo-label dataset, and a confidence threshold is set to divide the original pseudo-label dataset into multiple confidence subsets. For example, a confidence higher than 0.9 is a high confidence subset, 0.7-0.9 is a medium confidence subset, and a confidence lower than 0.7 is a low confidence subset.
[0035] Step S30, evaluating the accuracy of the pseudo labels of each original pseudo label sub-data set, and weighting each pseudo label according to the evaluation result and the confidence level to obtain a weighted pseudo label data set; It should be noted that the evaluation result refers to the result obtained after evaluating the accuracy of the pseudo-label, which may include indicators such as the accuracy of the pseudo-label; the weighted pseudo-label dataset refers to the pseudo-label dataset after weighted processing, and the weight of the pseudo-label reflects its accuracy and confidence.
[0036] It is understandable that since the accuracy of the pseudo-labels is usually not evaluated when using pseudo-labels for training, the model may learn wrong features. Therefore, step S30 is performed to evaluate the accuracy of the pseudo-labels and weight them, so as to avoid using inaccurate pseudo-labels to train the model, resulting in slow improvement or even deterioration of model performance, thereby improving the effectiveness of the pseudo-labels and further improving the training effect of the model.
[0037] For example, a small amount of annotated data is used as a validation set, the consistency between the pseudo-labels of each confidence subset and the true labels is calculated, the accuracy of the pseudo-labels is evaluated, and the pseudo-labels are weighted according to the accuracy evaluation and confidence. Pseudo-labels with high accuracy and high confidence are given higher weights, and vice versa. The weighted pseudo-labels are then combined with the corresponding data to form a weighted pseudo-label dataset.
[0038] In a feasible implementation, step S30 may include steps S31 to S33: Step S31, for any original pseudo-label sub-dataset, determining the accuracy of each pseudo-label in the original pseudo-label sub-dataset, and using the accuracy as an evaluation result of the accuracy of the pseudo-label in the original pseudo-label sub-dataset; It is understandable that since existing semi-supervised learning methods often do not consider the accuracy of pseudo-labels and directly use all pseudo-labels for training, this may cause the model to learn wrong information. Therefore, step S31 is performed to identify and reduce the negative impact of inaccurate pseudo-labels on model training by evaluating the accuracy of pseudo-labels, thereby improving the training efficiency and accuracy of the model and reducing the noise introduced by erroneous pseudo-labels.
[0039] In a feasible implementation manner, the step of determining the accuracy of each pseudo-label in the original pseudo-label sub-dataset in step S31 may include steps S311 to S313: Step S311, sampling the original pseudo-label sub-dataset to obtain a sample data set; It should be noted that the sample dataset refers to a part of the data randomly sampled from the original pseudo-label sub-dataset, which is used to represent the entire dataset.
[0040] It is understandable that directly processing the entire original pseudo-label sub-dataset will result in excessive computation and low efficiency, so step S311 is performed to avoid the problem of wasting computing resources and time consumption by sampling the original pseudo-label sub-dataset, thereby improving the effect of determining the accuracy of the pseudo-label.
[0041] Step S312, determining the correct number of each pseudo label in the sample data set, and calculating the pseudo label accuracy of the sample data set based on the correct number and the number of samples in the sample data set; Exemplarily, the correct number of each pseudo-label in the sample data set is determined by manual inspection, and the ratio of the correct number to the number of samples is used as the pseudo-label accuracy of the sample data set.
[0042] Step S313: Using the pseudo-label accuracy as the accuracy of each pseudo-label in the original pseudo-label sub-dataset.
[0043] In this implementation, by sampling the original pseudo-label sub-dataset to obtain a representative sample data set, and then determining the correct number of each pseudo-label in the sample data set and calculating the pseudo-label accuracy, and finally extending the accuracy to the entire original pseudo-label sub-dataset, it effectively avoids the computational efficiency and evaluation bias problems caused by directly processing the full data set, and achieves the effect of improving computational efficiency, accurately evaluating the quality of pseudo-labels, and improving model training effects and generalization capabilities.
[0044] Step S32, determining a weighted weight of the corresponding pseudo-label according to the evaluation result and each confidence level in the original pseudo-label sub-dataset, and weighting the pseudo-label based on the weighted weight; It is understandable that, since existing methods usually do not distinguish between the confidence and accuracy of pseudo-labels, all pseudo-labels have the same weight in training, which may reduce the performance of the model. Therefore, step S32 is performed to determine the weighted weights according to the confidence and accuracy, which can increase the influence of pseudo-labels with high confidence and high accuracy, and reduce the influence of pseudo-labels with low confidence and low accuracy, thereby enhancing the model's learning of high-quality pseudo-labels and improving the overall performance and generalization ability of the model.
[0045] In a feasible implementation manner, the step of determining the weighted weight of the corresponding pseudo-label according to the evaluation result and each confidence level in the original pseudo-label sub-dataset in step S32 may include steps S321-S322: Step S321, for any pseudo-label in the original pseudo-label sub-dataset, construct the data to be mapped based on the evaluation result and the confidence of the pseudo-label; It should be noted that the data to be mapped refers to a comprehensive indicator that combines the pseudo-label evaluation results and confidence, which is used to quantify the quality and reliability of the pseudo-label.
[0046] It is understandable that since the weight distribution of pseudo-labels often lacks comprehensive evaluation, relying solely on confidence or evaluation results may lead to unfair weight distribution and technical deviation. Therefore, performing step S321 can avoid the problem of inaccurate weight distribution caused by relying solely on evaluation results or confidence, thereby more accurately reflecting the reliability of pseudo-labels and improving the accuracy of weight distribution.
[0047] Exemplarily, for each pseudo-label in the original pseudo-label sub-dataset, its evaluation result and confidence are first minimum-maximum standardized, that is, they are scaled to a range between 0 and 1. The purpose of standardization is to make the evaluation result and confidence on the same scale to facilitate subsequent calculations. Then the product of the two is calculated as the data to be mapped, so that the data to be mapped comprehensively considers the evaluation result and confidence, and can more accurately reflect the quality of the pseudo-label. For example, assuming that the evaluation result of a pseudo-label is 0.8 and the confidence is 0.9, after standardization, their product is 0.72, that is, the data to be mapped is 0.72.
[0048] Step S322: Map the data to be mapped to a target weight, and use the target weight as a weighted weight of the pseudo label.
[0049] It should be noted that the target weight refers to the weight value finally assigned to the pseudo-label after the mapping process, which is used to adjust the influence of the pseudo-label on the model training process during model training.
[0050] It is understandable that since weight assignment is often subjective and inaccurate and lacks data-driven basis, step S322 is performed to avoid the inaccuracy of subjective assignment by mapping the data to be mapped to the target weight, thereby achieving objective and precise weight distribution.
[0051] Exemplarily, all the data to be mapped of the pseudo-label are converted into target weights through a linear mapping function. First, determine the minimum min x and maximum max x of all the data to be mapped. Then, for each data to be mapped x, the target weight y is calculated by the formula y=(x-min x) / (max x-min x+c), where c is a very small positive number, such as 1×10^-8, which is used to prevent the denominator from being zero and avoid a significant impact on the mapping result. For example, assuming that the data to be mapped is [0.72, 0.85, 0.60, 0.90], then min x=0.60, max x=0.90. For the case of x=0.72, y=0.4, that is, the target weight is 0.4, which can be used as the weighted weight of the pseudo-label.
[0052] In this implementation, the data to be mapped is constructed by combining the evaluation results and the confidence of the pseudo-label, and mapped to the target weight, thereby avoiding the one-sidedness and subjectivity of assigning weights to a single indicator in the existing scheme, achieving a comprehensive reflection of the quality of the pseudo-label and an objective and precise weight allocation, and further improving the model training effect and generalization ability.
[0053] Step S33: after traversing each original pseudo-label sub-dataset, construct a weighted pseudo-label data set based on each weighted pseudo-label.
[0054] It is understandable that since indiscriminately merging all pseudo-labels into the label data set may cause model overfitting or performance degradation, step S33 is performed to form a weighted pseudo-label data set based on the weighted pseudo-labels, which can better balance the label data and pseudo-label data and avoid the model from being disturbed by low-quality pseudo-labels, thereby achieving effective fusion of label data and pseudo-label data and improving the training effect and generalization ability of the model.
[0055] In this implementation, by evaluating the accuracy and confidence of the pseudo-labels and assigning weighted weights to each pseudo-label based on these evaluation results, a weighted pseudo-label dataset is constructed, thereby avoiding the problems of inefficient model training and insufficient accuracy caused by the uneven quality of pseudo-labels in existing semi-supervised learning, thereby improving the model training efficiency and accuracy, enhancing the model's ability to learn high-quality pseudo-labels, and improving the overall performance and generalization ability of the model.
[0056] Step S40: merging the weighted pseudo-label dataset and the label dataset to train the target model according to the merged new label dataset.
[0057] It should be noted that the merged new labeled dataset refers to the dataset obtained by merging the weighted pseudo-label dataset and the labeled dataset, which is used to train the target model.
[0058] It is understandable that, since in semi-supervised learning, usually only labeled datasets or pseudo-labeled datasets are used for training, and the combination of the two is not fully utilized, step S40 is performed to merge the weighted pseudo-labeled dataset and the labeled dataset to avoid the problem of insufficient data volume due to only using the labeled dataset or low data quality due to only using the pseudo-labeled dataset during the model training process, thereby improving the training effect and performance of the model by increasing the amount of reliable training data.
[0059] For example, the weighted pseudo-label dataset is merged with the original label dataset to form a new label dataset, in which each sample has a weight. The target model is trained using a weighted loss function (such as weighted cross entropy loss) so that the model pays more attention to samples with high weights. At the same time, pseudo-label generation and model training can be iteratively performed to gradually improve model performance.
[0060] This embodiment provides a sample labeling method, which integrates data confirmed by manual sampling into the semi-supervised learning process and introduces adaptive adjustment of pseudo-label weights in the labeling process to achieve the effect of increasing the weight of high-quality pseudo-labels and reducing the weight of low-quality pseudo-labels, so as to effectively improve the accuracy of pseudo-labels. The weighted pseudo-label dataset is then merged with the label dataset to further train the target model. This avoids reliance on a large amount of manually labeled data and greatly improves the accuracy of sample labeling, thereby reducing the cost of manual labeling while improving the accuracy of sample labeling.
[0061] In a feasible implementation manner, before step S10, steps S01 to S05 may also be included: Step S01, obtaining an original dataset to be labeled, and extracting part of the data from the original dataset to be labeled for multiple independent labeling to obtain multiple sets of candidate label datasets; It should be noted that the original dataset to be labeled refers to a dataset that has not yet been labeled and needs to be labeled for subsequent machine learning tasks; the candidate label dataset refers to a dataset of multiple labeled versions obtained by multiple independent labeling of the original dataset to be labeled.
[0062] It is understandable that, since existing solutions usually rely on single-subject labeling, this may lead to labeling deviations and inconsistencies. Therefore, performing step S01 through multiple independent labeling can avoid the subjective bias and errors of a single labeler, improve the diversity and reliability of the labeled data, and thus provide higher quality data for subsequent model training.
[0063] Exemplarily, a small portion of data is randomly selected from a large amount of unlabeled original data sets to be labeled, and labeled by two groups of labelers to form two sets of candidate label data sets of different versions.
[0064] Step S02, for any sample data in each group of candidate label data sets, compare the labels of the sample data to obtain the labeling evaluation result of the sample data; It should be noted that sample data refers to a single data instance in a dataset, such as an image, a piece of text, or an audio clip; the annotation evaluation result refers to the conclusion obtained after comparing and evaluating multiple annotation results of the sample data, which is used to determine whether the annotation quality meets the standards.
[0065] It is understandable that due to the lack of objective evaluation of annotation quality, low-quality annotation data is used for training, thus affecting model performance. Therefore, step S02 is performed to evaluate the consistency and accuracy of annotations by comparing the labels of multiple annotators, avoid the impact of low-quality annotation data on model training, achieve quality control of annotation data, and ensure that only high-quality annotation data is used for subsequent processing.
[0066] Exemplarily, when the candidate label data sets are two sets of label data of different versions, the label data are compared to obtain the consistency of the labels, and the data are used as the labeling evaluation results of the sample data.
[0067] Step S03, when the labeling evaluation result meets the standard, the sample data and the corresponding label are used as target label data; Step S04, if the labeling evaluation result does not meet the standard, re-label the sample data, and use the sample data and the corresponding secondary label as target label data; It is understandable that there is often a lack of effective processing methods for sample data that do not meet the labeling evaluation standards, which leads to these data being discarded or incorrectly used for training. Therefore, step S04 is performed to improve the low-quality labeled data through secondary labeling, ensuring that all sample data meet certain quality standards.
[0068] Exemplarily, when the labeling evaluation results do not meet the standards, the sample data is discussed or experts are invited to arbitrate to ensure the quality of the initial data set, and the sample data after secondary labeling is used as the target label data.
[0069] Step S05, after traversing each sample data, construct a label data set based on each target label data.
[0070] In this implementation, part of the data is extracted from the original data set to be labeled and labeled multiple times independently, and the labels are compared to obtain the labeling evaluation results. If the evaluation results do not meet the standards, secondary labeling is performed to finally construct a labeled data set, thereby avoiding the deviations and errors that may be introduced by a single labeling, improving the accuracy and consistency of the labeled data, and achieving the construction of a higher quality labeled data set.
[0071] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated in the following. Figure 2 After step S40, the sample labeling method further includes steps S50 to S60: Step S50, performing performance evaluation on the trained target model and comparing the performance evaluation result with the preset index; It should be noted that the performance evaluation result refers to the performance of the model on the test set, which can be quantified by metrics such as accuracy, recall, F1-score, etc.
[0072] Step S60, in the case where the comparison result does not reach the preset index, return to execute the step of training the target model based on the label data set based on the new combined label data set.
[0073] It can be understood that since the weighted pseudo-label data set under one optimization may be difficult to achieve good model training results, step S60 is performed. By using the new combined label data set, the model can obtain more training information and higher annotation quality, thereby improving its training effect and performance, avoiding the problem of lack of effective improvement measures when the model performance does not meet the standard, ensuring that the model can reach the preset performance index, and improving its accuracy and reliability.
[0074] In this embodiment, by implementing performance evaluation and comparing it with the preset index, and retraining the model with the new combined label data set when the index is not reached, the problem that the performance of the model does not meet the standard and cannot be effectively improved in the prior art is avoided, realizing the continuous optimization and improvement of the model performance, and ensuring the accuracy and reliability of the model.
[0075] In a feasible implementation manner, after the step of performing performance evaluation on the trained target model in step S50, the sample annotation method further includes steps S100 to S300: Step S100, obtain the first annotation cost of the label data set and the second annotation cost of the weighted pseudo-label data set; It should be noted that the first annotation cost refers to the resources, time, and money required to annotate the real label data set, which can be jointly measured by the number of label data, the annotation time of a single sample, and the salary of the annotation personnel. The second annotation cost refers to the computing resources and time required to generate the pseudo-label data set, which can be jointly measured by the number of pseudo-labels confirmed by sampling, the confirmation time of a single sample, and the salary of the annotation personnel.
[0076] Step S200, calculate the total annotation cost according to the first annotation cost and the second annotation cost, and determine the improvement status of the annotation efficiency based on the ratio of the total annotation cost to the performance evaluation result; It should be noted that the improvement status of the annotation efficiency refers to the improvement of the benefit of the annotation work by comparing the total annotation cost and the performance evaluation result.
[0077] It is understandable that since the existing solutions do not combine cost and performance to evaluate efficiency, it is impossible to measure the actual benefits of the labeling work. Therefore, step S200 is performed to determine whether the labeling efficiency is improved by calculating the ratio of the total labeling cost to the performance evaluation result, thereby avoiding the limitation of focusing only on performance or cost, thereby achieving comprehensive optimization of cost and performance.
[0078] Step S300, adjusting the number of samples of the to-be-annotated data set, the confidence partition threshold and the weighting strategy of the original pseudo-label sub-data set according to the labeling efficiency improvement status, and returning to execute the step of training the target model based on the labeled data set after the adjustment.
[0079] It should be noted that the confidence division threshold refers to the threshold that divides pseudo labels into different confidence intervals, which is used to screen pseudo labels of different qualities; the weighted strategy refers to the method of assigning different weights to different data sets in model training, which is used to balance the impact of real labels and pseudo labels.
[0080] It is understandable that the fixed number of samples in the dataset to be labeled, the confidence division threshold or weighting strategy of the original pseudo-label sub-dataset may cause the model to be unable to adapt to different data distributions and labeling qualities. Therefore, step S300 is performed to avoid the model performance bottleneck caused by fixed parameters by adjusting the number of samples, confidence threshold and weighting strategy according to the improvement of labeling efficiency, thereby achieving the flexibility and effectiveness of model training.
[0081] For example, firstly, based on the calculated labeling efficiency improvement, evaluate whether the current labeling strategy is effective: If the labeling efficiency is improved well, that is, the model performance is significantly improved and cost-effective, then you can consider increasing the number of samples in the dataset to be labeled to further improve the generalization ability and performance of the model. At the same time, lower the confidence partition threshold of the original pseudo-label sub-dataset to allow more pseudo-labels with lower confidence to enter the training process, thereby increasing the diversity and quantity of training data. In addition, you can increase the weight of pseudo-labels so that the model pays more attention to the information of these pseudo-labels during training.
[0082] If the labeling efficiency is not improved well, that is, the model performance is not significantly improved or the cost-effectiveness is low, then the number of samples in the dataset to be labeled needs to be reduced to reduce the labeling cost. At the same time, the confidence partition threshold of the original pseudo-label sub-dataset is increased to ensure that only pseudo-labels with higher confidence are used for training, thereby improving the quality of the training data. In addition, the weight of the pseudo-label can be reduced so that the model pays more attention to the real labeled data during the training process.
[0083] The adjusted parameters will be applied to the dataset to be labeled, the original pseudo-labeled sub-dataset, and the weighting strategy. Then, the step of training the target model based on the labeled dataset is returned to optimize the model performance. By continuously adjusting and optimizing these parameters, the labeling efficiency can be improved and the model performance can be optimized.
[0084] In this implementation, the annotation costs of the labeled dataset and the weighted pseudo-labeled dataset are calculated to evaluate the improvement in annotation efficiency. The number of samples in the dataset to be labeled, the confidence partition threshold and the weighting strategy of the original pseudo-labeled sub-dataset are adjusted according to the status. This avoids the inefficiency and cost waste caused by the inability to dynamically adjust the annotation strategy in traditional methods, thereby achieving the effect of reducing annotation costs and improving annotation efficiency while ensuring model performance.
[0085] For example, to help understand the implementation process of the sample labeling method obtained by combining this embodiment with the above-mentioned embodiment 1, please refer to Figure 3 , Figure 3 A brief flow chart of a sample annotation method is provided, specifically: First, accurately annotate the initial dataset L, i.e., the label dataset, and train the initial model M0, i.e., the target model, based on L, so as to use M0 to predict the unlabeled data U, i.e., the data to be labeled, and obtain the pseudo-label P of each unlabeled data and the confidence of the pseudo-label. Then, the pseudo-labels are stratified based on the confidence to obtain multiple original pseudo-label sub-datasets with different confidence intervals. Then, the pseudo-labels of each confidence interval are manually confirmed, and the accuracy of the pseudo-labels of each confidence interval is evaluated. Then, the pseudo-labels are weighted according to the confidence and accuracy. Finally, L and the weighted P are merged as new training data to iteratively train the model M1 using the new training data. At the same time, it is judged whether the stopping condition of the iteration is reached, based on whether the model performance target is reached or the upper limit of the labeling budget is reached.
[0086] For further information, please refer to Figure 4 , Figure 4 A schematic diagram of the pseudo-label sampling confirmation process of a sample annotation method is provided, specifically: First, input the pseudo-label dataset P, i.e. the original pseudo-label dataset, and then divide P into multiple intervals according to the confidence level, such as high confidence area, medium confidence area, and low confidence area. Then randomly select one or more samples from each confidence interval, and give the samples to the labeler to confirm whether the pseudo-label is correct. If it is correct, record the confirmation result. If it is incorrect, the labeler corrects the pseudo-label and records the corrected label, i.e. the target label data. After ensuring that all intervals have completed sampling confirmation, evaluate the accuracy of the pseudo-labels in each confidence interval, and adjust the pseudo-label weights or directly screen them out according to the accuracy.
[0087] For further information, please refer to Figure 5 , Figure 5 A schematic diagram of the task flow of annotators in a sample annotation method is provided, specifically: When the labeler receives a sample to be confirmed and the corresponding pseudo-label, he / she checks the sample and pseudo-label, and records the correct confirmation result if the pseudo-label is correct; if the pseudo-label is incorrect, he / she corrects the pseudo-label and records the confirmation result of the corrected label.
[0088] For further information, please refer to Figure 6 , Figure 6 A schematic diagram of a semi-supervised training cycle flow of a sample labeling method is provided, specifically: The semi-supervised learning system uses the labeled data set L to train the initial model so that the model predicts U, generates pseudo labels, and has pseudo labels and confidence. Then, stratification and sampling are performed according to the confidence, and the sampled samples and pseudo labels are given to manual confirmation, and the confirmation results are returned after manual confirmation. Then, the quality of the pseudo labels is evaluated, and each pseudo label is weighted according to the evaluation results. Finally, the model is iteratively trained using L and the weighted pseudo labels to output the final model.
[0089] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the sample annotation method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0090] This application also provides a sample marking device, please refer to Figure 7 , the sample labeling device comprises: The prediction module 10 is used to train the target model based on the label data set so that the target model predicts the label of the to-be-labeled data set, and obtains the pseudo label of each to-be-labeled data and the confidence of the pseudo label; A partitioning module 20 is used to construct an original pseudo-label dataset based on the dataset to be labeled and the corresponding pseudo-labels and confidences, and to partition the original pseudo-label dataset according to the confidences to obtain a plurality of original pseudo-label sub-datasets with different confidence intervals; A weighting module 30 is used to evaluate the accuracy of the pseudo labels of each original pseudo label sub-data set, and weight each pseudo label according to the evaluation result and the confidence level to obtain a weighted pseudo label data set; The training module 40 is used to merge the weighted pseudo-label data set and the label data set to train the target model according to the merged new label data set.
[0091] Optionally, the weighting module 30 is further used for: For any original pseudo-label sub-dataset, determine the accuracy of each pseudo-label in the original pseudo-label sub-dataset, and use the accuracy as an evaluation result of the accuracy of the pseudo-label in the original pseudo-label sub-dataset; Determining a weighted weight of a corresponding pseudo-label according to the evaluation result and each confidence level in the original pseudo-label sub-dataset, and weighting the pseudo-label based on the weighted weight; After traversing each original pseudo-label sub-dataset, a weighted pseudo-label data set is constructed based on each weighted pseudo-label.
[0092] Optionally, the weighting module 30 is further used for: Sampling the original pseudo-label sub-dataset to obtain a sample data set; Determining the correct number of each pseudo-label in the sample data set, and calculating the pseudo-label accuracy of the sample data set based on the correct number and the number of samples in the sample data set; The pseudo-label accuracy is used as the accuracy of each pseudo-label in the original pseudo-label sub-dataset.
[0093] Optionally, the weighting module 30 is further used for: For any pseudo-label in the original pseudo-label sub-dataset, constructing data to be mapped based on the evaluation result and the confidence of the pseudo-label; The data to be mapped is mapped to a target weight, and the target weight is used as a weighted weight of the pseudo label.
[0094] Optionally, the evaluation module 50 in the sample labeling device is used to: Perform performance evaluation on the trained target model and compare the performance evaluation results with the preset indicators; When the comparison result does not reach the preset index, the step of training the target model based on the labeled data set is returned to be executed based on the merged new labeled data set.
[0095] Optionally, the evaluation module 50 is further configured to: Obtaining a first labeling cost of the label data set and a second labeling cost of the weighted pseudo-label data set; Calculating a total annotation cost according to the first annotation cost and the second annotation cost, and determining an improvement in the annotation efficiency based on a ratio of the total annotation cost to the performance evaluation result; The number of samples of the to-be-annotated data set, the confidence partitioning threshold and the weighting strategy of the original pseudo-label sub-data set are adjusted according to the labeling efficiency improvement status, and after the adjustment, the step of training the target model based on the labeled data set is returned to be executed.
[0096] Optionally, the preparation module 60 in the sample labeling device is used to: Obtaining an original dataset to be labeled, and extracting part of the data from the original dataset to be labeled for multiple independent labelings to obtain multiple sets of candidate label datasets; For any sample data in each group of candidate label data sets, compare the labels of the sample data to obtain the labeling evaluation result of the sample data; When the labeling evaluation result meets the standard, the sample data and the corresponding label are used as target label data; If the labeling evaluation result does not meet the standard, the sample data is re-labeled, and the sample data and the corresponding secondary label are used as target label data; After traversing each sample data, a label data set is constructed based on each target label data.
[0097] The sample annotation device provided by the present application adopts the sample annotation method in the above embodiment, which can solve the technical problem of how to reduce the cost of manual annotation while ensuring the accuracy of annotation. Compared with the prior art, the beneficial effects of the sample annotation device provided by the present application are the same as those of the sample annotation method provided by the above embodiment, and other technical features in the sample annotation device are the same as those disclosed in the above embodiment method, which will not be repeated here.
[0098] The present application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the sample labeling method in the above-mentioned embodiment 1.
[0099] Reference below Figure 8 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present application. The electronic devices in the embodiments of the present application may include but are not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions: tablet computers), PMPs (Portable Media Players: portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0100] like Figure 8As shown, the electronic device may include a processing device 1001 (e.g., a central processing unit, a graphics processor, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 to a random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the electronic device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other via a bus 1005. An input / output interface 1006 is also connected to the bus. Generally, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although the electronic device with various systems is shown in the figure, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have instead.
[0101] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0102] The electronic device provided by the present application adopts the sample annotation method in the above embodiment, which can solve the technical problem of how to reduce the cost of manual annotation while ensuring the accuracy of annotation. Compared with the prior art, the beneficial effects of the electronic device provided by the present application are the same as the beneficial effects of the sample annotation method provided by the above embodiment, and the other technical features in the electronic device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0103] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0104] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0105] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the sample labeling method in the above-mentioned embodiment.
[0106] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM: Random Access Memory), a read-only memory (ROM: Read Only Memory), an erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency: Radio Frequency), etc., or any suitable combination of the above.
[0107] The computer-readable storage medium may be included in the electronic device, or may exist independently without being installed in the electronic device.
[0108] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by an electronic device, the electronic device: trains a target model based on a label data set so that the target model predicts the label of the data set to be labeled, and obtains a pseudo label of each data to be labeled and a confidence of the pseudo label; constructs an original pseudo label data set based on the data set to be labeled and the corresponding pseudo label and confidence, and divides the original pseudo label data set according to the confidence to obtain a plurality of original pseudo label sub-data sets with different confidence intervals; evaluates the pseudo label accuracy of each original pseudo label sub-data set, and weights each pseudo label according to the evaluation result and the confidence to obtain a weighted pseudo label data set; merges the weighted pseudo label data set and the label data set to train the target model according to the merged new label data set.
[0109] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0110] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0111] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.
[0112] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned sample annotation method, and can solve the technical problem of how to reduce the cost of manual annotation while ensuring the accuracy of annotation. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the sample annotation method provided in the above-mentioned embodiment, and will not be repeated here.
[0113] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A sample labeling method, characterized in that: The sample annotation method comprises: The target model is trained based on the labeled data set so that the target model predicts the labels of the to-be-labeled data set, and obtains pseudo labels of each to-be-labeled data set and the confidence of the pseudo labels; Based on the dataset to be labeled and the corresponding pseudo labels and confidences, an original pseudo label dataset is constructed, and the original pseudo label dataset is divided according to the confidences to obtain a plurality of original pseudo label sub-datasets with different confidence intervals; Evaluate the accuracy of the pseudo labels of each original pseudo label sub-dataset, and weight each pseudo label according to the evaluation result and the confidence level to obtain a weighted pseudo label data set; The weighted pseudo-label dataset and the label dataset are merged to train the target model according to the merged new label dataset.
2. The sample labeling method according to claim 1, characterized in that: The step of evaluating the accuracy of the pseudo labels of each original pseudo label sub-data set, and weighting each pseudo label according to the evaluation result and the confidence level to obtain a weighted pseudo label data set comprises: For any original pseudo-label sub-dataset, determine the accuracy of each pseudo-label in the original pseudo-label sub-dataset, and use the accuracy as an evaluation result of the accuracy of the pseudo-label in the original pseudo-label sub-dataset; Determining a weighted weight of a corresponding pseudo-label according to the evaluation result and each confidence level in the original pseudo-label sub-dataset, and weighting the pseudo-label based on the weighted weight; After traversing each original pseudo-label sub-dataset, a weighted pseudo-label data set is constructed based on each weighted pseudo-label.
3. The sample labeling method according to claim 2, characterized in that: The step of determining the accuracy of each pseudo label in the original pseudo label sub-dataset comprises: Sampling the original pseudo-label sub-dataset to obtain a sample data set; Determining the correct number of each pseudo-label in the sample data set, and calculating the pseudo-label accuracy of the sample data set based on the correct number and the number of samples in the sample data set; The pseudo-label accuracy is used as the accuracy of each pseudo-label in the original pseudo-label sub-dataset.
4. The sample labeling method according to claim 2, characterized in that: The step of determining the weighted weight of the corresponding pseudo label according to the evaluation result and each confidence level in the original pseudo label sub-dataset comprises: For any pseudo-label in the original pseudo-label sub-dataset, constructing data to be mapped based on the evaluation result and the confidence of the pseudo-label; The data to be mapped is mapped to a target weight, and the target weight is used as a weighted weight of the pseudo label.
5. The sample labeling method according to claim 1, characterized in that: After the step of merging the weighted pseudo-label dataset and the label dataset to train the target model according to the merged new label dataset, the method further includes: Perform performance evaluation on the trained target model and compare the performance evaluation results with the preset indicators; When the comparison result does not reach the preset index, the step of training the target model based on the labeled data set is returned to be executed based on the merged new labeled data set.
6. The sample labeling method according to claim 5, characterized in that: The step of evaluating the performance of the trained target model further includes: Obtaining a first labeling cost of the label data set and a second labeling cost of the weighted pseudo-label data set; Calculating a total annotation cost according to the first annotation cost and the second annotation cost, and determining an improvement in the annotation efficiency based on a ratio of the total annotation cost to the performance evaluation result; The number of samples of the to-be-annotated data set, the confidence partitioning threshold and the weighting strategy of the original pseudo-label sub-data set are adjusted according to the labeling efficiency improvement status, and after the adjustment, the step of training the target model based on the labeled data set is returned to be executed.
7. The sample labeling method according to claim 1, characterized in that: The step of training the target model based on the label data set also includes: Obtaining an original dataset to be labeled, and extracting part of the data from the original dataset to be labeled for multiple independent labelings to obtain multiple sets of candidate label datasets; For any sample data in each group of candidate label data sets, compare the labels of the sample data to obtain the labeling evaluation result of the sample data; When the labeling evaluation result meets the standard, the sample data and the corresponding label are used as target label data; If the labeling evaluation result does not meet the standard, the sample data is re-labeled, and the sample data and the corresponding secondary label are used as target label data; After traversing each sample data, a label data set is constructed based on each target label data.
8. A sample labeling device, characterized in that: The sample marking device comprises: A prediction module is used to train a target model based on a label data set so that the target model predicts the label of the to-be-labeled data set, and obtains a pseudo label of each to-be-labeled data set and a confidence level of the pseudo label; A partitioning module, configured to construct an original pseudo-label data set based on the data set to be labeled and the corresponding pseudo-labels and confidences, and to partition the original pseudo-label data set according to the confidences to obtain a plurality of original pseudo-label sub-data sets with different confidence intervals; A weighting module is used to evaluate the accuracy of the pseudo labels of each original pseudo label sub-dataset, and weight each pseudo label according to the evaluation result and the confidence level to obtain a weighted pseudo label data set; The training module is used to merge the weighted pseudo-label data set with the label data set to train the target model according to the merged new label data set.
9. An electronic device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the sample labeling method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the sample labeling method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Semi-supervised learning method based on pseudo label weighting
CN112232416A
Small-sample NL2SQL method based on semi-supervised learning and meta-learning
CN114817307A
Cited By
Model training method, device and equipment for overflow concentration in ore grinding process and medium
CN120670852A
Automatic training data screening method and device, electronic equipment and storage medium
CN120911634A
Pseudo tag optimization method and device, storage medium and computer equipment
CN121505389A