Sample learning methods, data annotation devices, electronic devices, and media
By constructing a semi-supervised learning model and conducting multiple iterations of verification, the problem of time-consuming and labor-intensive data annotation was solved, achieving efficient and accurate data annotation and reducing labor costs.
Patent Information
- Application Number
- CN202210095511.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-01-26
AI Technical Summary
Existing technologies for data annotation are time-consuming, labor-intensive, and inefficient. Manual annotation is costly and makes it difficult to achieve timely and accurate data annotation.
By constructing a semi-supervised learning model, accuracy is verified using labeled and unlabeled datasets. Combined with manual review, the model is iterated multiple times until its accuracy meets the preset conditions, and then the unlabeled dataset is updated.
It reduced manual annotation and review processes, saved labor costs, improved the efficiency and accuracy of data annotation, and enabled model updates and optimizations.
Smart Images

Figure CN114418096B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data annotation technology, and in particular to a sample learning method, data annotation device, electronic device, and medium. Background Technology
[0002] Typically, as people's demands for product equipment increase, when using product equipment for data annotation, users often hope to maintain both the timeliness and accuracy of data annotation in the product equipment.
[0003] Currently, in terms of data annotation technology, in order to accurately annotate data, it is often based on predefined annotation specifications and label systems, and manual annotation is used to annotate data line by line or frame by frame. This results in manual annotation being time-consuming, labor-intensive, and inefficient, with a single person only able to annotate a fixed amount of data per day, leading to high costs. Summary of the Invention
[0004] The first aspect of this application provides a sample learning method, which includes acquiring a labeled dataset and an unlabeled dataset; constructing a network model based on the labeled dataset; inputting the labeled dataset and the unlabeled dataset into the network model to verify the accuracy of the unlabeled dataset using the network model, thereby obtaining the model accuracy; when the model accuracy does not meet a preset condition, selecting several samples from the unlabeled dataset and manually reviewing the pre-labeled data corresponding to the several samples to update the unlabeled dataset; and repeating the step of inputting the labeled dataset and the unlabeled dataset into the network model again to verify the accuracy of the unlabeled dataset using the network model, through multiple iterations until the model accuracy meets the preset condition.
[0005] The second aspect of this application provides a data annotation method, including: obtaining an unlabeled dataset; and calling a network model as provided in the first aspect of this application to annotate the unlabeled dataset to obtain labeled data.
[0006] A third aspect of this application provides a data annotation device, which includes:
[0007] The acquisition module is used to acquire labeled and unlabeled datasets;
[0008] Build modules are used to construct network models based on labeled datasets;
[0009] The processing module is used to input labeled and unlabeled datasets into the network model, so as to use the network model to verify the accuracy of the unlabeled dataset and obtain the model accuracy.
[0010] The processing module is also used to select several samples from the unlabeled dataset when the model accuracy does not meet the preset conditions, and to manually review the pre-labeled data corresponding to the several samples in order to update the unlabeled dataset.
[0011] The processing module is also used to re-execute the step of inputting the labeled and unlabeled datasets into the network model to verify the accuracy of the unlabeled dataset using the network model, through multiple iterations until the model accuracy meets the preset conditions.
[0012] A fourth aspect of this application provides an electronic device comprising: a processor and a memory, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in the first or second aspect.
[0013] A fifth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods of the first or second aspect of this application.
[0014] The beneficial effects of this application are as follows: This application constructs a network model based on a labeled dataset, and then inputs the labeled dataset and the unlabeled dataset into the network model, enabling the unlabeled dataset to learn from samples in the network model based on the labeled dataset, thereby updating the unlabeled dataset. This reduces the need for manual labeling or review, saving manpower costs. Furthermore, by setting preset conditions for model accuracy, multiple iterations are made until the model accuracy meets the preset conditions, further enabling the updating of the network model. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart illustrating an embodiment of the sample learning method of this application;
[0017] Figure 2 This is a schematic diagram of the logical framework of a specific embodiment of the sample learning method of this application;
[0018] Figure 3 This application Figure 1 A flowchart illustrating a specific embodiment of step S12;
[0019] Figure 4 This application Figure 3A schematic diagram of the overall structure of the network module;
[0020] Figure 5 This application Figure 4 A schematic diagram of the network structure of the prediction classifier;
[0021] Figure 6 This application Figure 1 A flowchart illustrating a specific embodiment of step S13;
[0022] Figure 7 This application Figure 6 A diagram illustrating network training on an already labeled dataset;
[0023] Figure 8 This application Figure 6 A schematic diagram of the maximum difference network training in step S33 of a specific embodiment;
[0024] Figure 9 This application Figure 6 A schematic diagram of the minimum difference network training in step S33 of a specific embodiment;
[0025] Figure 10 This application Figure 1 A flowchart illustrating a specific embodiment of step S14;
[0026] Figure 11 This is a schematic diagram of the network training process in a single sample learning iteration cycle of this application;
[0027] Figure 12 This application Figure 1 A flowchart of a specific embodiment of step S15;
[0028] Figure 13 This is a flowchart illustrating the process after the sample learning method of this application meets the preset conditions;
[0029] Figure 14 This is a flowchart illustrating an embodiment of the data annotation method of this application;
[0030] Figure 15 This is a structural block diagram of a data annotation device provided in an embodiment of this application;
[0031] Figure 16 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0032] Figure 17 This application provides a schematic diagram of the structure of a computer-readable storage medium;
[0033] Figure 18 This is a schematic block diagram of the hardware architecture of the terminal in this application. Detailed Implementation
[0034] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0035] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0036] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0037] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0038] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrases "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0039] To illustrate the technical solution of this application, this application provides a sample learning method. Please refer to [link / reference needed]. Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the sample learning method of this application. The method specifically includes the following steps:
[0040] S11: Obtain the labeled and unlabeled datasets;
[0041] Data serves as the foundation for the development of Artificial Intelligence (AI), and labeled data holds a crucial position in the AI industry. In labeling scenarios, a portion of the acquired images is often labeled, resulting in labeled and unlabeled data.
[0042] By grouping labeled data into a single data pool, a labeled dataset is obtained. Similarly, by grouping unlabeled data into a separate data pool, an unlabeled dataset is obtained. When unlabeled data is labeled, the unlabeled dataset can be updated, and the updated unlabeled data can then be merged into the labeled dataset.
[0043] Of course, those skilled in the art can manually annotate a small portion of the unlabeled data to obtain labeled data, thereby providing an annotation template or prior for the annotation of the majority of unlabeled data; they can also use network models to annotate the unlabeled dataset, and then manually review the annotation. Of course, there are other annotation methods, which are not limited here.
[0044] S12: Construct a network model based on the labeled dataset;
[0045] Typically, a network model can be built based on a labeled dataset. This network model can be a semi-supervised learning model that uses semi-supervised learning (SSL) to pre-label the unlabeled dataset.
[0046] Specifically, semi-supervised learning is a learning method that combines supervised and unsupervised learning. Semi-supervised learning uses a large amount of unlabeled data, as well as labeled data, to perform pattern recognition tasks; when using semi-supervised learning, it requires as few people as possible to work, while still achieving relatively high accuracy.
[0047] S13: Input the labeled and unlabeled datasets into the network model to verify the accuracy of the unlabeled dataset and obtain the model accuracy.
[0048] In order to verify the accuracy of the model on the unlabeled dataset using the network model, which is a semi-supervised learning model, the model verifies the accuracy of the unlabeled dataset based on the reference of the labeled dataset.
[0049] Furthermore, after inputting the labeled dataset into the semi-supervised learning model, the preset model accuracy corresponding to the labeled dataset can be obtained. This preset model accuracy can be used to compare the model accuracy obtained after inputting the unlabeled dataset into the network model, thereby using the network model to verify the accuracy of the unlabeled dataset.
[0050] In addition, by inputting a portion of the labeled data from the labeled dataset into the network model, the preset model accuracy corresponding to the partial labeled dataset can be obtained, thereby realizing the collection of the validation dataset and thus achieving the purpose of accuracy verification.
[0051] S14: Determine whether the model accuracy meets the preset conditions;
[0052] Generally, network models include preset conditions to determine their accuracy. Often, if the model accuracy meets these preset conditions, it can be assumed that the network model has converged after pre-labeling the unlabeled dataset. Specifically, if the preset condition is a preset model accuracy, the model accuracy can be compared with this preset accuracy to determine whether the model accuracy meets the preset conditions.
[0053] If the model accuracy does not meet the preset conditions, proceed to step S15, which involves selecting several samples from the unlabeled dataset and manually reviewing the pre-labeled data corresponding to these samples to update the unlabeled dataset; then, the step of inputting the labeled and unlabeled datasets into the network model again to verify the accuracy of the unlabeled dataset using the network model is executed again, which means returning to step S13 and iterating multiple times until the model accuracy meets the preset conditions; if the model accuracy meets the preset conditions, proceed to step S16, which means confirming that the model accuracy meets the preset conditions.
[0054] Therefore, this application constructs a network model based on the labeled dataset, and then inputs the labeled and unlabeled datasets into the network model. This enables the unlabeled dataset to learn from samples in the network model based on the labeled dataset, thereby updating the unlabeled dataset. This reduces the need for manual annotation or review, saving manpower costs. Furthermore, by setting preset conditions for model accuracy, multiple iterations are performed until the model accuracy meets the preset conditions, further enabling the updating of the network model.
[0055] Furthermore, for a better understanding of the overall logic and framework of the sample learning method in this application, please refer to [link / reference]. Figure 2 , Figure 2 This is a schematic diagram of the logical framework of a specific embodiment of the sample learning method of this application.
[0056] First, a neural network model with a semi-supervised learning structure is constructed to obtain the semi-supervised learning model. Then, labeled and unlabeled datasets are input into the semi-supervised learning model for accuracy verification, obtaining the model accuracy. Simultaneously, pre-labeled data corresponding to all unlabeled data can also be obtained. If the target number of pre-labeled data is met, it can be determined whether the number of labels for the pre-labeled data has reached the target number. Of course, pre-labeled data can be labeled even if the target number is not met; the specific timing is not limited. The target number can be preset or less than or equal to the number of unlabeled data. If the target number is not met, the data is input into the fusion query module for sample selection. The fusion query module is based on the pre-labeled data (prediction results). Because the query strategy integrates uncertainty measurement and table edge measurement, it can actively learn to select several samples from the unlabeled data and manually review the pre-labeled data corresponding to these samples to update the unlabeled dataset. This process is repeated multiple times until the model accuracy meets the requirements.
[0057] Thus, by utilizing a semi-supervised learning model, unlabeled datasets can actively learn (AL), acquiring efficient features with a small number of labels through a periodic, learning-based approach. Specifically, after initializing the labeled dataset using machine learning, data with high information content or significant differences are selected from the unlabeled dataset for manual labeling. The manually labeled data is then added back to the dataset for retraining. This iterative process of data selection and training reduces training costs.
[0058] Furthermore, please refer to Figures 3 to 5 , Figure 3 This application Figure 1 A flowchart illustrating a specific embodiment of step S12. Figure 4 This application Figure 3 A schematic diagram of the overall structure of the network module. Figure 5 This application Figure 4 A schematic diagram of the network structure of the prediction classifier; the network model is built based on the labeled dataset, specifically including the following steps:
[0059] S21: Call the sample feature extraction module to extract features from the labeled dataset and obtain the first image features;
[0060] Before labeled data can accurately represent unlabeled data, there is usually a distributional bias between labeled and unlabeled data, especially when the number of labeled data points in the labeled dataset is small. Information-rich data should be located in the biased distribution boundary region, and the two classifiers in adversarial learning show greater classification differences at class boundaries. Based on this idea, such as... Figure 4 The overall structure of the network module in this method is divided into a sample feature extraction module and an edge sample mining module. The sample feature extraction module includes an encoder (encoder f).
[0061] Specifically, the labeled dataset is input into the sample feature extraction module, processed by the encoder, and the ResNet50 network structure is used. The output is a linear layer with 256 output units, which can obtain the first image features corresponding to the labeled dataset. This provides favorable prerequisites for building a semi-supervised learning model. Of course, after the semi-supervised learning model is built, it can also extract image features corresponding to the unlabeled dataset.
[0062] S22: Call the edge sample mining module, input the first image features into the first prediction classifier and the second prediction classifier respectively, and obtain the first prediction data of multiple categories and the second prediction data of multiple categories respectively;
[0063] The edge sample mining module includes a first predictor classifier (predictor h1) and a second predictor classifier (predictor h2). The second predictor classifier has the same network structure as the first predictor classifier, which is a Siamese neural network (SNN). An SNN is primarily a neural network architecture. Unlike a model that learns to classify its inputs, this neural network learns to distinguish between two inputs and learns the similarities between them.
[0064] Specifically, such as Figure 5 As shown, the edge sample mining module consists of two predictors, each composed of two fully connected layers, fc1 and fc2. The first layer, fc1, is a network layer with 64 linear units connected to the output of the encoder f (256 linear units), with ReLU and Dropout layers (sparse coefficient of 0.4) in between to prevent overfitting. The second layer, fc2, is connected to the output layer with N linear units representing the number of classes, outputting the predicted probability for each class. During training, the parameters of the two predictors are initialized relatively independently.
[0065] Therefore, by processing the first image features using the first predictive classifier, first predictive data for multiple categories can be obtained; by processing the first image features using the second predictive classifier, second predictive data for multiple categories can be obtained.
[0066] S23: Based on the first and second prediction data, construct the sample feature extraction module and the edge sample mining module to obtain the network model.
[0067] In this way, by comparing and learning from the first and second predicted data, we can find the similarities between them, that is, the most basic features; on the other hand, we can find the maximum difference between them, which can greatly promote the improvement of the model.
[0068] Furthermore, the network model includes a loss function, which comprises a first loss function and a second loss function. Labeled and unlabeled datasets are input into the network model to validate the model's accuracy on the unlabeled dataset, thus obtaining the model's accuracy. Please refer to [link to relevant documentation]. Figure 6 , Figure 6 This application Figure 1 Step S13 is a flowchart of a specific embodiment, which specifically includes the following steps:
[0069] S31: For the labeled dataset, call the first loss function to calculate and obtain the first prediction loss;
[0070] Assume dataset X = {X l ,X u}, where X u X is unlabeled data. l This is labeled data; network parameter Θ = {θ} e ,θ1,θ2}, where θ e Let θ1 and θ2 be the parameters of the network encoder f, and θ1 and θ2 be the parameters of the classifiers h1 and h2, respectively.
[0071] Specifically, Step 1: Train the network model using the labeled dataset, and calculate the first loss function L. cls The definition is shown in equation (1):
[0072]
[0073] in, M is the number of categories, y ic The sign function takes the value 1 if the true class of sample i is c, and 0 otherwise. Let i be the predicted probability that sample i belongs to category c; The predicted probabilities for each category.
[0074] S32: For the unlabeled dataset, the second loss function is called to calculate the second prediction loss;
[0075] For the unlabeled dataset, the second loss function is called to calculate the second prediction loss. As shown in equation (2):
[0076]
[0077] S33: Based on the first prediction loss and the second prediction loss, update the network parameters of the network module to obtain the updated network model.
[0078] Specifically, all labeled data is input into the network structure, and the cross-entropy loss function is used. This comprehensively considers the prediction bias of the two classifiers, that is, based on the first and second prediction losses, the backpropagation algorithm is used to evaluate all parameters Θ={θ} of the network model. e The network model is updated by updating θ1,θ2} to obtain the updated network model.
[0079] Furthermore, based on the first and second prediction losses, the network parameters of the network module are updated. Please refer to [link to relevant documentation]. Figure 7 , Figure 7 This application Figure 6 The diagram below illustrates the network training of the labeled dataset, specifically including two aspects:
[0080] On the one hand, Step 2: Maximize the predictive variance. See also... Figure 8 , Figure 8 This application Figure 6 Step S33, a schematic diagram of a specific embodiment of maximizing the difference network training, involves updating the first parameters of the first predictive classifier and the second parameters of the second predictive classifier based on the maximum value between the first and second predictive losses, and then stopping the updating of the encoder parameters. Figure 8 As shown, when the difference value is the largest, it indicates that there is a large deviation area in the data samples of the first and second predictive classifiers. Therefore, the network gradient is updated according to the backpropagation algorithm. In the gradient update process, the dashed part of Step 2: stop grad with X indicates that the parameters in the encoder are not updated, but only the first parameter of the first predictive classifier h1 and the second parameter of the second predictive classifier h2 are updated.
[0081] Among them, the largest prediction difference L max It can be expressed by equation (3):
[0082]
[0083] On the other hand, Step 3: Minimize the prediction discrepancy. See also... Figure 9 , Figure 9 This application Figure 6 Step S33, a schematic diagram of minimizing the difference network training in a specific embodiment, involves updating the encoder parameters based on the minimum value between the first prediction loss and the second prediction loss, and then ceasing the updating of the first and second parameters. Figure 9 As shown, when the difference value is the smallest, it indicates that there are many similarities between the data samples in the first predictive classifier and the second predictive classifier.
[0084] To ensure that the sample feature extraction module can still effectively represent all data, it is necessary to align the distribution of the labeled dataset and the unlabeled dataset. During the update process, the dashed part of Step 3: stop grad with X indicates that the first parameter of the first predictive classifier h1 and the second parameter of the second predictive classifier h2 are not updated, but only the parameters in the encoder are updated.
[0085] Among them, the minimum prediction difference L min It can be expressed by equation (4):
[0086]
[0087] Furthermore, please refer to Figure 10 and Figure 11 , Figure 10 This application Figure 1 A flowchart illustrating a specific embodiment of step S14. Figure 11 This is a schematic diagram of the network training process in a single sample learning iteration cycle of this application. The method also includes step 4: overall loss and iteration, which includes the following steps:
[0088] S41: Based on the first prediction loss, the maximum value, and the minimum value, determine whether the loss value of the loss function is less than or equal to a preset threshold;
[0089] like Figure 11 As shown, in a single AL iteration cycle, the processes of training labeled data, maximizing prediction discrepancies, and minimizing prediction discrepancies are repeated until the loss function L... total Until convergence. Specifically, the total functional loss L total It can be defined as shown in equation (5):
[0090]
[0091] Among them, L cls Indicates the first predicted loss; L max L represents the maximum prediction discrepancy, or the maximum value. min This represents the minimum prediction discrepancy, or the minimum value.
[0092] If the loss function is less than or equal to the preset threshold, proceed to step S42, which means that the loss function has converged and the model accuracy meets the preset conditions; if the loss function is greater than the preset threshold, proceed to step S43, which means that the model accuracy does not meet the preset conditions.
[0093] Furthermore, when the model accuracy does not meet the preset conditions, several samples are selected from the unlabeled dataset, and the pre-labeled data corresponding to these samples are manually reviewed to update the unlabeled dataset. Please refer to [link to relevant documentation]. Figure 12 , Figure 12 This application Figure 1 A flowchart of a specific embodiment of step S15 is shown, which specifically includes the following steps:
[0094] S51: Call the fusion query module to obtain the uncertainty measure and marginal distribution measure of the pre-labeled data;
[0095] After constructing a semi-supervised learning model through the above steps and training it with all the data (labeled and unlabeled datasets), the progress of the unlabeled dataset is verified by using the preset model accuracy of the labeled data obtained before the network model.
[0096] When the model accuracy does not meet the requirements, the network model will make predictions on all unlabeled data. Of course, the network model can also make predictions on all unlabeled data before judging the model accuracy. Here, there is no restriction on the timing of the pre-labeling of unlabeled data.
[0097] Therefore, when the requirements are not met, the fusion query module can be invoked to obtain the uncertainty measure and marginal distribution measure of the pre-labeled data. The uncertainty measure includes at least the confidence levels of the pre-labeled data for different categories.
[0098] S52: Based on uncertainty measurement, marginal distribution measurement and preset conditions, analyze and sort the pre-labeled data to select a number of samples from the unlabeled dataset, and manually review the pre-labeled data corresponding to the samples.
[0099] Based on uncertainty measures, marginal distribution measures, and preset conditions, the pre-labeled data is analyzed, sorted, and filtered to select valuable pre-labeled data, resulting in several samples selected from the unlabeled dataset. The specific process is as follows:
[0100] assumed For classifier h j The predicted probability for unlabeled samples, where M is the number of unlabeled samples; where For classifier h j For sample y i The predicted probability; the query strategy Q is shown in equation (6), and the dataset consisting of n samples is queried in each AL iteration cycle. As shown in equation (7);
[0101]
[0102]
[0103] Here, sort(Q)[:n] represents sorting Q in ascending order and taking the first n elements to form a set, and idx(.) represents the index of the set element.
[0104] Because the pre-labeled data obtained from training the network model may be inaccurate, manual review is necessary. Specifically, this can be done by manually reviewing the pre-labeled data corresponding to a number of samples.
[0105] S53: Store the reviewed pre-labeled data into the data pool containing the unlabeled dataset to update the unlabeled dataset.
[0106] After approval, the pre-labeled data is stored in the data pool containing the unlabeled dataset to update the unlabeled dataset. This allows the labeled data corresponding to several selected samples to be moved into the training sample set (e.g., ...). Figure 2 The data update process updates the labeled dataset, thus initiating the next AL iteration cycle.
[0107] Furthermore, please refer to Figure 13 , Figure 13 This is a flowchart illustrating the sample learning method of this application after meeting the preset conditions. The method also includes:
[0108] S61: When the model accuracy meets the preset conditions, pre-label the unlabeled dataset;
[0109] When a semi-supervised learning model performs accuracy verification on an unlabeled dataset, it is not necessary to pre-label all unlabeled data. In this way, the unlabeled dataset can be pre-labeled when the model accuracy meets or does not meet the preset conditions.
[0110] The difference is that when a semi-supervised learning model performs accuracy verification on an unlabeled dataset, all unlabeled data are pre-labeled once; while after determining whether the model accuracy meets or does not meet the preset conditions, all unlabeled data are pre-labeled twice. Both labeling methods are acceptable, and the choice depends on the specific needs. In fact, there is no restriction here.
[0111] S62: Determine whether the number of annotations in the pre-annotated data meets the preset number;
[0112] Regarding the number of annotations in the pre-annotated data, a preset number (also known as a target number) can be set to determine the quantity of the pre-annotated data.
[0113] If it is determined that the number of labels in the pre-labeled dataset does not meet the preset number, then proceed to step S63, which is to re-execute the step of inputting the labeled and unlabeled datasets into the network model to use the network model to verify the accuracy of the unlabeled dataset. For details, please refer to [link to relevant documentation]. Figure 1Step S13 in the process will not be repeated here. The process involves multiple iterations until the number of annotations meets the preset number. If it is determined that the number of annotations in the pre-annotated dataset meets the preset number, then the process proceeds to step S64, which means stopping the iteration.
[0114] In addition, this application also provides a data annotation method, please refer to [link / reference]. Figure 14 , Figure 14 This is a flowchart illustrating an embodiment of the data annotation method of this application. The method includes:
[0115] S71: Obtain the unlabeled dataset;
[0116] This step is as follows: Figure 1 The process of obtaining the unlabeled dataset in step S11 is similar and will not be repeated here.
[0117] S72: Call the network model described above to label the unlabeled dataset to obtain labeled data.
[0118] By utilizing updated or improved semi-supervised network learning models, unlabeled datasets can be actively learned. This is because semi-supervised learning models include a semi-supervised learning module and a fusion query module. The semi-supervised learning module is further divided into two sub-modules: sample feature extraction and edge sample mining.
[0119] By utilizing semi-supervised training methods, the sample feature extraction sub-model can effectively represent all data, including unlabeled data. The edge sample mining module trains two adversarial classifiers using all data, measuring the marginal distribution of the unlabeled dataset relative to the labeled dataset by the difference in their predictions, and measuring the information content of the sample by the minimum confidence of the predictions of a single classifier. The fusion query module innovatively integrates the marginal distribution measurement and the minimum confidence strategy for sample screening. Finally, the selected samples are manually reviewed and added to the labeled dataset. The entire process iterates through data screening and training. The training process innovatively adopts a phased stopping gradient update strategy to effectively ensure the stability of the feature extraction module and the consistency and difference in the sample predictions of the two classifiers in the edge sample mining module, preventing gradient collapsing.
[0120] Furthermore, by utilizing semi-supervised learning methods to learn from all samples, this approach uncovers marginal distribution samples of unlabeled data relative to labeled samples under unsupervised conditions. This overcomes the shortcomings of existing AI annotation methods that overly rely on manually labeled data for target model training. Simultaneously, addressing the instability of the target model and poor application in multi-class scenarios caused by the single reliance on uncertainty query strategies in current AL annotation schemes, this approach integrates marginal distribution metrics and uncertainty metrics for sample selection and manual annotation, effectively expanding the stability and application scenarios of current AL annotation schemes. This approach can be effectively applied to single-label and multi-label classification scenarios, significantly improving the training efficiency of AL algorithms while substantially reducing manual annotation costs. Moreover, this method is applicable to both single-label and multi-label classification AL, ensuring the accuracy of labeled data while effectively reducing manual workload and improving AI data annotation efficiency.
[0121] Furthermore, the sample feature extraction submodule utilizes CNNs to effectively represent all sample data. Its structure is not limited to the ResNet50 network listed in the scheme and can be replaced with any feature extraction network (ResNeXt, ResNet101 / 152, HRNet, EfficientNet, MobileNet, ShuffleNet, DenseNet, etc.). The edge sample mining submodule measures the edge distribution characteristics of unlabeled data relative to labeled data by the prediction difference between two classifiers. Its classifier construction is not limited to the two fully connected layers listed in the scheme; it can also be some more complex classifiers (such as multilayer perceptrons (MLPs)). Additionally, in the loss calculation for maximizing and minimizing the prediction difference, L... cls The loss function that calculates the difference between the predicted probability and the true class is not limited to the cross-entropy loss function (CE); other loss functions can also be used. d x is The calculation process for the difference between the predictions of two classifiers is not limited to the 1-norm; the number of correctly predicted classes can also be used as a metric.
[0122] In addition, this application also provides a data annotation device, please refer to [link / reference]. Figure 15 , Figure 15 This is a structural block diagram of a data annotation device provided in an embodiment of this application. The data annotation device 60 includes:
[0123] Module 61 is used to acquire labeled and unlabeled datasets;
[0124] Module 62 is used to build a network model based on the labeled dataset;
[0125] Processing module 63 is used to input labeled and unlabeled datasets into the network model, so as to use the network model to verify the accuracy of the unlabeled dataset and obtain the model accuracy;
[0126] The processing module 63 is also used to select several samples from the unlabeled dataset when the model accuracy does not meet the preset conditions, and to manually review the pre-labeled data corresponding to the several samples in order to update the unlabeled dataset.
[0127] The processing module 63 is also used to re-execute the step of inputting the labeled dataset and the unlabeled dataset into the network model to verify the accuracy of the unlabeled dataset using the network model, through multiple iterations until the model accuracy meets the preset conditions.
[0128] Therefore, this application constructs a network model based on the labeled dataset, and then inputs the labeled and unlabeled datasets into the network model. This enables the unlabeled dataset to learn from samples in the network model based on the labeled dataset, thereby updating the unlabeled dataset. This reduces the need for manual annotation or review, saving manpower costs. Furthermore, by setting preset conditions for model accuracy, multiple iterations are performed until the model accuracy meets the preset conditions, further enabling the updating of the network model.
[0129] In addition, this application also provides an electronic device, please refer to [link / reference needed]. Figure 16 , Figure 16 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 70 includes a processor 71 and a memory 72. The memory 72 stores a computer program 721. The processor 71 is used to execute the computer program 721 to perform the method as described above, which will not be repeated here.
[0130] In addition, this application also provides a computer-readable storage medium, please refer to [link to relevant documentation]. Figure 17 , Figure 17 This application provides a schematic diagram of the structure of a computer-readable storage medium 80, which stores a computer program 81. When the computer program 81 is executed by a processor, it can implement the method described above, which will not be repeated here.
[0131] Please see Figure 18 , Figure 18This is a schematic block diagram of the hardware architecture of the terminal of this application. The electronic device 900 can be a smart TV, industrial computer, tablet computer, mobile phone, or laptop computer, etc. This embodiment uses a mobile phone as an example. The structure of the terminal 900 may include a radio frequency (RF) circuit 910, a memory 920, an input unit 930, a display unit 940, a sensor 950, an audio circuit 960, a WiFi (wireless fidelity) module 970, a processor 980, and a power supply 990. Among them, the RF circuit 910, memory 920, input unit 930, display unit 940, sensor 950, audio circuit 960, and WiFi module 970 are respectively connected to the processor 980; the power supply 990 is used to provide power to the entire electronic device 900.
[0132] Specifically, the RF circuit 910 is used to receive and transmit signals; the memory 920 is used to store data instruction information; the input unit 930 is used to input information, and may include a touch panel 931 and other input devices 932 such as operation buttons; the display unit 940 may include a display panel; the sensor 950 includes an infrared sensor, a laser sensor, etc., used to detect user proximity signals, distance signals, etc.; the speaker 961 and the microphone 962 are connected to the processor 980 through the audio circuit 960 for receiving and transmitting sound signals; the WiFi module 970 is used to receive and transmit WiFi signals; and the processor 980 is used to process the mobile phone's data information.
[0133] The above description is only a part of the embodiments of this application and does not limit the scope of protection of this application. Any equivalent device or equivalent process transformation made based on the content of this application specification and drawings, or direct or indirect application in other related technical fields, are similarly included in the patent protection scope of this application.
Claims
1. A sample learning method, characterized in that, The method includes: Obtain the labeled and unlabeled datasets; Construct a network model based on the labeled dataset; The labeled dataset and the unlabeled dataset are input into the network model to verify the accuracy of the unlabeled dataset and obtain the model accuracy. When the model accuracy does not meet the preset conditions, several samples are selected from the unlabeled dataset, and the pre-labeled data corresponding to the several samples are manually reviewed to update the unlabeled dataset; The step of inputting the labeled dataset and the unlabeled dataset into the network model again to verify the accuracy of the unlabeled dataset using the network model is repeated multiple times until the model accuracy meets the preset condition. The step of constructing a network model based on the labeled dataset includes: The sample feature extraction module is invoked to extract features from the labeled dataset to obtain the first image features; The edge sample mining module is invoked, and the first image features are input into the first prediction classifier and the second prediction classifier respectively to obtain first prediction data and second prediction data of multiple categories respectively. The edge sample mining module includes a first prediction classifier and a second prediction classifier, and the second prediction classifier has the same network structure as the first prediction classifier. Based on the first prediction data and the second prediction data, the sample feature extraction module and the edge sample mining module are constructed to obtain the network model.
2. The method according to claim 1, characterized in that, The network model includes a loss function, which includes a first loss function and a second loss function. The step of inputting the labeled dataset and the unlabeled dataset into the network model to perform accuracy verification on the unlabeled dataset using the network model and obtain the model accuracy includes: For the labeled dataset, the first loss function is called to calculate the first prediction loss; For the unlabeled dataset, the second loss function is called to calculate the second prediction loss; Based on the first prediction loss and the second prediction loss, the network parameters of the network model are updated to obtain the updated network model.
3. The method according to claim 2, characterized in that, The sample feature extraction module includes an encoder; The step of updating the network parameters of the network model based on the first prediction loss and the second prediction loss includes: Based on the maximum value between the first prediction loss and the second prediction loss, the first parameters of the first prediction classifier and the second parameters of the second prediction classifier are updated, and the updating of the encoder parameters is stopped; or The encoder parameters are updated based on the minimum value between the first prediction loss and the second prediction loss, and the updating of the first and second parameters is stopped.
4. The method according to claim 3, characterized in that, The method further includes: Based on the first prediction loss, the maximum value, and the minimum value, determine whether the loss value of the loss function is less than or equal to a preset threshold; If the loss function is less than or equal to the model accuracy, the model accuracy is determined to be converged and the model accuracy meets the preset conditions. If the value is greater than the preset condition, then the model accuracy does not meet the preset condition.
5. The method according to claim 4, characterized in that, The network model also includes a fusion query module; When the model accuracy does not meet the preset conditions, a number of samples are selected from the unlabeled dataset, and the pre-labeled data corresponding to the number of samples are manually reviewed to update the unlabeled dataset, including: The fusion query module is invoked to obtain the uncertainty measure and marginal distribution measure of the pre-labeled data; Based on the uncertainty measure, the marginal distribution measure, and the preset conditions, the pre-labeled data is analyzed, sorted, and filtered to select several samples from the unlabeled dataset, and the pre-labeled data corresponding to the several samples is manually reviewed. The pre-labeled data after review is stored in the data pool containing the unlabeled dataset to update the unlabeled dataset.
6. The method according to any one of claims 1-4, characterized in that, The method further includes: When the model accuracy meets the preset condition, the unlabeled dataset is pre-labeled; If it is determined that the number of annotations in the pre-labeled dataset does not meet the preset number, the step of inputting the labeled dataset and the unlabeled dataset into the network model again to use the network model to verify the accuracy of the unlabeled dataset is executed again. This process is repeated multiple times until the number of annotations meets the preset number.
7. A data annotation method, characterized in that, The method includes: Obtain the unlabeled dataset; The network model described in any one of claims 1-6 is invoked to label the unlabeled dataset to obtain labeled data.
8. A data annotation device, characterized in that, The data annotation device includes: The acquisition module is used to acquire labeled and unlabeled datasets; The building module is used to construct a network model based on the labeled dataset; The processing module is used to input the labeled dataset and the unlabeled dataset into the network model, so as to use the network model to verify the accuracy of the unlabeled dataset and obtain the model accuracy; The processing module is further configured to select several samples from the unlabeled dataset when the model accuracy does not meet the preset conditions, and manually review the pre-labeled data corresponding to the several samples to update the unlabeled dataset. The processing module is further configured to re-execute the step of inputting the labeled dataset and the unlabeled dataset into the network model to verify the accuracy of the unlabeled dataset using the network model, through multiple iterations until the model accuracy meets the preset condition.
9. An electronic device, characterized in that, include: A processor and a memory, wherein the memory stores a computer program, and the processor is configured to execute the computer program to implement the method as claimed in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Data set tagging method and related apparatus
CN108062394A
Incremental learning method and system based on small number of labeled samples
CN112132179A
Image recognition method, device, equipment and storage medium
CN113569081A