Data labeling method, computer program product, storage medium and electronic device
By using the combination of the annotation model and user annotation, the data annotation method is automatically evaluated and marked, and the problems of inefficient data annotation and difficulty in ensuring accuracy in the prior art are solved, and efficient and accurate data annotation is achieved.
Patent Information
- Application Number
- CN202210551037.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-05-18
AI Technical Summary
In the prior art, data labeling mainly relies on manual completion, is inefficient, and the labeling accuracy is difficult to guarantee.
A data labeling method is proposed. By obtaining the data to be marked and its inference results, using the labeling model to make predictions, obtaining the labeling labels of some data marked by the user, estimating the labeling accuracy of the data not yet marked, and automatically marking it when the standards are met.
It significantly improves the efficiency of data labeling, ensures the accuracy of labeling results, and reduces the workload of manual labeling.
Smart Images

Figure CN115100484B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a data labeling method, a computer program product, a storage medium, and an electronic device. Background Art
[0002] Data labeling is an important task in the field of artificial intelligence. Data labeling means adding reliable label information to data. For example, in image classification tasks, data labeling can refer to marking the true category of the training image, so that the true category can be used as a supervisory signal to conduct supervised training on the image classification model, so that the model can learn the correspondence between the content of the training image and its true category, and then effectively classify other images.
[0003] However, current data labeling is basically entirely dependent on manual work, and its efficiency is very low. Summary of the invention
[0004] The purpose of the embodiments of the present application is to provide a data labeling method, a computer program product, a storage medium and an electronic device to improve the above-mentioned technical problems.
[0005] To achieve the above objectives, this application provides the following technical solutions:
[0006] In a first aspect, an embodiment of the present application provides a data labeling method, including: obtaining first data to be labeled and its inference result; wherein the inference result of the first data to be labeled is information obtained by predicting the first data to be labeled using a labeling model; obtaining a labeling label of second data to be labeled that is labeled by a user; wherein the second data to be labeled is part of the data to be labeled in the first data to be labeled; estimating whether the labeling accuracy of third data to be labeled that has not been labeled in the first data to be labeled has reached a standard based on the labeling label of the second data to be labeled; wherein the labeling accuracy of the third data to be labeled is the accuracy of the inference label determined based on the inference result of the third data to be labeled; if the labeling accuracy of the third data to be labeled has reached the standard, the inference label of the third data to be labeled is determined as the labeling label of the third data to be labeled.
[0007] The above method only needs to manually label the second data to be labeled in the first data to be labeled, and based on the labeling results, the labeling accuracy of the third data to be labeled that has not been labeled in the first data to be labeled can be evaluated, and the third data to be labeled can be automatically labeled when its accuracy meets the standard, without the need to manually label all the data to be labeled. Therefore, this method significantly improves the efficiency of data labeling, and the labeling accuracy is also effectively guaranteed. Among them, the labeling accuracy of the third data to be labeled reaches the expected standard (i.e., meets the standard), and since the second data to be labeled is manually labeled, when the user is in a normal state, its labeling accuracy can be considered to be close to or equal to 100%.
[0008] In a second aspect, an embodiment of the present application provides a computer program product, including computer program instructions, which, when read and executed by a processor, execute the method provided by the first aspect or any possible implementation of the first aspect.
[0009] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are read and executed by a processor, the method provided by the first aspect or any possible implementation of the first aspect is executed.
[0010] In a fourth aspect, an embodiment of the present application provides an electronic device, comprising: a memory and a processor, wherein the memory stores computer program instructions, and when the computer program instructions are read and run by the processor, the method provided by the first aspect or any possible implementation of the first aspect is executed. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0012] Figure 1 The process of the data annotation method provided in the embodiment of the present application is shown;
[0013] FIG. 2(A) and FIG. 2(B) show an implementation of a user annotation interface;
[0014] FIG3(A) and FIG3(B) show another implementation of the user annotation interface;
[0015] Figure 4 Shows Figure 1 A flow chart of an iterative implementation of the method in FIG.
[0016] Figure 5 The fourth principle of reducing the data to be annotated is shown;
[0017] Figure 6 The architecture and working principle of the data annotation system provided by the embodiment of the present application are shown;
[0018] Figure 7 Shows Figure 6 A way to implement model training and reasoning in the system;
[0019] Figure 8 The functional modules included in the data labeling device provided in the embodiment of the present application are shown;
[0020] Fig. 9 The structure of an electronic device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0021] In recent years, research on computer vision, deep learning, machine learning, image processing, image recognition and other technologies based on artificial intelligence has made important progress. Artificial Intelligence (AI) is an emerging science and technology that studies and develops theories, methods, technologies and application systems for simulating and extending human intelligence. Artificial intelligence is a comprehensive discipline involving many types of technologies such as chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, neural networks, etc. Computer vision, as an important branch of artificial intelligence, specifically allows machines to recognize the world. Computer vision technology usually includes face recognition, liveness detection, fingerprint recognition and anti-counterfeiting verification, biometric recognition, face detection, pedestrian detection, target detection, pedestrian recognition, image processing, image recognition, image semantic understanding, image retrieval, text recognition, video processing, video content recognition, 3D reconstruction, virtual reality, augmented reality, simultaneous positioning and map construction, computational photography, robot navigation and positioning and other technologies. With the research and progress of artificial intelligence technology, this technology has been applied in many fields, such as security control, urban management, traffic management, building management, park management, face access, face attendance, logistics management, warehouse management, robots, intelligent marketing, computational photography, mobile phone imaging, cloud services, smart homes, wearable devices, unmanned driving, automatic driving, smart medical care, face payment, face unlocking, fingerprint unlocking, identity verification, smart screen, smart TV, camera, mobile Internet, live broadcast, beauty, beauty, medical beauty, intelligent temperature measurement, etc. The data annotation method in the embodiment of this application also belongs to the category of artificial intelligence technology.
[0022] Before introducing the technical solutions in the embodiments of the present application, the meaning of data annotation and the significance of data annotation are briefly explained.
[0023] Data labeling is the process of labeling data through some means (for example, manual confirmation). The label can be understood as the value of a certain attribute (or several attributes) of the data. If the labeling effect is ideal, then the label should be the true value of the attribute.
[0024] In the field of artificial intelligence, data annotation is often performed on the training samples of a certain model (for example, a neural network model). The labeled training samples can be used for supervised training of the model, and the trained model can be used to perform specific tasks. For the convenience of description in the following text, it may be referred to as a task model. These tasks can be summarized as: inputting the task data into the task model, and after the task model is calculated, predicting the value of a certain attribute (or several attributes) of the task data. The predicted value may have actual physical meaning in specific problems.
[0025] For example, a possible training process is: input the training sample into the task model, obtain the prediction result of the model output for a certain attribute of the training sample, substitute the prediction result and the label of the training sample into the preset loss function to calculate the prediction loss, which represents the difference between the prediction result and the label of the training sample, and iteratively optimize the parameters of the task model until the loss converges.
[0026] If the label of the training sample represents the true value of a certain attribute of the training sample, then the training process can also be regarded as a process of adjusting the parameters of the task model so that the predicted result of the attribute of the training sample outputted by it is closer and closer to its true value. Therefore, after the task model is trained, it can also predict a more reasonable result for the attribute of the task data when processing task data of the same type as the training sample, that is, it can better complete the specific task.
[0027] For example, the task model can be an image classification model, which is used to predict the category of a specific object in an image, such as predicting which number between 0 and 9 the handwritten number in the image is. In this case, the training sample can be a pre-collected handwritten number image, and the label of the training sample is the real number in the pre-annotated image.
[0028] For another example, the task model may be a target detection model, which is used to detect the position of a specific object in an image, such as the position of a vehicle in an image. In this case, the training sample may be a pre-collected road image, and the label of the training sample is the actual position of the vehicle in the pre-annotated image.
[0029] For another example, the task model may be a part-of-speech analysis model, which is used to predict the part of speech of each word in a sentence, such as noun, verb, preposition, etc. In this case, the training samples may be pre-collected sentences, and the labels of the training samples are the true parts of speech of the words in the pre-annotated sentences.
[0030] In short, the scheme of this application does not limit the specific form of the data to be annotated, such as images, text, voice, etc.; nor does it limit the specific form of the label, which is related to the specific task to be performed by the task model, such as the real category and real position of the object in the image. In the following text, for simplicity, the data annotation method proposed in this application is mainly introduced by taking the case where the data to be annotated is an image as an example, but this should not be regarded as limiting the scope of protection of this application.
[0031] Furthermore, in the field of artificial intelligence, human cognitive results are usually taken as correct results, and the task model is only a simulation of the cognitive process of humans on a specific task. Therefore, current data annotation is usually done manually, and the manually annotated labels are taken as the true values of the attributes in the data. However, the training of the task model is likely to require a large number of labeled training samples, resulting in a huge workload for the annotators and very low annotation efficiency.
[0032] The data labeling method proposed in this application aims to combine manual labeling and machine automatic labeling to significantly improve the efficiency of data labeling and ensure that the accuracy of the labeling results can meet the task requirements.
[0033] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.
[0034] The terms "comprises," "comprising," or any other variation thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0035] The terms "first", "second", etc. are only used to distinguish one entity or operation from another entity or operation, and should not be understood as indicating or implying relative importance, nor should they be understood as requiring or implying any such actual relationship or order between these entities or operations.
[0036] Figure 1 The process of the data annotation method in the embodiment of the present application is shown. The method can be, but is not limited to, Fig. 9 The electronic device shown in FIG. 1 is executed. The specific structure of the electronic device can be referred to in the following text about Fig. 9 Reference Figure 1 , data annotation methods include:
[0037] Step S110: Obtain the first data to be labeled and its inference result.
[0038] The first data to be labeled is a batch of data to be labeled. For example, the first data to be labeled can be a training sample that has not yet been labeled. After being labeled, it can be used as a formal training sample for the task model. The inference result of the first data to be labeled is the information obtained by predicting the first data to be labeled using the labeling model, or after the first data to be labeled is input into the labeling model, the model output obtained is the inference result of the first data to be labeled. For example, if the labeling model is a classification model, the inference result of the first data to be labeled can be the probability predicted by the model that the first data to be labeled belongs to each category. If the labeling model is a regression model, the inference result of the first data to be labeled can be the regression value predicted by the model.
[0039] The annotation model is the model used in the data annotation stage. The annotation model can have the same input and output as the task model mentioned above. For example, if the task model is an image classification model, its input is an image, and its output is the probability that the image belongs to each category (for example, for handwritten digit recognition, it is the probability that the image belongs to the numbers 0 to 9, and there is only one handwritten digit in the image by default), then the annotation model can also be an image classification model, its input is an image, and its output is the probability that the image belongs to each category. For another example, if the task model is a target detection model, its input is an image, and its output is the position of each target in the image (for example, for vehicle detection, it is the position of each vehicle in the image), then the annotation model can also be a target detection model, its input is an image, and its output is the position of each target in the image, and so on.
[0040] However, the annotation model and the task model may be, but not necessarily, the same model. That is, they may have the same structure but different parameters, or simply different structures. There are many factors that lead to the two being designed as different models:
[0041] For example, in a certain scenario, we hope to give the data to be labeled as accurately as possible, that is, we do not need to consider resource consumption too much when labeling the data. At this time, we can use a more complex model as the labeling model; but for the task model, it is likely to be deployed in an actual production environment (for example, on a mobile phone or camera), and the resources it can use are strictly limited, so we can only use a simpler model.
[0042] For another example, as can be seen in the following text, in some implementations of the data annotation method, the annotation model will only be trained using manually annotated data, and will not be trained using machine-annotated data. Therefore, directly using it to perform tasks will not achieve very good results, because the training data it has "seen" is very limited. The task model may use both manually annotated data and machine-annotated data during training, so it will perform tasks better because it has "seen" more training data. How to train the annotation model will not be elaborated here.
[0043] For simplicity, in the examples given below, the annotation model and the task model can be considered to be the same type of models (for example, both are image classification models) and have the same form of input and output, but this should not be regarded as a limitation on the scope of protection of this application.
[0044] The "obtaining" in step S110 generally refers to obtaining through various means: for example, the first data to be labeled and the inference results thereof may have been obtained and stored somewhere before executing step S110, and the first data to be labeled and the inference results thereof may be directly read when step S110 is executed; for another example, the first data to be labeled and the inference results thereof may not have been obtained before executing step S110, and the first data to be labeled is obtained from somewhere (for example, the inference result pool mentioned later) when step S110 is executed, and the inference results of the first data to be labeled are calculated using the labeling model, and so on.
[0045] Step S120: Obtain the annotation label of the second data to be annotated that is annotated by the user.
[0046] The second data to be labeled is part of the data to be labeled in the first data to be labeled. The second data to be labeled can be sampled from the first data to be labeled, and the sampling method is not limited, for example, it can be random sampling, or it can be sampling according to a fixed pattern, etc. There may be no repeated data in the second data to be labeled to avoid repeated labeling.
[0047] After obtaining the second data to be annotated, it can be provided to the user for manual annotation, where the user refers to the annotator. There are many ways to provide the second data to be annotated to the user: for example, if the device used by the user for data annotation (hereinafter referred to as the user terminal) and the device that performs step S120 (hereinafter referred to as the current device) are not the same device, the second data to be annotated can be sent from the current device to the user terminal and displayed on the user annotation interface of the user terminal (for example, a web page or client interface) for the user to annotate; for another example, if the user terminal and the current device are the same device, the second data to be annotated can be directly displayed on the user annotation interface of the current device for the user to annotate, and so on.
[0048] It should be understood that although the user must somehow know the content of the second data to be annotated in order to annotate it, the way of obtaining the information is not limited to displaying it in the user annotation interface. For example, the content of the second data to be annotated can also be played by voice (if it can be played by voice), or the second data to be annotated can be printed out for the user to view, etc.
[0049] The user labels the second data to be labeled, which means manually labeling it. This label is called the labeling label. If the manual labeling is assumed to be accurate, the labeling label is also the final label of the second data to be labeled. For example, if the first data to be labeled is used to train a classification model (for example, the task model and the labeling model are both classification models), the user can use the category of the first data to be labeled that he or she thinks is as its labeling label; for another example, if the first data to be labeled is used to train a regression model (for example, the task model and the labeling model are both regression models), the user can use the regression value of the first data to be labeled that he or she thinks is as its labeling label. When introducing Figures 2(A), 2(B), 3(A) and 3(B), the user's labeling process will be further explained, which will not be expanded here.
[0050] After the user completes the annotation, he can trigger the confirmation operation representing the completion of the annotation (for example, click "Submit" on the user annotation interface), and return the annotation result, that is, the annotation label of the second data to be annotated, to the current device (if the user has already annotated on the current device, it can be considered that it is returned to the main program that executes the data annotation process), so that the current device obtains the annotation label of the second data to be annotated. After executing step S120, the second data to be annotated is annotated.
[0051] Step S130: Estimate whether the labeling accuracy of the third data to be labeled has reached the standard according to the labeling label of the second data to be labeled.
[0052] The third data to be labeled is the data that has not been labeled in the first data to be labeled. Since the second data to be labeled is only part of the data to be labeled in the first data to be labeled, after the second data to be labeled is labeled, there must be data that has not been labeled in the first data to be labeled, which may generate the third data to be labeled. However, it should be noted that all the data that has not been labeled in the first data to be labeled may, but not necessarily, belong to the third data to be labeled. There are different implementation methods here. For example, as will be known later, in some implementation methods, some unlabeled data that are evaluated to have a low labeling accuracy will be excluded from the third data to be labeled.
[0053] The labeling accuracy of the third data to be labeled is the accuracy of the inference label determined according to the inference result of the third data to be labeled. Since the third data to be labeled belongs to the first data to be labeled, and the inference result of the first data to be labeled has been obtained in step S110, the inference result of the third data to be labeled is known when executing step S130.
[0054] The inference label of the third data to be labeled is a label obtained based on the inference result of the third data to be labeled and combined with a specific rule. The inference label is no different from the label label introduced above in form, except that the inference label is a label automatically given by the machine, not a manually labeled label, so its accuracy is usually not 100% (the manually labeled label is assumed to be correct). The labeling accuracy can be estimated based on the label label of the second data to be labeled. The principle is that the second data to be labeled also has an inference label, and the rule for obtaining the inference label is the same as that of the third data to be labeled. Therefore, the difference between the inference label of the second data to be labeled and the label (correct labeling result) can reflect the accuracy of the inference label calculated according to this rule. There are many rules for obtaining inference labels, and examples will be given later, which will not be expanded here.
[0055] It should be pointed out that the concept of inference label introduced above is only to explain the meaning of the labeling accuracy of the third data to be labeled. It does not mean that the inference label of the third data to be labeled must be calculated when estimating the labeling accuracy, nor does it mean that the third data to be labeled must be labeled with inference labels when estimating the labeling accuracy. The definition of the labeling accuracy can be understood as follows: if the third data to be labeled is labeled according to its inference label, the accuracy of the labeling result meets the estimated accuracy. In short, there is no contradiction between the third data to be labeled having an inference label and being in a state of not being labeled.
[0056] The accuracy of the third data to be labeled meets the standard, which means that the estimated accuracy of its labeling is greater than the target accuracy. The target accuracy can be an accuracy set according to the training requirements of the task model before starting data labeling, for example, it can be 99%, 98%, etc. If it is estimated that the accuracy of the third data to be labeled meets the standard, step S140 is executed. If it is estimated that the accuracy of the third data to be labeled does not meet the standard, different processing methods may be adopted: for example, more second data to be labeled are sampled, and steps S120 to S130 are executed to re-estimate the labeling accuracy of the third data to be labeled. Of course, the third data to be labeled may change at this time; for example, the labeling model is retrained, and steps S110 to S130 are executed to re-estimate the labeling accuracy of the third data to be labeled; for example, the third data to be labeled is directly manually labeled, and so on.
[0057] Step S140: Determine the inferred label of the third data to be labeled as the labeled label of the third data to be labeled.
[0058] Since the labeling accuracy of the third data to be labeled can meet the requirements, the third data to be labeled can be labeled directly according to the inference label, and the labeling label of the third data to be labeled is also its final label. If the inference label of the third data to be labeled has not been calculated before, the inference label of the third data to be labeled needs to be calculated first in step S140. Since the third data to be labeled is automatically labeled, its labeling efficiency is much higher than manual labeling, and its labeling accuracy is also guaranteed, for example, it is sufficient to meet the training requirements of the task model.
[0059] After executing step S140, the second data to be labeled and the third data to be labeled in the first data to be labeled have been labeled, but it is not ruled out that there is still data in the first data to be labeled that has not been labeled. For these data, different processing methods can be adopted: for example, direct manual labeling; for example, retraining the labeling model and then labeling (at this time the inference result may change, thereby changing the labeling result), and so on.
[0060] Brief summary Figure 1The method in this paper only needs to manually label the second data to be labeled in the first data to be labeled, and based on the labeling result, the labeling accuracy of the third data to be labeled that has not been labeled in the first data to be labeled can be evaluated, and the third data to be labeled can be automatically labeled when its accuracy meets the standard, without the need to manually label all the data to be labeled. Therefore, this method significantly improves the efficiency of data labeling, and the labeling accuracy is also effectively guaranteed, for example, it can meet the training requirements of the task model. Among them, the labeling accuracy of the third data to be labeled reaches the expected standard, and since the second data to be labeled is manually labeled, its labeling accuracy can be considered to be close to or equal to 100% when the user is in a normal state.
[0061] Note that the inventors have found that the data to be labeled are often huge in number, and the first data to be labeled is likely to be just a batch of data. Even if not every batch of the first data to be labeled can trigger the execution of step S140, there are still many batches of the first data to be labeled that can trigger the execution of step S140, that is, automatic labeling is almost always executed, and the above beneficial effects are not only achieved in a few special cases.
[0062] In addition, it should be pointed out that although data annotation and task model training are always linked together in the above explanation, in fact the two are not bound together, that is, the labeled data (the second data to be labeled and the third data to be labeled) are not necessarily used to train the task model. For example, they may only be used for display and storage purposes, that is, their use is not limited.
[0063] When describing step S120, it is mentioned that in order to obtain the annotation label of the second data to be annotated, the second data to be annotated can be first displayed on the user annotation interface and annotated by the user. Since the second data to be annotated has not been annotated before, in one implementation, the second data to be annotated with unknown labels can be displayed on the user annotation interface, and the annotation label is completely given by the user. This implementation is relatively simple, but the workload of the user is relatively large.
[0064] Therefore, in another implementation, the predicted label of the second data to be labeled can be determined based on the inference result of the second data to be labeled, and not only the second data to be labeled but also its predicted label should be displayed on the user labeling interface.
[0065] Among them, the predicted label is a label obtained based on the inference result of the second data to be labeled and combined with specific rules. The predicted label is no different in form from the labeled label introduced above, except that the predicted label is not a manually labeled label. There are many specific rules for obtaining the predicted label, and examples will be given later. I will not expand on them here, but it should be noted that the specific rules for calculating the predicted label and the specific rules for calculating the inference label may be the same or different, so the predicted label and the inference label of the same second data to be labeled may be the same or different.
[0066] The predicted label can represent the annotation result predicted according to the inference result of the annotation model. If the inference result of the annotation model is accurate enough, the predicted label may be very close to the annotation label. Therefore, the predicted label can be used as important reference information for users when performing manual annotation, so that users can directly confirm the predicted label as the annotation label without modifying the predicted label displayed on the user annotation interface. Alternatively, most of the predicted labels can be confirmed as the annotation labels by only slightly modifying the predicted label displayed on the user annotation interface (modified to the annotation label that the user considers appropriate), thereby significantly improving the user's annotation efficiency.
[0067] Optionally, if the annotation model is a classification model, the predicted label of the second data to be labeled is the predicted category of the second data to be labeled. At this time, the second data to be labeled can be classified and displayed on the user annotation interface according to the predicted category of the second data to be labeled.
[0068] 2(A), the labeling model is an emotion classification model, the second data to be labeled is a face image, the inference result of the second data to be labeled can be the probability that the face in the image belongs to each emotion category (for example, happy, sad, surprised, etc.), and the predicted category of the second data to be labeled can be the emotion category to which the image belongs, such as happy, sad, surprised, etc.
[0069] In Figure 2(A), face images predicted to be in the happy category are displayed in one row, face images predicted to be in the sad category are displayed in another row, and so on. After the face images are classified and displayed, it is easy for the user to confirm the emotion category to which each face image belongs, and it is also easier to pick out images with incorrect prediction categories and change their categories, because a small amount of heterogeneous data will be more obvious in a large amount of similar data with common characteristics. For example, for the second face image in the happy category, the curvature direction of its mouth corner is obviously different from that of other face images, and the user can quickly confirm that its predicted category is wrong.
[0070] Furthermore, the second data to be labeled and its predicted category can be displayed in the first area of the user labeling interface, and the user can perform a data selection operation in the first area. If a piece of second data to be labeled is selected by the user, it indicates that the predicted category of the second data to be labeled is not recognized by the user.
[0071] For example, referring to Figure 2 (B), the upper half of the user annotation interface is the first area. If the user uses a PC to perform annotation, the data selection operation can be a single click of the left mouse button. Assuming that the user believes that the predicted category of the second face image in the happy category is wrong, the user can move the cursor on the screen to this face image and then click the left mouse button. At this time, the words "unknown category" will appear under the face image, indicating that the category of the image is to be determined.
[0072] In response to the data selection operation triggered in the first area, the user annotation interface may display the selected second data to be annotated in the second area of the interface, and display the candidate categories of the selected second data to be annotated in the second area for the user to select when manually annotating.
[0073] For example, referring to FIG2(B), the lower half of the user annotation interface is the second area, where the second face image in the happy category is displayed, and all possible emotion categories are displayed on the right for the user to select as the annotation label. In particular, the unknown category can also be used as a candidate category. If the user selects the unknown category, it means that the annotation label of the face image is unknown, or it can be regarded as not having been annotated.
[0074] It should be understood that the layout of the first area and the second area in the user annotation interface is not limited to the layout shown in FIG. 2(A) and FIG. 2(B), and the first area and the second area may even overlap.
[0075] The beneficial effect of labeling in the above manner is that in the first area, the user only needs to pay attention to whether the predicted category of a certain second data to be labeled is correct, that is, only need to make a judgment (judge "yes" or "no"), without making a selection (selecting the correct label), and in the second area, the user only needs to make a selection without making a judgment. Thus, the user can first select all the second data to be labeled with incorrect predicted categories in the first area, and then uniformly modify the categories in the second area (change to the categories in the labeling label), that is, in each area, the user only needs to operate according to a single labeling logic, without switching back and forth between two thinking logics. Experiments show that this method can significantly improve the user's labeling speed.
[0076] It should be understood that it is not necessary to divide the annotation interface into the first area and the second area. For example, only the first area can be set. After the user selects each second data to be annotated in the first area, the category is changed directly at its display position. However, at this time, the user needs to switch back and forth between the two thinking logics of judgment and selection, and the annotation efficiency may be reduced.
[0077] In addition, even if the annotation model is a classification model, the second data to be annotated does not have to be displayed according to the predicted category on the user annotation interface. For example, all categories of second data to be annotated can be displayed mixedly, and the predicted category of each second data to be annotated can also be displayed at the display position.
[0078] Furthermore, the inventors have found that when users are labeling, they are likely to judge whether their predicted labels are correct or what their labeling labels are based on a part of the second data to be labeled. For example, in Figure 2 (B), users may mainly judge the emotional category based on the curvature of the corners of the mouth of the face image (upward curvature means happiness, downward curvature means sadness), and do not pay much attention to the rest of the face image. Therefore, if the user labeling interface can focus the display content on the key part of the user's attention (i.e., the position where the user's attention is focused) when displaying the second data to be labeled, it will be more conducive to users to quickly and accurately give labeling labels.
[0079] When the second data to be annotated is an image, possible methods for displaying the data based on user attention are as follows:
[0080] The second data to be annotated and its predicted label are displayed in the first area of the user annotation interface, and the reference image is displayed in the third area of the user annotation interface. The reference image is an image selected from all the second data to be annotated currently to be annotated, which can be selected by the user or automatically selected by the program (for example, automatically selecting the first image in the list).
[0081] The user can perform an image transformation operation on the reference image in the third area, the purpose of which can be to display the user's attention-focused area in the reference image that is closely related to the annotation result in a more prominent manner. The image transformation operation includes at least one of image translation, image rotation, image scaling, and frame selection of a local area of the image.
[0082] In response to the image transformation operation triggered in the third area, the user annotation interface (or background) can record the image transformation operation. Since each operation corresponds to specific operation parameters, recording the operation is actually recording these operation parameters. For example, the translation amount can be recorded for image translation, the rotation angle can be recorded for image rotation, the zoom multiple can be recorded for image scaling, the position of the local area can be recorded for the selection of the local area of the image, and so on.
[0083] Finally, the recorded image transformation operation can be applied to each image in the second data to be annotated, and the second data to be annotated after the image transformation operation is refreshed and displayed in the first area. Thus, each image in the second data to be annotated is also displayed in a more prominent manner in the same manner as the reference image, which is convenient for the user to observe and annotate.
[0084] For example, referring to FIG3(A), the left part of the user annotation interface is the third area, the upper half of the right part is the first area, and the lower half of the right part is the second area. There is a fixed-size window (gray) in the third area, and the first face image in the happy category is selected as the reference image and displayed in the window.
[0085] Afterwards, the user can perform at least one of the following operations in the fixed window: translation, rotation, scaling, or selecting a local area of the image. The operation parameters will be recorded, and the part that exceeds the fixed window size during the transformation process will not be displayed. Referring to Figure 3(B), the user performed a translation operation to move the face in the image to the center of the window. At this time, the offset of the image center relative to the window center can be recorded as the translation parameter; the user also performed a zoom operation to make the face in the image fill the entire window. At this time, the image magnification can be recorded as the zoom parameter; the user also selected the mouth area of the image that needs to be focused on when annotating, as shown in the thick dashed box. At this time, the position of the mouth area can be recorded as the parameter for selecting the local area.
[0086] Finally, a transformation matrix can be generated based on the recorded translation parameters and scaling parameters. By applying the transformation matrix to each face image in the second data to be annotated, the face can be moved to the center of the screen and enlarged. Then, based on the recorded position of the mouth area, the area at the same position is framed from each face, and the framed part is further enlarged until it fills the original display area of the second image to be annotated (it can also be a window of a fixed size). The final refresh display effect is shown in Figure 3 (B). It can be seen that in the first area, the face image has been replaced by a partial enlarged image of its mouth. Based on these enlarged images, the user can see the curvature direction of the mouth corner more clearly, so that the wrong predicted label can be quickly selected, or the correct annotation label can be given.
[0087] Note that, first, the image transformation operation, when applied to each image in the second data to be labeled, is not necessarily the same as the effect when applied to the reference image. For example, in Figure 3(B), for the operation of selecting a local area of the image, when applied to the reference image, only the position of the mouth area is given, while when applied to each image in the second data to be labeled, not only the position of the mouth area is determined, but also the mouth area is enlarged and displayed.
[0088] Second, the layout of the first area and the third area in the user annotation interface is not limited to the layout shown in FIG. 3(A) and FIG. 3(B) , and the first area and the third area may even overlap.
[0089] Third, although the first area, the second area, and the third area are shown in FIG. 3(A) and FIG. 3(B), in the above scheme of displaying data based on user attention, whether to set the second area is optional.
[0090] Fourth, the above scheme for displaying data based on user attention has nothing to do with whether the annotation model is a classification model. For example, if the annotation model is a target detection model, this scheme can also be applied to facilitate users to more accurately mark the location of the target.
[0091] Based on the above embodiment, the following continues to introduce the estimation method of whether the labeling accuracy of the third data to be labeled meets the standard in step S130. In one implementation, step S130 can be performed according to the following process:
[0092] According to the label of the second data to be labeled, combined with specific rules, the fourth data to be labeled is determined from the first data to be labeled; if the fourth data to be labeled contains third data to be labeled that have not yet been labeled, then according to the number of correctly labeled second data to be labeled in the fourth data to be labeled, it is estimated whether the labeling accuracy of the third data to be labeled meets the standard.
[0093] Among them, the fourth data to be labeled may include the second data to be labeled that is correctly labeled or the third data to be labeled that has not been labeled, or may include both types of data to be labeled. Which data is specifically included is related to the design of the above-mentioned specific rules. The second data to be labeled is correctly labeled if: the label of the second data to be labeled is the same as the inference label of the second data to be labeled determined based on the inference result of the second data to be labeled. The meaning of the inference label and the label label has been explained in the previous text and will not be repeated.
[0094] Let's first consider the case where the fourth data to be labeled contains the second data to be labeled that are correctly labeled and the third data to be labeled that have not been labeled. At this time, the fourth data to be labeled is equivalent to defining a range from the first data to be labeled, and the inference labels of some of the data to be labeled within this range have been manually confirmed to be correct. Since the rules for determining the inference labels of each data to be labeled are unified, the accuracy of the inference labels of the remaining data to be labeled can be estimated based on the number of correctly labeled data within the range.
[0095] If the fourth data to be labeled only contains the third data to be labeled that have not been labeled, it is equivalent to the case where the number of correctly labeled second data to be labeled in the fourth data to be labeled is 0, and it can be handled uniformly with the case where the number of correctly labeled second data to be labeled is not 0. If the fourth data to be labeled only contains the correctly labeled second data to be labeled, it means that this batch of fourth data to be labeled has been labeled and there is no data that needs to be automatically labeled. It should be understood that the rules for determining the fourth data to be labeled should not make the fourth data to be labeled always contain only the third data to be labeled that have not been labeled, because the labeling accuracy will inevitably be difficult to meet the standard, nor should it always contain only the correctly labeled second data to be labeled, because this is equivalent to never performing automatic labeling.
[0096] Further, when the above implementation is adopted in step S130, Figure 1 The method in can also be implemented as Figure 4 The iterative process in:
[0097] Step S210: Obtain the first data to be labeled and its inference result.
[0098] Step S210 is similar to step S110 and will not be described again. Figure 4 From the overall point of view, it is an iterative process, but step S210 does not participate in the iteration. Steps S220 to S260 are an iterative implementation of steps S120 to S140, or each round of iteration logically includes steps S220 to S260 (of course, not all steps S220 to S260 will be executed in each round of iteration).
[0099] Step S220: sampling the second data to be labeled from the fourth data to be labeled and provided to the user for labeling in this round of iteration.
[0100] The fourth data to be labeled in step S220 can be considered as the fourth data to be labeled at the end of the previous round of iteration, or the fourth data to be labeled at the beginning of this round of iteration. In each round of iteration, the fourth data to be labeled may be updated, but not necessarily (see step S230). For the first round of iteration, since there is no previous round, the fourth data to be labeled at the beginning of the first round of iteration is the first data to be labeled.
[0101] There is no limit to the method of sampling the second data to be labeled from the fourth data to be labeled. The second data to be labeled must be resampled in each round of iteration. Optionally, for those second data to be labeled that have been sampled in previous rounds, since they have been labeled by users in previous iterations, they do not need to be sampled repeatedly. That is, the second data to be labeled each time can be sampled from the third data to be labeled that has not been labeled in the fourth data to be labeled after the previous round of iteration, thereby avoiding repeated labeling of data.
[0102] Step S230: obtaining the annotation label of the second data to be annotated that is annotated by the user, and screening the fourth data to be annotated according to the annotation label of the second data to be annotated.
[0103] The sampled second data to be labeled can be sent to the user for labeling and obtain its labeling label. The specific implementation method can be referred to the previous text and will not be repeated here.
[0104] If the fourth data to be labeled only include the second data to be labeled that are correctly labeled and / or the third data to be labeled that have not been labeled, then after the second data to be labeled in this round are labeled, the current fourth data to be labeled may no longer meet the requirement, so it is necessary to screen the fourth data to be labeled according to specific rules to obtain new fourth data to be labeled. The new fourth data to be labeled only include the second data to be labeled that are correctly labeled and / or the third data to be labeled that have not been labeled, so that the labeling accuracy of the third data to be labeled can continue to be effectively evaluated.
[0105] As for the filtered out data to be labeled, including the second data to be labeled with errors, and may also include some data to be labeled that have not yet been labeled, there is no limit to the processing method. For example, the second data to be labeled with errors only means that there is an error in its inference label, but there is no error in its annotation label (the manual annotation is accurate by default), so it can be considered that it has been labeled. For the data to be labeled that have not yet been labeled, they can be re-inferred and re-annotated after the annotation model is updated, and so on.
[0106] Note that in the steps after step S230, all references to the fourth data to be labeled refer to the new fourth data to be labeled obtained after screening. For simplicity, it is still referred to as the fourth data to be labeled.
[0107] Step S240: determining whether the fourth data to be labeled includes the third data to be labeled that has not been labeled.
[0108] If the fourth data to be labeled contains the third data to be labeled that has not been labeled, then continue to execute step S250. If the fourth data to be labeled does not contain the third data to be labeled that has not been labeled, that is, the fourth data to be labeled only contains the second data to be labeled that is correctly labeled, then automatic labeling cannot be performed at this time, and thus the labeling of the current batch of the fourth data to be labeled can be terminated. Since the parts of the first data to be labeled except the fourth data to be labeled are all generated by the screening operation in step S230, if this part of the data is not labeled temporarily, the labeling of the current batch of the first data to be labeled can also be terminated at this time.
[0109] Step S250: Estimate whether the labeling accuracy of the third data to be labeled meets the standard according to the number of the second data to be labeled that are correctly labeled in the fourth data to be labeled.
[0110] If the labeling accuracy of the third data to be labeled has reached the standard, then continue to execute step S260. If the labeling accuracy of the third data to be labeled has not reached the standard, then the next round of iteration can be started, that is, jump to step S220 to continue. Note that the second data to be labeled correctly labeled in step S250 is not limited to the second data to be labeled labeled in this round of iteration (that is, not limited to the second data to be labeled sampled in step S220), and the second data to be labeled labeled in previous iterations should all be taken into consideration.
[0111] Step S260: Determine the inferred label of the third data to be labeled as the labeled label of the third data to be labeled.
[0112] Step S260 is similar to step S140 and will not be described again. After step S260 is executed, the labeling of the fourth batch of data to be labeled is completed, and if the first data to be labeled other than the fourth data to be labeled is not labeled temporarily, the labeling of the current batch of first data to be labeled can also be completed at this time.
[0113] Note that although step S260 is nominally an iterative step (for example, this step can be included in the code of the iterative part), in fact, during the entire iterative process, this step is only executed once at most when the conditions are met, and the iteration will be exited after execution.
[0114] The key to the above iterative process is to continuously screen the fourth data to be labeled, so that the third data to be labeled in the fourth data to be labeled has a higher and higher labeling accuracy estimate until it reaches the target (of course, it may not reach the target in the end). Figure 4 The process in Figure 1 One way to implement the method in Figure 4 For any steps not mentioned in the Figure 1 Description of the relevant steps in .
[0115] The following uses the case where the labeling model is a classification model as an example to illustrate two implementation methods of iterative steps S220 to S260. At this time, the inference result of the first data to be labeled in step S210 includes the probability that the first data to be labeled belongs to each category.
[0116] Method 1
[0117] Before executing the first round of iteration, a label candidate pool corresponding to each category is created based on the inference results of the first data to be labeled. A label candidate pool can be understood as a data set or a storage space. For example, if the classification results of the labeling model have a total of nc categories, then nc label candidate pools need to be created. Label candidate pool i corresponds to category i, where i is the category number, for example, i can be any integer between 0 and nc-1.
[0118] Each annotation candidate pool includes the probability that the first data to be labeled and its inference results belong to the category corresponding to the annotation candidate pool. For example, the probability that the first data to be labeled and its inference results belong to category i is contained in the annotation candidate pool i.
[0119] Note that the so-called inclusion of the first data to be labeled in the labeling candidate pool does not mean that a copy of the first data to be labeled must be made for each labeling candidate pool. It is also possible to save a copy of the identifier of the first data to be labeled (for example, the serial number of each first data to be labeled) in the labeling candidate pool. Since the corresponding first data to be labeled can be found according to the identifier of the first data to be labeled, it is equivalent to saving a copy of the first data to be labeled in the labeling candidate pool. For other data pools mentioned later, if there are different data pools storing the same data, it can be understood similarly.
[0120] After creating the annotation candidate pool, each iteration can include the following steps:
[0121] Step A1: Sampling the second data to be labeled provided to the user for labeling in this iteration from the fourth data to be labeled in all labeling candidate pools.
[0122] In step A1, the fourth data to be labeled in each labeling candidate pool can be considered as the fourth data to be labeled in the labeling candidate pool at the end of the previous round of iteration, or the fourth data to be labeled in the labeling candidate pool at the beginning of this round of iteration. In each round of iteration, the fourth data to be labeled in each labeling candidate pool may be updated, but not necessarily (see step C1). For the first round of iteration, since there is no previous round, at the beginning of the first round of iteration, the fourth data to be labeled in each labeling candidate pool is the first data to be labeled in the labeling candidate pool.
[0123] There is no limit to the method of sampling the second data to be labeled from the fourth data to be labeled in all the candidate pools for labeling, and the second data to be labeled must be resampled in each round of iteration. For example, the fourth data to be labeled in all the candidate pools for labeling can be taken as a union to remove duplicate data to be labeled, and then the second data to be labeled to be labeled in this round can be sampled from the union. Furthermore, the third data to be labeled that has not been labeled in the fourth data to be labeled in all the candidate pools for labeling can be taken as a union, and then the second data to be labeled to be labeled in this round can be sampled from the union, thereby avoiding repeated sampling of the already labeled data.
[0124] Step B1: Obtain the annotation category of the second data to be annotated that is annotated by the user.
[0125] Since the problem at this time is classification, the labeling category of the second data to be labeled is the labeling label of the second data to be labeled mentioned in step S230.
[0126] For each annotation candidate pool, perform the following steps, taking any annotation candidate pool i as an example:
[0127] Step C1: For each second data to be labeled obtained in step A1, taking any second data to be labeled j as an example, if the labeling category of the second data to be labeled j is different from the category corresponding to the labeling candidate pool i (i.e., category i), and the second data to be labeled j is included in the fourth data to be labeled in the labeling candidate pool i, then the fourth data to be labeled in the labeling candidate pool i is reduced to only the data to be labeled that meets the following conditions: the probability of belonging to category i in the inference result is greater than the probability of belonging to category i in the inference result of the second data to be labeled j.
[0128] The reduction in step C1 is an implementation of the screening in step 230. Category i is the inference label of all the first data to be labeled in the label candidate pool i. Therefore, if the label category of the second data to be labeled j is different from category i, it means that the inference label of the second data to be labeled j is wrong. According to the definition of the fourth data to be labeled in the previous text, it can only contain the correctly labeled second data to be labeled and / or the third data to be labeled that has not been labeled, so the second data to be labeled j should be removed from the current fourth data to be labeled.
[0129] In addition, for each fourth data to be labeled in the labeling candidate pool i (the second data to be labeled itself is also the fourth data to be labeled), except for the incorrectly labeled second data to be labeled j, if the probability of belonging to category i in its inference result is no greater than the probability of belonging to category i in the inference result of the second data to be labeled j, it indicates that the inference label of the fourth data to be labeled can only be the same, or even less reliable than the inference label of the second data to be labeled j (which is already an incorrect label). If it is retained in the fourth data to be labeled, it will make it difficult to further improve the labeling accuracy of the third data to be labeled, and thus such fourth data to be labeled can also be removed from the current fourth data to be labeled.
[0130] It is not difficult to see that, except for the removed data to be labeled, the remaining data in the fourth data to be labeled meets the conditions mentioned in step C1.
[0131] Figure 5 The fourth principle of reducing the data to be labeled is shown. Figure 5 The thin arrow in the figure indicates the second data to be labeled correctly (not limited to the second data to be labeled in this round), the thick arrow indicates the second data to be labeled incorrectly, and the third data to be labeled that has not been labeled is not shown.
[0132] Reference Figure 5 , assuming that the second data to be labeled j is the first wrongly labeled second data to be labeled from the left, the fourth data to be labeled is located on the left side of the second data to be labeled j, including several correctly labeled second data to be labeled and the third data to be labeled not shown, because the probability of belonging to category i in the inference result is greater than the probability of belonging to category i in the inference result of the second data to be labeled j ( Figure 5 The bottom is the changing trend of the probability of belonging to category i), so it is retained after reduction. The second data to be labeled j itself, and the part of the fourth data to be labeled located on the right side of the second data to be labeled j, including an incorrectly labeled second data to be labeled, several correctly labeled second data to be labeled, and several unlabeled data to be labeled that are not shown. Since the probability of belonging to category i in its inference result is not greater than the probability of belonging to category i in the inference result of the second data to be labeled j, it is removed after reduction, so that the second half of the rectangular box representing the fourth data to be labeled becomes a dotted line.
[0133] Note that according to the reduction method described above, when removing the incorrectly labeled second data to be labeled j, the second data to be labeled k may also be removed "incidentally". Therefore, in step C1, it is also required to confirm whether the second data to be labeled j is still included in the fourth data to be labeled in the labeling candidate pool i before reduction to avoid repeated removal.
[0134] Reference Figure 5, assuming that the second data to be labeled k currently to be processed is the second wrongly labeled second data to be labeled from the left. Since the second data to be labeled k has been removed when processing the second data to be labeled j before, the second data to be labeled k is currently not in the fourth data to be labeled, and will not cause the fourth data to be labeled to be reduced.
[0135] In addition, the fourth data to be labeled that is removed during the reduction process does not necessarily have to be deleted from the labeled data pool i, and it may also continue to remain in the labeled data pool i.
[0136] In the current round of iteration steps after step C1, all references to the fourth data to be labeled refer to the fourth data to be labeled obtained after reduction. For the sake of simplicity, it is still referred to as the fourth data to be labeled.
[0137] Step D1: Determine whether the fourth data to be labeled in the labeled data pool i contains the third data to be labeled that has not been labeled.
[0138] If the fourth data to be labeled contains the third data to be labeled that has not been labeled, then continue to execute step E1. If the fourth data to be labeled does not contain the third data to be labeled that has not been labeled, that is, the fourth data to be labeled only contains the second data to be labeled that is correctly labeled, then automatic labeling cannot be performed at this time, so the iterative process can be jumped out, and the labeling of the labeled data pool i is ended.
[0139] Step E1: Estimate whether the labeling accuracy of the third data to be labeled meets the standard according to the number of the second data to be labeled that are correctly labeled in the fourth data to be labeled in the labeled data pool i.
[0140] Here is a formula to estimate whether the annotation accuracy meets the standard: 1 / (1+w 2 / Ci)>acc.
[0141] Among them, the left side of the greater than sign is the estimation formula of the annotation accuracy, and the right side acc is the target accuracy mentioned when explaining step S130. If the left side is greater than the right side, it means that the annotation accuracy meets the standard. In the expression on the left, Ci is the number of correctly labeled second data to be labeled in the fourth data to be labeled in the labeling candidate pool i (note that this Ci second data to be labeled is not necessarily in this round. If the standard correct second data to be labeled in the previous iterations still belongs to the fourth data to be labeled in this round, it also contributes to Ci), w is a preset value related to the confidence level. The larger the value of w, the higher the confidence level of the estimated accuracy. For example, when w=2, the confidence level is 95%.
[0142] It should be understood that the estimation of the labeling accuracy may also use other formulas, which are not limited in this application. In addition, judging whether the labeling accuracy meets the standard does not mean that the labeling accuracy estimate must be explicitly calculated. For example, if the above formula is deformed, we can get Ci>w 2 / (1 / acc-1), that is, judgment can also be made by comparing Ci with a certain threshold.
[0143] If the labeling accuracy of the third data to be labeled has reached the standard, then continue to execute step F1; if the labeling accuracy of the third data to be labeled has not reached the standard, then the next round of iteration can be started, that is, jump to step A1 to continue execution.
[0144] Step F1: Determine category i as the labeling category of the third data to be labeled.
[0145] According to the description in step C1, category i is the inference label of the third data to be labeled, and the label category of the third data to be labeled is the label label of the third data to be labeled mentioned in step S260. After step F1 is executed, the iteration process can be exited to end the labeling of the labeled data pool i.
[0146] Optionally, when creating a candidate pool for annotation, method 1 may further sort the first data to be annotated in each candidate pool for annotation in descending order according to the probability of belonging to the category corresponding to the candidate pool in the inference result. In this case, step C1 may be implemented as follows:
[0147] For each second data to be labeled obtained in step A1, if the labeling category of the second data to be labeled j is different from category i, and the sorting index of the second data to be labeled j is not greater than the maximum sorting index of the fourth data to be labeled in the labeling candidate pool i, then the fourth data to be labeled in the labeling candidate pool i is reduced to only the data to be labeled that meets the following conditions: its sorting index is less than the sorting index of the second data to be labeled j.
[0148] Among them, the sorting index of the data to be labeled can refer to its sorting number in the labeling candidate pool i. Since the first data to be labeled in the labeling candidate pool i is sorted in descending order according to the probability corresponding to category i in its reasoning result, if the sorting index of a fourth data to be labeled is smaller than the sorting index of the second data to be labeled j, it indicates that the probability of belonging to category i in its reasoning result is greater than the probability of belonging to category i in the reasoning result of the second data to be labeled j, which is consistent with the condition in step C1.
[0149] It is not difficult to see that, according to the above method, the distribution range of the sorting index of the fourth data to be labeled is reduced in one direction. For example, if the sorting index is counted incrementally from 0, then at the beginning, the distribution range of the sorting index of the fourth data to be labeled is [0, m-1], m is the number of the first data to be labeled, after one round of iteration, the distribution range of the sorting index of the fourth data to be labeled is [0, m1-1], m1≤m, after two rounds of iteration, the distribution range of the sorting index of the fourth data to be labeled is [0, m2-1], m2≤m1, and so on.
[0150] Furthermore, to avoid repeated removal of data, step C1 also requires that the condition "the second data to be labeled j is included in the fourth data to be labeled in the labeling candidate pool i" should be met before reduction. For the case where the labeling candidate pool is sorted in descending order according to probability, this condition can be equivalent to that the sorting index of the second data to be labeled j is not greater than the maximum sorting index of the fourth data to be labeled in the labeling candidate pool i.
[0151] It can be seen that if the data in the annotation candidate pool is sorted first, when performing the reduction of the fourth data to be labeled, only a simple sorting index comparison is required instead of a probability value comparison, thereby improving processing efficiency.
[0152] Assuming that the sorting index of the data to be labeled starts counting from 0, Li represents the number of the fourth data to be labeled in the labeling candidate pool i (the maximum sorting index of the fourth data to be labeled is Li-1), the above implementation of step C1 can be further simplified as follows:
[0153] If the labeling category of the second data to be labeled j is different from the category i, the number Li of the fourth data to be labeled in the labeling candidate pool i is reduced to min(index(j,i),Li).
[0154] Among them, index(j,i) represents the sorting index of the second to-be-annotated data j in the annotation candidate pool i. The meaning of this formula is briefly explained below:
[0155] According to the foregoing, since the distribution range of the sorting index of the fourth data to be labeled is reduced in one direction during the reduction process, the number Li of the fourth data to be labeled is reduced to min(index(j,i),Li), which is equivalent to the distribution range of the sorting index of the fourth data to be labeled being reduced to [0, min(index(j,i),Li)-1]. Obviously, any number in the interval [0, min(index(j,i),Li)-1] is less than index(j,i), that is, the sorting index of the fourth data to be labeled after reduction is less than the sorting index of the second data to be labeled j.
[0156] Moreover, Li will only change when index(j,i)<Li or index(j,i)≤Li-1, otherwise Li will still maintain its original value. That is, reduction is possible only when the sorting index of the second data to be labeled j is not greater than the maximum sorting index Li-1 of the fourth data to be labeled in the labeling candidate pool i.
[0157] Optionally, before executing step B1, method 1 can also determine the predicted category (i.e., predicted label) of the second data to be labeled based on the inference result of the second data to be labeled, and display the second data to be labeled and its predicted category on the user labeling interface to improve the efficiency of user labeling.
[0158] If the first data to be labeled in the labeling candidate pool is sorted in descending order according to the probability of belonging to the category corresponding to the labeling candidate pool in its inference result, then in each round of iteration, the predicted category of the second data to be labeled can be determined in the following way, taking any second data to be labeled j as an example:
[0159] First, the sorting index of the second to-be-annotated data j in each annotation candidate pool is obtained, for example, index(j,i), where i ranges from 0 to nc-1, that is, a total of nc sorting indexes are obtained.
[0160] Then, each sorting index is normalized by the number of the fourth data to be labeled in the corresponding labeling candidate pool to obtain the normalized sorting index. For example, index(j,i) can be normalized to index(j,i) / Li. Note that since the step of giving the predicted label is performed before step B1, the reduction of the fourth data to be labeled in this iteration has not yet been performed, so Li should be the number of the fourth data to be labeled at the beginning of this iteration.
[0161] Finally, the category corresponding to the minimum sorting index among all normalized sorting indexes is determined as the predicted category of the second data to be labeled j. For example, argmin_i(index(j,i) / Li) is used as the predicted category of the second data to be labeled j. Here, argmin_i represents the value of i when the value in the brackets is calculated to be the minimum.
[0162] The prediction category determined in this way, since the sorting index is normalized by the number of the fourth data to be labeled in the labeling candidate pool, is equivalent to weakening the difference in distribution of data of different categories, which is beneficial to improving the accuracy of the prediction category and further improving the efficiency of manual labeling.
[0163] Method 2
[0164] Method 2 does not need to create a labeling candidate pool, or it can be considered that there is only one labeling candidate pool in Method 2, and all the first data to be labeled are in the labeling candidate pool. Each round of iteration may include the following steps:
[0165] Step A2: sampling the second data to be labeled from the fourth data to be labeled and provided to the user for labeling in this round of iteration.
[0166] The fourth data to be labeled at the beginning of the first round of iteration is the first data to be labeled. Step A2 is similar to step S220 and step A1, and will not be described again.
[0167] Step B2: Obtain the annotation category of the second data to be annotated that is annotated by the user.
[0168] Since the problem at this time is classification, the labeling category of the second data to be labeled is the labeling label of the second data to be labeled mentioned in step S230.
[0169] Step C2: For each second data to be labeled obtained in step A2, taking any second data to be labeled j as an example, if the labeling category of the second data to be labeled j is different from the category corresponding to the maximum probability in its reasoning result, and the second data to be labeled j is included in the fourth data to be labeled, then the fourth data to be labeled is reduced to only the data to be labeled that meets the following conditions: the maximum probability in its reasoning result is greater than the maximum probability in the reasoning result of the second data to be labeled j.
[0170] The reduction in step C2 is an implementation method of the screening in step 230. The inference labels of all first data to be labeled are the categories corresponding to the maximum value of the probability of the first data to be labeled belonging to each category in the inference result (referred to as the maximum probability in the inference result). Therefore, if the labeled category of the second data to be labeled j is different from the category corresponding to the maximum probability in its inference result, it means that the inference label of the second data to be labeled j is wrong. According to the definition of the fourth data to be labeled in the previous text, it can only contain the correctly labeled second data to be labeled and / or the third data to be labeled that has not been labeled, so the second data to be labeled j should be removed from the current fourth data to be labeled.
[0171] In addition, for each fourth data to be labeled (the second data to be labeled itself is also the fourth data to be labeled) except the incorrectly labeled second data to be labeled j, if the maximum probability in its inference result is not greater than the maximum probability in the inference result of the second data to be labeled j, it indicates that the inference label of the fourth data to be labeled can only be the same, or even less reliable than the inference label of the second data to be labeled j (which is already an incorrect label). If it is retained in the fourth data to be labeled, it will make it difficult to further improve the labeling accuracy of the third data to be labeled, and thus such fourth data to be labeled can also be removed from the current fourth data to be labeled.
[0172] It is not difficult to see that, except for the removed data to be labeled, the remaining data in the fourth data to be labeled meets the conditions mentioned in step C2.
[0173] In the current round of iteration steps after step C2, all references to the fourth data to be labeled refer to the fourth data to be labeled obtained after reduction. For the sake of simplicity, it is still referred to as the fourth data to be labeled.
[0174] Step D2: Determine whether the fourth data to be labeled contains the third data to be labeled that has not been labeled.
[0175] If the fourth data to be labeled contains the third data to be labeled that has not been labeled, then continue to execute step E2. If the fourth data to be labeled does not contain the third data to be labeled that has not been labeled, that is, the fourth data to be labeled only contains the second data to be labeled that is correctly labeled, then automatic labeling cannot be performed at this time, so the iterative process can be jumped out and the labeling of the first data to be labeled is ended.
[0176] Step E2: Estimate whether the labeling accuracy of the third data to be labeled meets the standard according to the number of the second data to be labeled that are correctly labeled in the fourth data to be labeled.
[0177] The formula for estimating whether the annotation accuracy meets the standard can be: 1 / (1+w 2 / C)>acc. Wherein, C is the number of correctly annotated second data to be annotated in the fourth data to be annotated (note that these C second data to be annotated are not necessarily in this round. If the second data to be annotated that are correct in the previous iterations still belong to the fourth data to be annotated in this round, they also contribute to C), w is a preset value related to the confidence level, and acc is the target accuracy mentioned in the description of step S130. It is not difficult to see that this formula is the same as the previous formula 1 / (1+w 2 / Ci)>acc is similar, except that Ci is replaced by C, because method 2 can be regarded as having only one annotation candidate pool.
[0178] It should be understood that other formulas may be used to determine whether the labeling accuracy meets the standard, and this application is not limited to this.
[0179] If the labeling accuracy of the third data to be labeled has reached the standard, then continue to execute step F2; if the labeling accuracy of the third data to be labeled has not reached the standard, then the next round of iteration can be started, that is, jump to step A2 to continue execution.
[0180] Step F2: Determine the category corresponding to the maximum probability in the inference result of the third data to be labeled as the labeled category of the third data to be labeled.
[0181] According to the description in step C2, the category corresponding to the maximum probability in the inference result of the third data to be labeled is the inference label of the third data to be labeled, and the labeling category of the third data to be labeled is the labeling label of the third data to be labeled mentioned in step S260. After step F2 is executed, the iterative process can be jumped out, and the labeling of the first data to be labeled is ended.
[0182] Optionally, before starting iteration, in method 2, the first data to be labeled may be sorted in descending order according to the maximum probability of the first data to be labeled belonging to each category in the inference result. In this case, step C2 may be implemented as follows:
[0183] For each second data to be labeled obtained in step A2, if the labeled category of the second data to be labeled j is different from the category corresponding to the maximum probability in its inference result, and the sorting index of the second data to be labeled j is not greater than the maximum sorting index in the fourth data to be labeled, then the fourth data to be labeled is reduced to only the data to be labeled that meets the following conditions: its sorting index is less than the sorting index of the second data to be labeled j.
[0184] Among them, the sorting index of the data to be labeled can refer to its sorting number in the first data to be labeled. Since the first data to be labeled is sorted in descending order according to the maximum probability in its reasoning result, if the sorting index of a fourth data to be labeled is smaller than the sorting index of the second data to be labeled j, it indicates that the maximum probability in its reasoning result is greater than the maximum probability in the reasoning result of the second data to be labeled j, that is, it is consistent with the condition in step C2. It is not difficult to see that according to the above method, the distribution range of the sorting index of the fourth data to be labeled is reduced in one direction.
[0185] Furthermore, to avoid repeated removal of data, step C2 also requires that the condition "the second data to be labeled j is included in the fourth data to be labeled" should be met before reduction. For the case where the first data to be labeled is sorted in descending order according to probability, this condition can be equivalent to that the sorting index of the second data to be labeled j is not greater than the maximum sorting index of the fourth data to be labeled.
[0186] It can be seen that if the first data to be labeled are sorted first, when reducing the fourth data to be labeled, only a simple sorting index comparison is required instead of a probability value comparison, thereby improving processing efficiency.
[0187] Assuming that the sorting index of the data to be labeled is counted incrementally from 0, and L represents the number of the fourth data to be labeled (in this case, the maximum sorting index of the fourth data to be labeled is L-1), the above implementation of step C2 can be further simplified as follows:
[0188] If the labeled category of the second data to be labeled j is different from the category corresponding to the maximum probability in its inference result, the number of the fourth data to be labeled L is reduced to min(index(j), L). Wherein, index(j) represents the sorting index of the second data to be labeled j in the first data to be labeled. The formula can be understood by referring to the formula min(index(j,i), Li) in method 1, and will not be repeated.
[0189] Optionally, before executing step B2, method 2 can also determine the predicted category (i.e., predicted label) of the second data to be labeled based on the inference result of the second data to be labeled, and display the second data to be labeled and its predicted category on the user labeling interface to improve the efficiency of user labeling.
[0190] If the first data to be labeled is sorted in descending order according to the maximum probability in its inference results, the prediction category of the second data to be labeled can be determined in the following way (regardless of the iteration round), taking any second data to be labeled j as an example:
[0191] The category corresponding to the maximum probability in the inference result of the second data to be labeled j is determined as the predicted category of the second data to be labeled j. Determining the predicted category in this way is very simple and fast.
[0192] Simply comparing Method 1 and Method 2, Method 2 is simpler to implement and does not require the creation of multiple annotation candidate pools. The advantage of Method 1 is that there is a annotation candidate pool for each category of data, and the number of fourth data to be annotated in each annotation candidate pool is different, which is equivalent to adaptively setting a discrimination threshold for each category to determine which data can be annotated as that category, rather than setting the same discrimination threshold for each category (e.g., probability value 1 / nc), which can overcome the bias in the data to be annotated to a certain extent (i.e., the data distribution of each category is different) and improve the accuracy of annotation.
[0193] It should be understood that even in the case where the labeling model is a classification model, the iterative steps S220 to S260 are not limited to the two implementation methods of method 1 and method 2. For example, in a certain implementation method of method 2, the first data to be labeled is sorted in descending order according to the maximum probability in the inference result, and the maximum probability can also be replaced by other indicators.
[0194] When explaining the steps of method 1 and method 2, any parts not mentioned can be referred to Figure 1 and Figure 4 An explanation of the relevant steps in the method, or a reference to each other.
[0195] According to the previous explanation, Figure 4 The data labeling method in only gives a labeling accuracy estimate for the third data to be labeled that has not been labeled in the current fourth data to be labeled. For other data to be labeled, their labeling accuracy is unknown, and their labels are given only when the labeling accuracy of the third data to be labeled meets the standard. When the labeling accuracy does not meet the standard, the label of the third data to be labeled is unknown, and the labels of other data to be labeled that have not been labeled (filtered out in step S230) are also unknown.
[0196] That is, it cannot be guaranteed that every piece of data to be labeled has a current reasonable labeling accuracy and a current reasonable label at any time, which brings difficulties to some application scenarios: for example, after labeling for a period of time, the user has to stop labeling for some reason, and hopes to obtain as much data as possible with a labeling accuracy greater than 96% (the target accuracy is 99%) and its labels for training the task model based on the current labeling progress. However, according to the method introduced before, at most, it can only give an estimated labeling accuracy of the third piece of data to be labeled, which is difficult to meet user needs.
[0197] In order to solve the above problems, in an improved solution, two attributes are added to each piece of data to be labeled: the labeling accuracy attribute and the labeling label attribute. The labeling accuracy attribute stores the current labeling accuracy estimate of the data to be labeled, and the labeling label attribute stores the current label estimate of the data to be labeled. In this way, users can end labeling at any time, and filter whether each piece of data to be labeled can be used as a training sample for the task model based on the value of the labeling accuracy attribute, and label the training sample based on the value of its labeling label attribute.
[0198] Next, we will use the steps in method 1 to illustrate how to calculate the labeling accuracy attribute and labeling label attribute of each piece of data to be labeled. Since method 1 is aimed at classification problems, the labeling label attribute can also be renamed as the labeling category attribute.
[0199] First, perform the initialization step. Before performing the iteration step, you can set initial values for the two attributes of the first data to be labeled. For example, for the labeling accuracy attribute, its initial value can be set to 0, 1 / nc, etc. For the labeling label attribute, its initial value can be set to unknown label (which can be regarded as a special label), or it can also be set to its inference label, etc.
[0200] Note that if these two items of a certain first data to be labeled already have values and are not initial values, there is no need to initialize them. The reason is that these data to be labeled may have been labeled before (such as a previous batch of first data to be labeled), and the values of these two attributes were calculated during the previous labeling, but they were reduced during the previous labeling and were not given the final labeling tags. Therefore, they are added to the current batch of first data to be labeled and re-labeled (because the labeling model may have been optimized at this time, its inference results for the same data to be labeled may also be different from the previous ones, and re-labeling may produce new results).
[0201] After executing step B1, the value of the annotation accuracy attribute of each second to-be-annotated data may be set to the first value, and the value of its annotation category attribute may be set to the annotation category given by the user. The first value may be a preset value, for example, if the annotation category given by the user is assumed to be correct, the first value may be 1, and of course other values may be taken by the first value.
[0202] When executing step C1, some of the data to be labeled may be removed. The removed data to be labeled include three components: first, the second data to be labeled, whose labeling accuracy attribute value is the first value, and the labeling label attribute value is the label given by the user; second, the data that has been reduced in the first round of iteration and has not been labeled. Due to initialization, its labeling accuracy attribute and labeling label attribute are their respective initial values; third, the data that has been reduced in non-first round of iteration and has not been labeled. According to the following content, its labeling accuracy attribute and labeling label attribute may be assigned specific values.
[0203] After executing step D1, if the fourth data to be labeled includes the third data to be labeled that has not been labeled, the labeling accuracy of the third data to be labeled is obtained. If the labeling accuracy of the third data to be labeled is calculated when making a judgment in step E1, the calculation result in E1 can be directly obtained here. If the labeling accuracy of the third data to be labeled is not calculated when making a judgment in step E1 (as mentioned above, it is not necessary to explicitly calculate the labeling accuracy estimate to estimate whether the labeling accuracy meets the standard), then the labeling accuracy of the third data to be labeled needs to be calculated at this time.
[0204] For each third data to be labeled, taking any third data to be labeled s as an example, if the value of the labeling accuracy attribute of the third data to be labeled s is less than the obtained labeling accuracy, then the value of the labeling accuracy attribute of the third data to be labeled s is set to the obtained labeling accuracy, and the value of its labeling category attribute is set to category i (taking the case in the labeling candidate pool i as an example).
[0205] For example, according to the formula 1 / (1+w 2 / Ci) can estimate the labeling accuracy of the third data s to be labeled in the labeling candidate pool i, and record the value of the labeling accuracy attribute of the third data s to be labeled as p(s), and the value of the labeling category attribute as c(s). If p(s)<1 / (1+w 2 / Ci), then p(s) is set to 1 / (1+w 2 / Ci), set c(s) to i, otherwise the values of p(s) and c(s) remain unchanged. In particular, if p(s)>acc, the value of c(s) is the final label of the third data to be labeled.
[0206] Note that since the same piece of data s to be labeled will be retained in each labeling candidate pool, the labeling accuracy may be estimated in multiple labeling candidate pools, but according to the above formula, the maximum value p(s) in the calculation result will be automatically selected.
[0207] It should be understood that even Figure 4 The steps in are not implemented by method 1, and the values of the labeling accuracy attribute and the labeling category attribute of each first to-be-labeled data can also be calculated in a similar manner, which will not be elaborated in detail.
[0208] The present application also provides a data annotation system, which can be, but is not limited to, deployed in Fig. 9 The electronic device shown is used to execute the steps of the data labeling method and any implementation method thereof introduced above, Figure 6 The architecture and working principle of the data annotation system 300 are shown. The data annotation method provided in the embodiment of the present application will be further introduced in conjunction with the data annotation system 300.
[0209] Reference Figure 6 , the data annotation system 300 includes:
[0210] Two functional modules: a model training and reasoning module 310 and a labeling strategy module 330; wherein the functional module can be understood as a program module of a data labeling program;
[0211] Four data pools: a labeling result pool 322, an inference data candidate pool 324, a training data pool 326 and an inference result pool 328; wherein a data pool can be understood as a data set, or can also be understood as a storage space.
[0212] And, an interface: user annotation interface 340 .
[0213] Combination Figure 6 , the working engineering of the system is summarized as follows:
[0214] Step 1: The model training and reasoning module 310 samples at least zero batches of training data from the training data pool 326 to train the labeling model, and uses the labeling model obtained after training to reason about at least one batch of data to be labeled sampled from the reasoning data candidate pool 324 to obtain the reasoning results of at least one batch of data to be labeled.
[0215] The sampling method in step 1 is not limited. The training data pool 326 is the data source of the training data. The training data in the training data pool 326 includes training samples and their annotated labels (for example, images and their annotated categories). How to obtain the training data will be described later. A batch of data includes at least one piece of data. If the sampled training data is a zero batch, it actually means that no data sampling is performed at this time, and thus no training of the annotation model is performed. The original annotation model continues to be used. It is just for the sake of simplicity in expression that the case of extracting zero batches of training data is combined with the case of extracting at least one batch of training data. It is possible to extract zero batches of training data. For example, it is believed that the performance of the current annotation model is good enough and there is no need to adjust its parameters to avoid wasting computing resources. The sampled training data does not need to be removed from the training data pool 326. Regarding the training process of the annotation model, reference can be made to the prior art, which is not explained in detail here.
[0216] The inference data candidate pool 324 is the data source of the data to be labeled, and the data to be labeled is unlabeled data. The sampled data to be labeled does not need to be temporarily removed from the inference data candidate pool 324. Regarding the inference process of the labeling model, reference can be made to the prior art, which will not be explained in detail here.
[0217] Optionally, when the data annotation system 300 is initialized, the model training and reasoning module 310 will initialize the annotation model, including building the structure of the model, setting the initial parameters of the model, setting the loss function used by the model, etc.
[0218] Then, the model training and reasoning module 310 can write all the data to be labeled into the reasoning data candidate pool 324, so that the initial data in the reasoning data candidate pool 324 is available, so that the data sampling mentioned in step 1 can be supported. Then, the labeling strategy module 330 can read a preset number (for example, 30) of data to be labeled from the reasoning data candidate pool 324, and directly hand over these data to be labeled to the user on the user labeling interface 340 and obtain the labeling label returned by the user labeling interface 340. Note that this part of the data to be labeled does not contain a predicted label, because the labeling model has not been trained at this time and the predicted label cannot be given. Afterwards, the labeling strategy module 330 writes this part of the data to be labeled and its labeling label into the training data pool 326, so that the initial data in the training data pool 326 is available, so that the data sampling mentioned in step 1 can be supported.
[0219] Figure 7 This shows an implementation method for model training and reasoning in step 1. Figure 7 Step 1 can be implemented as at least one model training / inference cycle, each model training / inference cycle includes at least zero model training (one model training corresponds to each batch of training data sampled from the training data pool 326), and at least zero model inference using the labeled model obtained after training (one model inference corresponds to each batch of data to be labeled sampled from the inference data candidate pool 324).
[0220] However, it should be noted that the number of training and inference in each model training / inference cycle cannot be zero, and at least one model inference should be included in all model training / inference cycles, otherwise the inference result of the data to be labeled cannot be obtained. For example, two model training / inference cycles can be implemented as 1 model training, 1 model inference, 1 model training, 1 model inference, or 0 model training, 1 model inference, 1 model training, 0 model inference, or 0 model training, 1 model inference, 0 model training, 2 model inferences, and so on.
[0221] Step 2: The model training and reasoning module 310 writes at least one batch of data to be labeled and its reasoning results in step 1 into the reasoning result pool 328.
[0222] Optionally, when the data annotation system 300 is initialized, the model training and reasoning module 310 may initialize the reasoning result pool 328 to be empty.
[0223] Step 3: The labeling strategy module 330 reads the first to-be-labeled data and its inference results from the inference result pool 328 .
[0224] The number of the first data to be labeled does not exceed the total number of the data to be labeled currently in the inference result pool 328 , and the read first data to be labeled and its inference result can be considered to be removed from the inference result pool 328 .
[0225] There are many possible implementations of step 3: in terms of the timing of reading, the labeling strategy module 330 can read the first data to be labeled and its inference results from the inference result pool 328 at fixed intervals, or wait until the labeling of the previous batch of first data to be labeled is completed (the end condition has been introduced above) and then read a new batch of first data to be labeled and its inference results from the inference result pool 328, etc. In terms of the amount of data read, the labeling strategy module 330 can read all the current data to be labeled and its inference results in the inference result pool 328 as the inference results of the first data to be labeled, or always read a fixed number of first data to be labeled and its inference results from the inference result pool 328 (unless the amount of data in the inference result pool 328 is insufficient), etc.
[0226] Steps 1 to 3 may also be considered as an implementation of step S110.
[0227] Step 4: The labeling strategy module 330 executes the labeling process of the first data to be labeled.
[0228] Step 4 corresponds to steps S120 to S140 . During the labeling process, the second data to be labeled and its predicted label (optional) need to be displayed on the user labeling interface 340 , and the labeling label of the second data to be labeled is obtained through the user labeling interface 340 .
[0229] For the second data to be annotated and its annotated label, the annotation strategy module 330 can write it into the training data pool 326 to increase the amount of training data therein, thereby improving the performance of the annotation model when step 1 is subsequently executed again. The third data to be annotated and its annotated label (obtained in step S140) that are automatically annotated can also be written into the training data pool 326, but this is not necessary, because after all, the automatic annotation results have not been manually confirmed, and the annotation accuracy is only estimated by the algorithm, which may have a negative impact on the performance of the annotation model.
[0230] For the second and third data to be annotated whose annotated labels have been determined, the annotation strategy module 330 can remove them from the inference data candidate pool 324 to avoid repeated inference and repeated annotation. However, for those data to be annotated whose annotated labels have not been determined, they can continue to be retained in the inference data candidate pool 324, allowing them to be inferred and annotated again. With the optimization of the annotation model, different results may be inferred for these data, so that they can be successfully annotated.
[0231] In addition, when the data annotation system 300 is initialized, the model training and reasoning module 310 can write all the data to be annotated and a special label (representing an unknown label) into the annotation result pool 322. When executing step 4, for the second data to be annotated and the third data to be annotated whose annotation labels have been determined, the annotation strategy module 330 can use their annotation labels to modify the corresponding special labels in the annotation result pool 322. Optionally, the annotation accuracy attribute and the annotation label attribute of the data to be annotated mentioned above can also be saved in the annotation result pool 322.
[0232] In an alternative solution, the labeling result pool 322 may also be initialized to be empty, and the labeling strategy module 330 is responsible for writing the second data to be labeled and the third data to be labeled whose labeling labels have been determined, as well as their labeling labels, into the labeling result pool 322 .
[0233] Steps 1 to 4 can be executed repeatedly. In one implementation (referred to as synchronous mode), when executing steps 3 to 4, steps 1 to 2 are not allowed to be executed. In another implementation (referred to as asynchronous mode), when executing steps 3 to 4, steps 1 to 2 are allowed to be executed simultaneously, that is, after the labeling strategy module 330 reads the first data to be labeled from the inference result pool 328, the model training and inference module 310 may write new data to be labeled and its inference results into the inference result pool 328 without waiting for the first data to be labeled to be labeled.
[0234] In the asynchronous mode, since the model training and reasoning module 310 and the labeling strategy module 330 can work independently according to their own logic, the labeling efficiency is higher than the synchronous mode. However, the asynchronous mode may cause a problem: although according to the previous introduction, for the second data to be labeled and the third data to be labeled whose label labels have been determined, the labeling strategy module 330 will remove them from the inference data candidate pool 324, but before removal, these data may be sampled from the inference data candidate pool 324 by the model training and reasoning module 310 and inferred, causing them to enter the inference result pool 328 again, thereby causing repeated labeling.
[0235] To improve this problem, for the second data to be annotated and the third data to be annotated whose annotated labels have been determined, the annotation strategy module 330 may also remove them from the inference data candidate pool 324 if they exist in the inference data candidate pool 324 .
[0236] Furthermore, in asynchronous mode, since the labeling strategy module 330 may write new training data to the training data pool 326 at any time, if step 1 is implemented in a training / inference cycle, the training data in the training data pool 326 may be different in different rounds of training / inference cycles.
[0237] It should be understood that executing steps 1 to 4 through the model training and reasoning module 310 and the labeling strategy module 330 is only an optional method, and the division of program modules is relatively free. For example, steps 1 to 4 can also be executed by the same module.
[0238] Next, based on the above embodiments, two marking modes provided in the embodiments of the present application are further introduced:
[0239] Optionally, the data annotation process can be performed in two annotation modes, one is a semi-automatic annotation mode, and the other is a manual annotation mode. Among them, the semi-automatic annotation mode refers to a mode in which user annotation is combined with automatic annotation. For example, the data annotation method introduced above is a semi-automatic annotation mode, in which the second data to be annotated is annotated by the user, and the third data to be annotated with an annotation accuracy that meets the standard is automatically annotated by the machine. The manual annotation mode refers to a mode that is completely annotated by the user, that is, all the data to be annotated are handed over to the user for annotation. An example of the manual annotation mode will be given later.
[0240] When the conditions for switching the annotation mode are met, the two annotation modes can be switched between each other. Specifically, when the conditions for switching the manual annotation mode are met, the semi-automatic annotation mode is switched to the manual annotation mode; when the conditions for switching the semi-automatic annotation mode are met, the manual annotation mode is switched to the semi-automatic annotation mode.
[0241] For example, if the third data to be labeled in a labeling candidate pool is automatically labeled, it is called the execution of a skip strategy. If M consecutive (M is greater than a certain numerical threshold) labeling candidate pools do not execute the skip strategy, but the labeling is ended because the fourth data to be labeled in the labeling candidate pool only contains the second data to be labeled correctly, it means that the data at this time is more difficult for the machine to label (or it means that the performance of the labeling model at this time does not meet the requirements of automatic labeling), and thus the semi-automatic labeling mode can be switched to the manual labeling mode.
[0242] However, it should be noted that even if the semi-automatic annotation mode is continued, the annotation can be completed. However, the annotation method in the manual annotation mode can usually be designed to be simpler than the semi-automatic annotation mode (for example, there is no need to create several annotation candidate pools), so that the annotation efficiency can be improved after the mode is switched.
[0243] For another example, if the labeling strategy module 330 reads all the data to be labeled and their inference results from the inference result pool 328 each time as the first data to be labeled and their inference results, then when the labeling strategy module 330 detects that the number of data to be labeled in the inference result pool is less than a certain quantity threshold, it can switch from the semi-automatic labeling mode to the manual labeling mode.
[0244] Because if the model training and reasoning module 310 performs reasoning at a roughly constant speed, the number of unlabeled data in the reasoning result pool 328 is small, indicating that the labeling has entered the final stage, and the remaining unlabeled data in the final stage is likely to be data that has not been successfully labeled in previous labeling processes (for example, has been reduced), that is, difficult data. It can be reasonably inferred that these data cannot trigger, or can only trigger a small amount of skip strategy, and it is better for the user to directly label them.
[0245] Wherein, if the labeling model is a classification model, and the classification result is nc categories, and method 1 is used for semi-automatic labeling, the quantity threshold can be, but is not limited to, set to (w 2 / (1 / acc-1))×nc×h, where h is a preset multiple, for example, 3, 5, etc.
[0246] Next, we will take the case where the annotation model is a classification model as an example to introduce a method for implementing the manual annotation mode:
[0247] For a batch of data to be labeled, assuming it is called the fifth data to be labeled, the inference result of the fifth data to be labeled includes the probability that the fifth data to be labeled belongs to each category. The labeling method of the fifth data to be labeled in the manual labeling mode is:
[0248] First, the fifth data to be labeled and its inference result are obtained, wherein the inference result of the fifth data to be labeled is obtained by inferring the fifth data to be labeled using the labeling model. For example, the fifth data to be labeled and its inference result can be read from the inference result pool 328.
[0249] Then, the fifth data to be labeled is sorted in descending order according to the maximum probability in its reasoning results, and the fifth data to be labeled and its predicted categories are provided to the user for labeling in batches (for example, 100 data in each batch) in ascending order of the sorting index; wherein the predicted category of each fifth data to be labeled is the category corresponding to the maximum probability in its reasoning result.
[0250] Obtain the labeling category of each batch of fifth data to be labeled by the user until all the fifth data to be labeled are labeled. Alternatively, after each batch of fifth data to be labeled is completed, recheck whether the semi-automatic labeling mode switching conditions are currently met. If the switching conditions are met, switch from manual labeling mode to semi-automatic labeling mode, and change the remaining fifth data to be labeled that have not yet been labeled to semi-automatic labeling.
[0251] The key to the above labeling method is to label in batches according to the sorting index. The labeling difficulty of each batch of data to be labeled is from low to high (the lower the probability in the inference result, the less sure the labeling model is about the category to which the data belongs, and thus the higher the labeling difficulty). The user first labels the low-difficulty data with higher efficiency. It is possible that after the low-difficulty data is labeled, the remaining data does not need to be labeled by the user because the semi-automatic labeling mode switching conditions are met. Therefore, users can avoid having to label the more difficult data entirely by themselves.
[0252] It should be understood that it is not necessary to use two marking modes. For example, the semi-automatic marking mode may be used all the time for marking.
[0253] In some application scenarios, it is hoped to distinguish between simple data sets and difficult data sets in the data to be labeled, so that the performance of the task model can be evaluated on these two data sets respectively, and then its optimization direction during training can be determined.
[0254] The data annotation method provided in the embodiment of the present application can be used to divide simple data sets into difficult data sets:
[0255] For example, data annotated in the semi-automatic annotation mode are considered to belong to the simple data set, and data annotated in the manual annotation mode are considered to belong to the difficult data set. For another example, after all the data to be annotated are annotated, the value of the annotation accuracy attribute is judged. If it is greater than a certain threshold, the data to be annotated is considered to belong to the simple data set, otherwise it is considered to belong to the difficult data set, and so on.
[0256] Figure 8 FIG. 4 shows a possible structure of a data labeling device 400 provided in an embodiment of the present application. Figure 8 , the data labeling device 400 includes:
[0257] The data acquisition unit 410 is used to acquire first data to be labeled and its inference result; wherein the inference result of the first data to be labeled is information obtained by predicting the first data to be labeled using the labeling model;
[0258] The manual labeling unit 420 is used to obtain a labeling label of the second data to be labeled by the user; wherein the second data to be labeled is part of the data to be labeled in the first data to be labeled;
[0259] The accuracy evaluation unit 430 is used to estimate whether the labeling accuracy of the third data to be labeled that has not been labeled in the first data to be labeled has reached the standard according to the labeling label of the second data to be labeled; wherein the labeling accuracy of the third data to be labeled is the accuracy of the inference label determined according to the inference result of the third data to be labeled;
[0260] The automatic labeling unit 440 is configured to determine the inferred label of the third data to be labeled as the labeling label of the third data to be labeled when the labeling accuracy of the third data to be labeled has reached a standard.
[0261] The data labeling device 400 provided in the embodiment of the present application, its implementation principle and the technical effects produced have been introduced in the aforementioned method embodiment, and it can implement all the contents in the aforementioned method embodiment. For the sake of brief description, the specific implementation process of the functions corresponding to each unit included in the device can refer to the corresponding content in the aforementioned method embodiment, which will not be repeated here.
[0262] Fig. 9 The structure of the electronic device 500 provided in the embodiment of the present application is shown. Fig. 9 The electronic device 500 includes: a processor 510, a memory 520 and a communication interface 530. These components are interconnected and communicate with each other through a communication bus 540 and / or other forms of connection mechanisms (not shown).
[0263] Among them, the processor 510 includes one or more (only one is shown in the figure), which can be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 510 can be a general-purpose processor, including a central processing unit (CPU), a micro control unit (MCU), a network processor (NP) or other conventional processors; it can also be a dedicated processor, including a graphics processor (GPU), a neural network processor (NPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. In addition, when there are multiple processors 510, some of them can be general-purpose processors and the other part can be dedicated processors.
[0264] The memory 520 includes one or more (only one is shown in the figure), which can be, but not limited to, random access memory (Random Access Memory, RAM), read only memory (Read Only Memory, ROM), programmable read-only memory (Programmable Read-Only Memory, PROM), erasable programmable read-only memory (Erasable Programmable Read-Only Memory, EPROM), electrically erasable programmable read-only memory (Electric Erasable Programmable Read-Only Memory, EEPROM), etc.
[0265] The processor 510 and other possible components can access the memory 520, read and / or write data therein. In particular, one or more computer program instructions can be stored in the memory 520, and the processor 510 can read and execute these computer program instructions to implement the data labeling method provided in the embodiment of the present application.
[0266] The communication interface 530 includes one or more (only one is shown in the figure), which can be used to communicate directly or indirectly with other devices to exchange data. The communication interface 530 can include an interface for wired and / or wireless communication.
[0267] Understandably, Fig. 9 The structure shown is for illustration only. The electronic device 500 may also include Fig. 9 More or fewer components as shown, or with Fig. 9 For example, when the electronic device 500 does not need to communicate with other devices, it may not include the communication interface 530.
[0268] Fig. 9 Each component shown in can be implemented by hardware, software or a combination thereof. The electronic device 500 may be a physical device, such as a server, a PC, a laptop, a tablet computer, a mobile phone, etc., or a virtual device, such as a virtual machine, a container, etc. Moreover, the electronic device 500 is not limited to a single device, but may also be a combination of multiple devices or a cluster consisting of a large number of devices.
[0269] The present application also provides a computer-readable storage medium on which computer program instructions are stored. When these computer program instructions are read and executed by a processor, the data labeling method provided by the present application is executed. For example, the computer-readable storage medium can be implemented as Fig. 9 The memory 520 in the electronic device 500.
[0270] An embodiment of the present application also provides a computer program product, which includes computer program instructions. When these computer program instructions are read and executed by a processor, the data labeling method provided by the embodiment of the present application is executed.
[0271] The above description is only an embodiment of the present application and is not intended to limit the protection scope of the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A data labeling method, characterized in that: include: Acquire first data to be labeled and its inference result; wherein the inference result of the first data to be labeled is information obtained by predicting the first data to be labeled using a labeling model, and the labeling model is an image classification model or an image regression model; Obtaining a label for second data to be labeled that is labeled by a user; wherein the second data to be labeled is part of the data to be labeled in the first data to be labeled; Estimate whether the labeling accuracy of the third data to be labeled that has not been labeled in the first data to be labeled has reached the standard according to the labeling label of the second data to be labeled; wherein the labeling accuracy of the third data to be labeled is the accuracy of the inference label determined according to the inference result of the third data to be labeled; If the labeling accuracy of the third data to be labeled has reached the standard, the inference label of the third data to be labeled is determined as the labeling label of the third data to be labeled.
2. The data labeling method according to claim 1, characterized in that: The estimating whether the labeling accuracy of the third data to be labeled in the first data to be labeled has reached the standard according to the labeling label of the second data to be labeled, comprises: Determine fourth data to be labeled in the first data to be labeled according to the label of the second data to be labeled; wherein the fourth data to be labeled includes the second data to be labeled correctly labeled and / or the third data to be labeled that has not been labeled, and the second data to be labeled correctly labeled means that: the label of the second data to be labeled is the same as the inference label of the second data to be labeled determined according to the inference result of the second data to be labeled; If the fourth data to be labeled includes third data to be labeled that has not been labeled, then it is estimated whether the labeling accuracy of the third data to be labeled meets the standard according to the number of correctly labeled second data to be labeled in the fourth data to be labeled.
3. The data labeling method according to claim 2, characterized in that: The steps of obtaining the label of the second data to be labeled by the user until determining the inference label of the third data to be labeled as the label of the third data to be labeled are performed in an iterative manner, and each round of iteration includes the following steps: Sampling the second data to be labeled provided to the user for labeling in this round of iteration from the fourth data to be labeled; wherein, for the first round of iteration, the fourth data to be labeled is the first data to be labeled; Acquire the annotation label of the second data to be annotated by the user, and filter the fourth data to be annotated according to the annotation label of the second data to be annotated; After screening, if the fourth data to be labeled contains third data to be labeled that have not been labeled, then according to the number of correctly labeled second data to be labeled in the fourth data to be labeled, it is estimated whether the labeling accuracy of the third data to be labeled meets the standard; If the labeling accuracy of the third data to be labeled has reached the standard, the inference label of the third data to be labeled is determined as the labeling label of the third data to be labeled.
4. The data labeling method according to claim 3, characterized in that: The labeling model is an image classification model, and the inference result of the first data to be labeled includes the probability that the first data to be labeled belongs to each category; Before executing the first round of iteration, the method further includes: creating a labeling candidate pool corresponding to each category according to the inference result of the first data to be labeled; wherein each labeling candidate pool contains the probability that the first data to be labeled belongs to the category corresponding to the labeling candidate pool in the first data to be labeled and its inference result; Each iteration includes the following steps: The second data to be labeled provided to the user for labeling in this round of iteration is sampled from the fourth data to be labeled in all labeling candidate pools; wherein, for the first round of iteration, the fourth data to be labeled in each labeling candidate pool is the first data to be labeled in the labeling candidate pool; Acquire the annotation category of the second data to be annotated by the user; wherein the annotation category of the second data to be annotated is the annotation label of the second data to be annotated; For each annotation candidate pool, perform the following steps: For each second data to be labeled, if the labeling category of the second data to be labeled is different from the category corresponding to the labeling candidate pool, and the second data to be labeled is included in the fourth data to be labeled in the labeling candidate pool, then the fourth data to be labeled in the labeling candidate pool is reduced to only the data to be labeled that meets the following conditions: the probability of its inference result belonging to the category corresponding to the labeling candidate pool is greater than the probability of the inference result of the second data to be labeled belonging to the category corresponding to the labeling candidate pool; After the fourth data to be labeled in the labeling candidate pool is reduced, the fourth data to be labeled includes the third data to be labeled that has not been labeled, then according to the number of the second data to be labeled that are correctly labeled in the fourth data to be labeled, it is estimated whether the labeling accuracy of the third data to be labeled meets the standard; If the labeling accuracy of the third data to be labeled has reached the standard, the category corresponding to the labeling candidate pool is determined as the labeling category of the third data to be labeled; wherein, the category corresponding to the labeling candidate pool is the inference label of the third data to be labeled, and the labeling category of the third data to be labeled is the labeling label of the third data to be labeled.
5. The data labeling method according to claim 4, characterized in that: In each annotation candidate pool, the first to-be-annotated data are sorted in descending order according to the probability of belonging to the category corresponding to the annotation candidate pool in the inference result thereof; If the second category to be labeled is different from the category corresponding to the labeling candidate pool, and the second data to be labeled is included in the fourth data to be labeled in the labeling candidate pool, the fourth data to be labeled in the labeling candidate pool is reduced to only the data to be labeled that meets the following conditions: the probability of belonging to the category corresponding to the labeling candidate pool in the inference result is greater than the probability of belonging to the category corresponding to the labeling candidate pool in the inference result of the second data to be labeled, including: If the annotation category of the second data to be labeled is different from the category corresponding to the annotation candidate pool, and the sorting index of the second data to be labeled is not greater than the maximum sorting index of the fourth data to be labeled in the annotation candidate pool, then the fourth data to be labeled in the annotation candidate pool is reduced to only the data to be labeled that meets the following conditions: its sorting index is less than the sorting index of the second data to be labeled.
6. The data labeling method according to claim 4 or 5, characterized in that: Each piece of the first data to be annotated includes an annotation accuracy attribute and an annotation category attribute. After obtaining the annotation category of the second data to be annotated annotated by the user, the method further includes: The value of the labeling accuracy attribute of each second to-be-labeled data is set to the first value, and the value of its labeling category attribute is set to the labeling category; For each annotation candidate pool, the following steps are also performed: After the fourth data to be labeled in the labeling candidate pool is reduced, if the fourth data to be labeled includes the third data to be labeled that has not been labeled, obtaining the labeling accuracy of the third data to be labeled; wherein the labeling accuracy is estimated based on the number of correctly labeled second data to be labeled in the fourth data to be labeled; For each third piece of data to be labeled, if the value of the labeling accuracy attribute of the third piece of data to be labeled is less than the labeling accuracy, the value of the labeling accuracy attribute of the third piece of data to be labeled is set to the labeling accuracy, and the value of its labeling category attribute is set to the category corresponding to the labeling candidate pool.
7. The data annotation method according to claim 3, characterized in that: The labeling model is an image classification model, and the inference result of the first data to be labeled includes the probability that the first data to be labeled belongs to each category; Each iteration includes the following steps: Sampling the second data to be labeled provided to the user for labeling in this round of iteration from the fourth data to be labeled; wherein, for the first round of iteration, the fourth data to be labeled is the first data to be labeled; Acquire the annotation category of the second data to be annotated by the user; wherein the annotation category of the second data to be annotated is the annotation label of the second data to be annotated; For each second data to be labeled, if the labeling category of the second data to be labeled is different from the category corresponding to the maximum probability in its inference result, and the second data to be labeled is included in the fourth data to be labeled, then the fourth data to be labeled is reduced to only the data to be labeled that meets the following conditions: the maximum probability in its inference result is greater than the maximum probability in the inference result of the second data to be labeled; After the fourth data to be labeled is reduced, if the fourth data to be labeled includes the third data to be labeled that has not been labeled, then according to the number of the second data to be labeled that are correctly labeled in the fourth data to be labeled, estimating whether the labeling accuracy of the third data to be labeled meets the standard; If the labeling accuracy of the third data to be labeled has reached the standard, the category corresponding to the maximum probability in the inference result of the third data to be labeled is determined as the labeling category of the third data to be labeled; wherein, the category corresponding to the maximum probability in the inference result of the third data to be labeled is the inference label of the third data to be labeled, and the labeling category of the third data to be labeled is the labeling label of the third data to be labeled.
8. The data labeling method according to claim 7, characterized in that: The first data to be labeled are sorted in descending order according to the maximum probability among the probabilities that the first data to be labeled belongs to each category in the inference results; If the labeling category of the second data to be labeled is different from the category corresponding to the maximum probability in the inference result, and the second data to be labeled is included in the fourth data to be labeled, the fourth data to be labeled is reduced to only the data to be labeled that meets the following conditions: the maximum probability in the inference result is greater than the maximum probability in the inference result of the second data to be labeled, including: If the labeled category of the second data to be labeled is different from the category corresponding to the maximum probability in its inference result, and the sorting index of the second data to be labeled is not greater than the maximum sorting index in the fourth data to be labeled, then the fourth data to be labeled is reduced to only the data to be labeled that meets the following conditions: its sorting index is less than the sorting index of the second data to be labeled.
9. The data labeling method according to claim 1, characterized in that: Before obtaining the annotation label of the second data to be annotated by the user, the method further includes: Determining a predicted label of the second data to be labeled according to the inference result of the second data to be labeled; The second data to be labeled and its predicted label are displayed on the user labeling interface.
10. The data labeling method according to claim 9, characterized in that: The labeling model is an image classification model, the inference result of the first data to be labeled includes the probability that the first data to be labeled belongs to each category, the first data to be labeled is labeled by performing at least one round of iteration, and before performing the first round of iteration, a labeling candidate pool corresponding to each category is created according to the inference result of the first data to be labeled, and in each labeling candidate pool, the first data to be labeled is sorted in descending order according to the probability of belonging to the category corresponding to the labeling candidate pool in the inference result; When executing each round of iteration, the predicted category of any second data to be labeled is determined in the following manner, where the predicted category of the second data to be labeled is the predicted label of the second data to be labeled: Obtain the sorting index of the second to-be-annotated data in each annotation candidate pool; Normalize each sorting index by the number of fourth to-be-annotated data in the corresponding annotation candidate pool to obtain a normalized sorting index; The category corresponding to the minimum sorting index among all normalized sorting indexes is determined as the predicted category of the second data to be labeled.
11. The data labeling method according to claim 9, characterized in that: The labeling model is an image classification model, the inference result of the first data to be labeled includes the probability that the first data to be labeled belongs to each category, and the first data to be labeled is labeled by performing at least one round of iteration; Determine the predicted category of any one piece of second data to be labeled in the following manner, where the predicted category of the second data to be labeled is the predicted label of the second data to be labeled: The category corresponding to the maximum probability in the inference result of the second data to be labeled is determined as the predicted category of the second data to be labeled.
12. The data labeling method according to any one of claims 9 to 11, characterized in that: The annotation model is an image classification model, the predicted label of the second data to be annotated is a predicted category of the second data to be annotated, and the displaying of the second data to be annotated and its predicted label on the user annotation interface includes: The second data to be labeled are classified and displayed on the user labeling interface according to the predicted category of the second data to be labeled.
13. The data labeling method according to claim 12, characterized in that: The second data to be annotated and its predicted category are displayed in the first area of the user annotation interface; The method further comprises: In response to a data selection operation for the second data to be annotated that is triggered in the first area, displaying the selected second data to be annotated in the second area of the user annotation interface, and displaying candidate categories of the selected second data to be annotated in the second area; If a piece of the second data to be labeled is selected by the user, it indicates that the predicted category of the second data to be labeled is not recognized by the user.
14. The data labeling method according to claim 9, characterized in that: The second data to be annotated is an image, and the second data to be annotated and its predicted label are displayed in the first area of the user annotation interface; The method further comprises: Displaying a reference image in a third area of the user annotation interface, the reference image being an image selected from the second data to be annotated; Recording the image transformation operation triggered in the third region for the reference image, applying the image transformation operation to each image in the second data to be annotated, and refreshing and displaying the second data to be annotated after the image transformation operation is performed in the first region; The image transformation operation includes at least one of image translation, image rotation, image scaling, and frame selection of a local area of an image.
15. The data labeling method according to claim 1, characterized in that: The obtaining of the first to-be-annotated data and the inference result thereof includes: Sampling at least zero batches of training data from a training data pool to train the labeling model, and using the labeling model to infer at least one batch of data to be labeled sampled from an inference data candidate pool to obtain an inference result of the at least one batch of data to be labeled; Writing the at least one batch of data to be labeled and the inference results thereof into an inference result pool; The first data to be labeled and its inference result are read from the inference result pool.
16. A computer program product, characterized in that The method comprises computer program instructions, and when the computer program instructions are read and executed by a processor, the method according to any one of claims 1 to 15 is executed.
17. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are read and executed by a processor, the method according to any one of claims 1 to 15 is executed.
18. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores computer program instructions, and when the computer program instructions are read and executed by the processor, the method according to any one of claims 1 to 15 is executed.
Citation Information
Patent Citations
User tag extension labeling method and device, equipment and storage medium
CN113139141A
System and method for active machine learning
US20190318261A1