Data annotation method, device, computing device and storage medium
The AI platform displays difficult examples and their attributes in the unlabeled image set on the display interface, and users confirm the labeling results, which solves the problems of high data labeling time and labor costs and improves the optimization efficiency and labeling efficiency of the AI model.
Patent Information
- Application Number
- CN202010605206.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-29
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2040-06-29
AI Technical Summary
In existing technologies, data labeling requires high time and labor costs, resulting in low efficiency in AI model optimization.
The AI platform mines multiple difficult examples from the unlabeled image set and displays the difficult examples and their attributes to the user on the display interface. The user confirms the labeling results, reducing the need to label unlabeled images.
This improves data annotation efficiency, thereby improving the optimization efficiency of AI models. Users can more quickly identify and understand difficult examples, reducing the number of difficult examples to handle.
Smart Images

Figure CN113935389B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, computing device, and storage medium for data annotation. Background Art
[0002] With the development of Internet technology, artificial intelligence (AI) technology has been applied to more and more fields. At the same time, many AI-related problems such as data acquisition, data processing, data labeling, and model training need to be solved urgently. This has inspired people to establish AI platforms to provide developers and users with convenient AI development environments and convenient development tools. AI platforms can be applied to multiple scenarios, such as image classification and object detection. Among them, image classification is an artificial intelligence technology that classifies and recognizes images. Object detection is an artificial intelligence technology that accurately locates and classifies objects in images or videos. AI platforms require a large amount of sample data to train AI models.
[0003] In related technologies, to reduce the amount of sample data to be labeled, users first label a small amount of sample data. The AI platform then trains the AI model based on this small amount of labeled sample data. The AI platform then uses the semi-trained AI model to infer the labeled sample data and obtain inference results. The user then confirms the inference results one by one, obtaining a large amount of labeled sample data, which is then used to further optimize and train the AI model.
[0004] In this way, since each unlabeled sample data needs to be confirmed, the time and labor costs of data labeling are high, which in turn leads to low efficiency in AI model optimization. Summary of the Invention
[0005] This application provides a method for providing data annotation, which can reduce data annotation and improve the optimization efficiency of AI models.
[0006] In a first aspect, the present application provides a method for providing data annotation, the method comprising: determining multiple difficult examples and difficult example attributes of multiple difficult examples in an unlabeled image set, wherein the difficult example attributes include a difficult example coefficient; displaying at least one difficult example among the multiple difficult examples and the corresponding difficult example attributes to a user through a display interface; and obtaining the annotation result after the user confirms at least one difficult example based on the difficult example coefficients of the multiple difficult examples in the display interface.
[0007] The solution shown in the present application, the method can be executed by an AI platform, and the AI platform can mine difficult examples in the unlabeled image set to obtain multiple difficult examples and difficult example attributes of multiple difficult examples in the unlabeled image set. The AI platform displays at least one difficult example among the multiple difficult examples and the difficult example attributes corresponding to the at least one difficult example to the user through a display interface. The user can confirm the at least one difficult example displayed, and the AI platform can obtain the annotation results after confirmation. Since the AI platform can provide difficult examples to users, users only need to confirm the difficult examples without having to annotate a large number of unlabeled images, so the annotation efficiency can be improved, and then the efficiency of optimizing the AI model can be improved. It also provides users with difficult example attributes, so that users can better understand difficult examples and further improve annotation efficiency.
[0008] In one possible implementation, the attributes of difficult examples also include the causes of the difficult examples; displaying at least one difficult example among multiple difficult examples and the corresponding attributes of the difficult examples through a display interface also includes: displaying the causes of the difficult examples corresponding to at least one difficult example and suggestion information corresponding to the causes of the difficult examples through a display interface; wherein the suggestion information indicates a processing method that can be performed to reduce the difficulty of the corresponding causes of the difficult examples.
[0009] In the solution presented in this application, the attribute of a difficult example can also include the reason for the difficulty. When the AI platform displays a difficult example on the display interface, in addition to displaying the difficulty coefficient of at least one difficult example, it can also display the reason for at least one difficult example and the corresponding suggestion information. This can help users better understand the difficult examples, enable users to more quickly confirm the annotation results, and provide users with treatment methods to reduce the difficulty of the examples, so that users can take appropriate measures to quickly reduce the difficulty of the examples.
[0010] In a possible implementation, the hard case attribute further includes a hard case type, where the hard case type includes false positive, missed positive, or predicted normal, so that the hard case can be quickly distinguished.
[0011] In one possible implementation, displaying at least one of the multiple difficult examples and the corresponding difficult example attributes on a display interface includes: displaying different difficult example types using different display colors or different display lines on the display interface. In this way, difficult examples can be distinguished more quickly.
[0012] In one possible implementation, determining multiple hard examples and hard example attributes of the multiple hard examples in the unlabeled image set includes: determining multiple hard examples and hard example attributes of the multiple hard examples in the unlabeled image set according to a hard example mining algorithm.
[0013] In one possible implementation, after obtaining the labeling result after the user confirms at least one difficult example based on the difficulty coefficients of multiple difficult examples in the display interface, it also includes: obtaining a weight for the difficulty cause of each difficult example in the at least one difficult example; analyzing the labeling result, if the target difficult example in the at least one difficult example is confirmed as a correct difficult example by the user, then increasing the weight corresponding to the difficulty cause of the target difficult example; in the labeling result, if the target difficult example is confirmed as an incorrect difficult example by the user, then reducing the weight corresponding to the difficulty cause of the target difficult example; and updating the difficult example mining algorithm according to the updated weight.
[0014] In the solution shown in this application, each difficult example cause in the difficult example mining algorithm corresponds to a weight. After the AI platform obtains the annotation results confirmed by the user, the AI platform can analyze the annotation results. If the target difficult example in at least one difficult example is confirmed by the user as a correct difficult example (the correct difficult example is the difficult example that the user believes is a difficult example among the difficult examples mined by the difficult example mining algorithm), the weight corresponding to the difficult example cause of the target difficult example can be increased; in the annotation results, if the target difficult example is confirmed by the user as an incorrect difficult example (the incorrect difficult example is the difficult example that the user believes is not a difficult example among the difficult examples mined by the difficult example mining algorithm), the weight corresponding to the difficult example cause of the target difficult example can be reduced, and then the AI platform updates the difficult example mining algorithm according to the updated weight. In this way, the weights of each difficult example cause in the difficult example mining algorithm can be updated according to the annotation results confirmed by the user, so that the accuracy of the difficult example mining algorithm in mining difficult examples can be rapidly improved.
[0015] In one possible implementation, at least one difficult example among multiple difficult examples and the corresponding difficult example attributes are displayed to the user through a display interface, including: obtaining the filtering range and / or display order of the difficulty example coefficient selected by the user on the display interface; and displaying at least one difficult example and the corresponding difficult example attributes through the display interface according to the filtering range and / or display order of the difficulty example coefficient.
[0016] The solution shown in this application provides an AI platform that provides a method for filtering and displaying difficulty examples. Users can select the filtering range and / or display order of the difficulty coefficient in the display interface. The AI platform can then display at least one difficulty example and the corresponding difficulty attribute based on the filtering range and / or display order of the difficulty coefficient. This provides a more flexible display method.
[0017] In one possible implementation, the method further includes: when at least one difficult example is a plurality of difficult examples, displaying statistical information corresponding to the unlabeled image set through a display interface, wherein the statistical information includes one or more of the distribution information of non-difficult examples and difficult examples in the unlabeled image set, the distribution information of difficult examples for each difficult example cause in the unlabeled image set, the distribution information of difficult examples for each difficult example coefficient range in the unlabeled image set, or the distribution information of difficult examples for each difficult example type in the unlabeled image set. In this way, the user can determine the reasoning performance of the current AI model through the distribution information of non-difficult examples and difficult examples, or the user can understand the distribution of difficult examples in the causes of difficult examples, so that the user can take targeted measures to reduce difficult examples, etc.
[0018] In a second aspect, a method for updating a difficult example mining algorithm is provided, the method comprising: determining multiple difficult examples in a set of unlabeled images using a difficult example mining algorithm; displaying at least one of the multiple difficult examples to a user through a display interface; and updating the difficult example mining algorithm based on the labeling result after the user confirms at least one difficult example through the display interface.
[0019] The solution shown in this application can be executed by an AI platform. The AI platform can use a difficult example mining algorithm to determine multiple unlabeled difficult examples. The AI platform can display at least one difficult example from the multiple difficult examples to the user in a display interface. The user can confirm the at least one difficult example in the display interface. The AI platform obtains the annotation result after the user confirms the at least one difficult example in the display interface, and then the AI platform updates the difficult example mining algorithm based on the annotation result. In this way, since the difficult example mining algorithm can be reversely updated based on the annotation result after the user confirms, the accuracy of difficult example prediction can be made increasingly higher.
[0020] In one possible implementation, the method further includes: determining the causes of difficulty of multiple difficult examples using a difficult example mining algorithm, and providing suggestion information corresponding to the causes of difficulty to the user through a display interface, wherein the suggestion information indicates possible processing methods for reducing the difficulty of the corresponding causes of difficulty.
[0021] In the solution presented in this application, the AI platform can also use a difficult case mining algorithm to determine the reasons for the difficulty of multiple difficult cases, and display the reasons and corresponding suggestions in the display interface. This allows users to better understand the difficult cases and take appropriate measures to reduce them, making the trained AI model have better predictive performance.
[0022] In one possible implementation, the difficult example mining algorithm is updated based on the labeling results after the user confirms at least one difficult example through a display interface, including: obtaining a weight for the difficulty cause of each difficult example in at least one difficult example; analyzing the labeling results, if the target difficult example in at least one difficult example is confirmed as a correct difficult example by the user, then increasing the weight corresponding to the difficulty cause of the target difficult example; in the labeling results, if the target difficult example is confirmed as an incorrect difficult example by the user, then reducing the weight corresponding to the difficulty cause of the target difficult example; and updating the difficult example mining algorithm based on the updated weight.
[0023] In the solution shown in this application, each difficult example cause in the difficult example mining algorithm corresponds to a weight. After the AI platform obtains the annotation results confirmed by the user, the AI platform can analyze the annotation results. If the target difficult example in at least one difficult example is confirmed by the user as a correct difficult example (the correct difficult example is the difficult example that the user believes is a difficult example among the difficult examples mined by the difficult example mining algorithm), the weight corresponding to the difficult example cause of the target difficult example can be increased; in the annotation results, if the target difficult example is confirmed by the user as an incorrect difficult example (the incorrect difficult example is the difficult example that the user believes is not a difficult example among the difficult examples mined by the difficult example mining algorithm), the weight corresponding to the difficult example cause of the target difficult example can be reduced, and then the AI platform updates the difficult example mining algorithm according to the updated weight. In this way, the weights of each difficult example cause in the difficult example mining algorithm can be updated according to the annotation results confirmed by the user, so that the accuracy of the difficult example mining algorithm in mining difficult examples can be rapidly improved.
[0024] In one possible implementation, the method further includes: displaying, via a display interface, attributes of the at least one difficult case corresponding to the at least one difficult case, where the attributes include one or more of a difficulty coefficient, a reason for the difficulty, and a type of the difficult case. This allows the user to be presented with more information about the difficult case, enabling the user to better understand the difficult case and improving verification efficiency.
[0025] In a third aspect, a data labeling apparatus is provided, the apparatus comprising: a hard example mining module configured to determine a plurality of hard examples in an unlabeled image set and hard example attributes of the plurality of hard examples, wherein the hard example attributes include a hard example coefficient;
[0026] The user input / output (I / O) module is used to: display at least one of the multiple difficult examples and the corresponding difficult example attributes to the user through a display interface; obtain the annotation result after the user confirms the at least one difficult example in the display interface according to the difficulty coefficients of the multiple difficult examples. In this way, since difficult examples can be provided to users, users only need to confirm the difficult examples without having to annotate a large number of unannotated images, so the annotation efficiency can be improved, and then the efficiency of optimizing the AI model can be improved. In addition, the attributes of difficult examples are provided to users, so that users can better understand difficult examples and further improve the annotation efficiency.
[0027] In a possible implementation, the difficult example attribute further includes a difficult example reason; and the user I / O module is further configured to:
[0028] The display interface displays a difficulty cause corresponding to the at least one difficult example and suggestion information corresponding to the difficulty cause; wherein the suggestion information indicates a possible treatment method for reducing the difficulty of the example corresponding to the difficulty cause. This allows the user to better understand the difficult example, allowing the user to more quickly confirm the annotation results, and provides the user with treatment methods for reducing the difficulty of the example, allowing the user to take appropriate measures to quickly reduce the difficulty of the example.
[0029] In a possible implementation, the hard case attribute further includes a hard case type, where the hard case type includes false positive, missed positive, or predicted normal. In this way, it is possible to quickly distinguish which hard case it is.
[0030] In a possible implementation, the user I / O module is further configured to display different difficult example types using different display colors or different display lines via the display interface.
[0031] In a possible implementation, the hard example mining module is specifically configured to determine, according to a hard example mining algorithm, a plurality of hard examples in the unlabeled image set and hard example attributes of the plurality of hard examples.
[0032] In one possible implementation, the device further includes a model training module, which is used to: obtain the labeling result after the user confirms the at least one difficult example in the display interface based on the difficulty coefficients of the multiple difficult examples, and then obtain a weight for the difficulty reason of each difficult example in the at least one difficult example; analyze the labeling result, if the target difficult example in the at least one difficult example is confirmed as a correct difficult example by the user, then increase the weight corresponding to the difficulty reason of the target difficult example; if the target difficult example is confirmed as an incorrect difficult example by the user, then reduce the weight corresponding to the difficulty reason of the target difficult example; and update the difficult example mining algorithm based on the updated weight. In this way, the weight of each difficulty reason in the difficult example mining algorithm can be updated based on the labeling result after the user's confirmation, so that the accuracy of the difficult example mining algorithm in mining difficult examples can be rapidly improved.
[0033] In one possible implementation, the user I / O module is configured to: obtain a filter range and / or display order of a difficulty coefficient selected by a user on the display interface; and display the at least one difficulty example and the corresponding difficulty attribute on the display interface based on the filter range and / or display order of the difficulty coefficient. This provides a more flexible display method.
[0034] In one possible implementation, the user I / O module is further used to: when the at least one difficult example is one of the multiple difficult examples, display statistical information corresponding to the unlabeled image set through the display interface, wherein the statistical information includes one or more of the distribution information of non-difficult examples and difficult examples in the unlabeled image set, the distribution information of difficult examples for each difficult example cause in the unlabeled image set, the distribution information of difficult examples for each difficult example coefficient range in the unlabeled image set, or the distribution information of difficult examples for each difficult example type in the unlabeled image set. In this way, the user can determine the reasoning performance of the current AI model through the distribution information of non-difficult examples and difficult examples, or the user can understand the distribution of difficult examples in the causes of difficult examples, so that the user can take targeted measures to reduce difficult examples, etc.
[0035] In a fourth aspect, a device for updating a hard example mining algorithm is provided, the device comprising: a hard example mining module for determining a plurality of hard examples in an unlabeled image set using a hard example mining algorithm;
[0036] A user input / output I / O module, configured to display at least one difficult example from the plurality of difficult examples to a user via a display interface;
[0037] A model training module is used to update the difficult example mining algorithm according to the labeling result after the user confirms the at least one difficult example through the display interface.
[0038] In this way, since the difficult example mining algorithm can be reversely updated according to the annotation results after the user's confirmation, the accuracy of difficult example prediction can be made higher and higher.
[0039] In one possible implementation, the difficult example mining module is further configured to determine the difficulty causes of the multiple difficult examples using the difficult example mining algorithm, and the user I / O module is further configured to provide the user with suggestion information corresponding to the difficulty causes via the display interface, wherein the suggestion information indicates possible treatment methods for reducing the difficulty of the examples corresponding to the difficulty causes. In this way, the user can better understand the difficult examples and take appropriate measures to reduce them, thereby improving the predictive performance of the trained AI model.
[0040] In one possible implementation, the model training module is configured to: obtain a weight for the difficulty cause of each of the at least one difficult example; analyze the labeling results; if a target difficult example in the at least one difficult example is confirmed by the user as a correct difficult example, increase the weight corresponding to the difficulty cause of the target difficult example; if the target difficult example is confirmed by the user as an incorrect difficult example, decrease the weight corresponding to the difficulty cause of the target difficult example; and update the difficult example mining algorithm based on the updated weights. In this way, the weights of the various difficulty causes in the difficult example mining algorithm can be updated based on the labeling results confirmed by the user, thereby rapidly improving the accuracy of the difficult example mining algorithm in mining difficult examples.
[0041] In one possible implementation, the user I / O module is configured to display, via the display interface, attributes of the difficulty case corresponding to the at least one difficult case, where the attributes include one or more of a difficulty coefficient, a difficulty reason, and a difficulty type. This allows the user to be presented with more information about the difficult case, enabling the user to better understand the difficult case and improving verification efficiency.
[0042] In a fifth aspect, the present application also provides a computing device, comprising a memory and a processor, wherein the memory is used to store a set of computer instructions; the processor executes the set of computer instructions stored in the memory, so that the computing device executes the method provided in the first aspect or any possible implementation of the first aspect.
[0043] In a sixth aspect, the present application also provides a computing device, comprising a memory and a processor, wherein the memory is used to store a set of computer instructions; the processor executes the set of computer instructions stored in the memory, so that the computing device executes the method provided in the second aspect or any possible implementation of the second aspect.
[0044] In a seventh aspect, the present application provides a computer-readable storage medium storing computer program code. When the computer program code is executed by a computing device, the computing device performs the method provided in the aforementioned first aspect or any possible implementation of the first aspect. The storage medium includes, but is not limited to, volatile memory, such as random access memory, and non-volatile memory, such as flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0045] In an eighth aspect, the present application provides a computer-readable storage medium storing computer program code. When the computer program code is executed by a computing device, the computing device performs the method provided in the aforementioned second aspect or any possible implementation of the second aspect. The storage medium includes, but is not limited to, volatile memory, such as random access memory, and non-volatile memory, such as flash memory, a hard disk, or a solid-state drive.
[0046] Ninth aspect, the present application provides a computer program product, comprising computer program code. When the computer program code is executed by a computing device, the computing device performs the method provided in the aforementioned first aspect or any possible implementation of the first aspect. The computer program product may be a software installation package. When the method provided in the aforementioned first aspect or any possible implementation of the first aspect is required, the computer program product may be downloaded and executed on the computing device.
[0047] Tenth aspect, the present application provides a computer program product, comprising computer program code. When the computer program code is executed by a computing device, the computing device performs the method provided in the aforementioned second aspect or any possible implementation of the second aspect. The computer program product may be a software installation package. When the method provided in the aforementioned second aspect or any possible implementation of the second aspect is required, the computer program product may be downloaded and executed on a computing device. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A schematic diagram of the structure of an AI platform 100 provided in an embodiment of the present application;
[0049] Figure 2 A schematic diagram of an application scenario of an AI platform 100 provided in an embodiment of the present application;
[0050] Figure 3 A schematic diagram of the deployment of an AI platform 100 provided in an embodiment of the present application;
[0051] Figure 4 A schematic diagram of the structure of a computing device 400 for deploying the AI platform 100 provided in an embodiment of the present application;
[0052] Figure 5 A flowchart of a method for providing data annotation provided in an embodiment of the present application;
[0053] Figure 6 A schematic diagram of a data upload interface provided in an embodiment of the present application;
[0054] Figure 7 A schematic diagram of an interface for exporting difficult examples provided in an embodiment of the present application;
[0055] Figure 8 A schematic diagram of an interface for starting smart annotation provided in an embodiment of the present application;
[0056] Figure 9 A schematic diagram of a display interface for multiple difficult examples provided in an embodiment of the present application;
[0057] Figure 10 A schematic diagram of a display interface for a single difficult example provided in an embodiment of the present application;
[0058] Figure 11 A flowchart of a data annotation method provided in an embodiment of the present application;
[0059] Figure 12 A flowchart of a method for updating a difficult example mining algorithm provided in an embodiment of the present application;
[0060] Figure 13 A schematic diagram of the structure of a data annotation device provided in an embodiment of the present application;
[0061] Figure 14 A schematic diagram of the structure of a data annotation device provided in an embodiment of the present application;
[0062] Figure 15 A schematic diagram of the structure of a device for updating a difficult example mining algorithm provided in an embodiment of the present application;
[0063] Figure 16 A schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0064] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0065] Artificial intelligence (AI) is gaining popularity, and machine learning is a key enabler of AI. It's permeating various industries, including medicine, transportation, education, and finance. Not only technical professionals, but even those without AI expertise across various industries are eager to use AI and machine learning to accomplish specific tasks.
[0066] To facilitate understanding of the technical solutions and embodiments provided in this application, the following describes in detail concepts such as AI models, AI model training, difficult examples, difficult example mining, and AI platforms:
[0067] AI models,It is a type of mathematical algorithm model that uses machine learning ideas to solve practical problems. The AI model includes a large number of parameters and calculation formulas (or calculation rules). The parameters in the AI model are numerical values that can be obtained by training the AI model through a training image set. For example, the parameters of the AI model are the calculation formulas or weights of the calculation factors in the AI model. The AI model also contains some hyper parameters. Hyper parameters are parameters that cannot be obtained by training the AI model through a training image set. Hyper parameters can be used to guide the construction of the AI model or the training of the AI model. There are many types of hyper parameters. For example, the number of iterations of AI model training, the learning rate, the batch size, the number of layers of the AI model, and the number of neurons in each layer. In other words, the difference between the hyper parameters and parameters of the AI model is that the values of the hyper parameters of the AI model cannot be obtained by analyzing the training images in the training image set, while the values of the parameters of the AI model can be modified and determined based on the analysis of the training images in the training image set during the training process.
[0068] AI models vary widely, and a widely used category is the neural network model. Neural network models are mathematical algorithmic models that mimic the structure and function of biological neural networks (the central nervous system of animals). A neural network model can include multiple neural network layers with different functions, each with parameters and calculation formulas. Different layers within a neural network model have different names depending on the calculation formula or function. For example, a layer that performs convolution calculations is called a convolutional layer, and convolutional layers are often used to extract features from input signals (such as images). A neural network model can also be constructed by combining multiple existing neural network models. Neural network models with different structures can be used in different scenarios (such as classification and recognition) or provide different results when used in the same scenario. Differences in neural network model structure include one or more of the following: a different number of layers within the neural network model, a different order of layers, and different weights, parameters, or calculation formulas within each layer. The industry already has a variety of highly accurate neural network models for applications such as recognition or classification. Some neural network models can be trained with specific training image sets and then used independently to complete a task or combined with other neural network models (or other functional modules) to complete a task.
[0069] With the exception of neural network models, most other AI models need to be trained before they can be used to complete a task.
[0070] Training AI models,This refers to using existing samples to fit the AI model to the patterns in the existing samples through a certain method, thereby determining the parameters of the AI model. For example, training an AI model for image classification or detection and recognition requires preparing a training image set. Depending on whether the training images in the training image set are labeled (i.e., whether the images have a specific type or name), the training of the AI model can be divided into supervised training and unsupervised training. When supervising the AI model, the training images in the training image set used for training are labeled. When training the AI model, the training images in the training image set are used as the input of the AI model, and the labels corresponding to the training images are used as the reference for the output values of the AI model. A loss function is used to calculate the loss between the output value of the AI model and the labels corresponding to the training images. The parameters of the AI model are adjusted based on the loss value. The AI model is iteratively trained with each training image in the training image set, and the parameters of the AI model are continuously adjusted until the AI model can accurately output the same output value as the labels corresponding to the training images based on the input training images. When an AI model is trained unsupervised, the training images in the training image set are unlabeled. The training images in the training image set are fed sequentially into the AI model, which then gradually identifies the associations and underlying patterns between the training images in the training image set until it can determine or identify the type or features of the input images. For example, in clustering, after receiving a large number of training images, an AI model used for clustering can learn the features of each training image, as well as the associations and differences between them, and automatically classify the training images into multiple categories. Different AI models can be used for different task types. Some AI models can be trained only with supervised learning, some only with unsupervised learning, and some with both supervised and unsupervised learning. The trained AI model can then be used to complete a specific task. Generally speaking, AI models in machine learning require supervised learning. Supervised training allows the AI model to more specifically learn the associations between the training images in the labeled training image set and the corresponding annotations, resulting in a higher accuracy when the trained AI model is used to predict other input inference images.
[0071] The following example uses supervised learning to train a neural network model for image classification. To train a neural network model for image classification, images are first collected based on the task and a training image set is constructed. This training image set contains three categories of images: apple, pear, and banana. The collected training images are stored in three folders, each named according to their category. The folder names represent the labels of all images within the folders. After the training image set is constructed, a neural network model capable of image classification (such as a convolutional neural network (CNN)) is selected and fed into the CNN. The convolution kernels of each layer in the CNN extract and classify the images, and finally output a confidence score that the image belongs to each category. A loss function is used to calculate a loss value based on the confidence score and the corresponding image label. The parameters of each CNN layer are then updated based on the loss value and the CNN structure. The training process continues until the loss value output by the loss function converges or all images in the training image set have been used for training.
[0072] Loss Function , is a function used to measure the degree to which the AI model has been trained (that is, it is used to calculate the difference between the result predicted by the AI model and the real target). In the process of training the AI model, because we hope that the output of the AI model is as close as possible to the value we really want to predict, we can compare the predicted value of the current AI model based on the input image with the real target value (that is, the annotation of the input image), and then update the parameters in the AI model according to the difference between the two (of course, there is usually an initialization process before the first update, that is, pre-configuring the initial values for the parameters in the AI model). Each training uses the loss function to judge the difference between the value predicted by the current AI model and the real target value, and updates the parameters of the AI model until the AI model can predict the real target value or a value very close to the real target value, and then the AI model is considered to be trained.
[0073] After the AI model is trained, it can be used to perform inference on images and obtain inference results. For example, in an image classification scenario, the specific inference process is: the image is input into the AI model, the convolution kernels in each layer of the AI model extract features from the image, and based on the extracted features, the image's category is output. In target detection (also known as object detection), the image is input into the AI model, the convolution kernels in each layer of the AI model extract features from the image, and based on the extracted features, the location and category of the bounding box of each target in the image are output. In scenarios covering both image classification and target detection, the image is input into the AI model, the convolution kernels in each layer of the AI model extract features from the image, and based on the extracted features, the image's category is output, as well as the location and category of the bounding box of each target in the image. It should be noted that some AI models have stronger inference capabilities, while others have weaker ones. Strong inference capabilities of an AI model mean that when the AI model is used to infer an image, the accuracy of the inference results is greater than or equal to a certain value. The weak reasoning ability of the AI model means that when the AI model is used to reason about an image, the accuracy of the reasoning results is lower than a certain value.
[0074] Data Annotation , is the process of adding labels to unlabeled data in the corresponding scenario. For example, unlabeled data is unlabeled images. In the image classification scenario, the category of the unlabeled image is added. In the object detection scenario, the location information and category of the object in the unlabeled image are added.
[0075] Hard example, It refers to the input data of the model when the output result of the initial AI model or the trained AI model is wrong or has a high error rate during the training of the initial AI model or the reasoning of the trained AI model. For example, during the training of the AI model, the input data whose loss function value between the prediction result and the label result during training is greater than a certain threshold is regarded as a difficult example. During the reasoning process of the AI model, an image in the reasoning image set is input into the AI model, and the error rate of the output reasoning result is higher than the target threshold, then the corresponding input image is a difficult example. In one scenario, the AI model can also be used to intelligently annotate unlabeled images. The process of using the AI model for intelligent annotation is actually the reasoning process of the AI model. Input images that are incorrectly annotated or have a high error rate are determined to be difficult examples.
[0076] Difficult Example Mining , refers to the method of determining an image as a hard example.
[0077] AI platform,It is a platform that provides AI developers and users with a convenient AI development environment and convenient development tools. The AI platform has built-in various AI models or AI sub-models that solve different problems. The AI platform can search and establish applicable AI models based on user needs. Users only need to determine their needs in the AI platform and prepare the training image set according to the prompts and upload it to the AI platform. The AI platform can train an AI model for the user that can be used to achieve the user's needs. Alternatively, the user prepares his or her own algorithm and training image set according to the prompts and uploads it to the AI platform. Based on the user's own algorithm and training image set, the AI platform can train an AI model that can be used to achieve the user's needs. Users can use the trained AI model to complete their specific tasks.
[0078] The AI platform in the embodiment of the present application introduces difficult example mining technology, so that the AI platform forms a closed-loop process of AI model construction, training, reasoning, difficult example mining, retraining, and re-reasoning.
[0079] It should be noted that the AI model mentioned above is a general term, and AI models include deep learning models, machine learning models, etc.
[0080] Figure 1 This is a schematic diagram of the structure of the AI platform 100 in the embodiment of the present application. It should be understood that Figure 1 This is only an exemplary structural diagram of the AI platform 100, and this application does not limit the division of modules in the AI platform 100. Figure 1 As shown, the AI platform 100 includes a user input / output (I / O) module 101, a hard example mining module 102, a model training module 103, an inference module 104, and a data preprocessing module 105. Optionally, the AI platform may also include an AI model storage module 106 and a data storage module 107.
[0081] The following briefly describes the functions of each module in the AI platform 100:
[0082] User I / O module 101: Receives task objectives input or selected by the user, receives the user's training image set, and so on. User I / O module 101 is also used to receive user confirmation of difficult examples and obtain one or more annotated images from the user. Furthermore, user I / O module 101 is used to provide AI models to other users. User I / O module 101 can be implemented, for example, using a graphical user interface (GUI) or a command line interface (CLI). For example, the GUI may display that the AI platform 100 can provide users with various AI services (such as image classification and object detection). Users can select a task objective on the GUI, such as image classification, and can then upload multiple unannotated images to the AI platform's GUI. After receiving the task objective and multiple unannotated images, the GUI communicates with model training module 103. Based on the user's determined task objective, model training module 103 selects or searches for an AI model that can be used to complete the user's task objective. The user I / O module 101 is further configured to receive difficult examples output by the difficult example mining module 102 and provide a GUI for the user to confirm the difficult examples.
[0083] Optionally, the user I / O module 101 may also be configured to receive user input of expectations regarding the effectiveness of the AI model in achieving the task objective, for example, inputting or selecting that the accuracy of the final AI model for face recognition must be greater than 99%.
[0084] Optionally, the user I / O module 101 may also be used to receive an AI model input by a user, etc. For example, a user may input an initial AI model in the GUI based on their own task objectives.
[0085] Optionally, the user I / O module 101 may also be configured to receive surface features and deep features of unlabeled images in an unlabeled image set input by a user. In image classification scenarios, surface features include one or more of image resolution, image aspect ratio, image red, green, and blue (RGB) mean and variance, image brightness, image saturation, or image clarity. Deep features refer to abstract features of an image extracted using a convolution kernel in a feature extraction model (such as a CNN). For target detection scenarios, surface features include surface features of bounding boxes and surface features of images. Surface features of bounding boxes may include the aspect ratio of each bounding box in a single-frame image, the proportion of the area of each bounding box in a single-frame image to the image area, the degree of edge formation of each bounding box in a single-frame image, the stacking diagram of each bounding box in a single-frame image, the brightness of each bounding box in a single-frame image, or the blurriness of each bounding box in a single-frame image. Surface features of images may include one or more of the resolution of the image, the aspect ratio of the image, the mean and variance of the RGB values of the image, the brightness of the image, the saturation of the image, or the clarity of the image, the number of boxes in a single-frame image, or the variance of the area of boxes in a single-frame image. Deep features refer to the abstract features of an image extracted using convolution kernels in feature extraction models (such as CNN).
[0086] Optionally, the user I / O module 101 may also be configured to provide a GUI for users to label training images in the training image set.
[0087] Optionally, the user I / O module 101 may also be configured to provide various pre-built initial AI models for the user to select. For example, the user may select an initial AI model on the GUI based on their task objectives.
[0088] Optionally, the user I / O module 101 may also be used to receive various configuration information of the user regarding the initial AI model, training images in the training image set, etc.
[0089] The difficult example mining module 102 is used to determine difficult examples and difficult example attributes in the unlabeled image set received by the user I / O module 101. The difficult example mining module 102 can communicate with the reasoning module 104 and the user I / O module 101. The difficult example mining module 102 can obtain the reasoning results of the reasoning module 104 on the unlabeled images in the unlabeled image set from the reasoning module 104, and mine difficult examples and difficult example attributes in the unlabeled image set based on the reasoning results. The difficult example mining module 102 can also provide the mined difficult examples and difficult example attributes to the user I / O module 101, where the difficult example attributes are used to guide the user to label and confirm the difficult examples.
[0090] Optionally, the hard example mining module 102 may also be configured to obtain, from the user I / O module 101 , surface features and deep features of unlabeled images in the unlabeled image set input by the user.
[0091] Model training module 103: used to train the AI model. The model training module 103 can communicate with the user I / O module 101, the inference module 104, and the AI model storage module 106. The specific process is as follows:
[0092] The initial AI model includes an untrained AI model and an AI model that has been trained but not optimized based on difficult examples. An untrained AI model refers to an AI model that has not been trained using a training image set, and the parameters in the constructed AI model are all preset values. An AI model that has been trained but not optimized based on difficult examples refers to an AI model that can be used for reasoning but has not been optimized based on difficult examples. It can include two types, one is the initial AI model selected directly by the user in the AI model storage module 105, and the other is an AI model obtained by training the constructed AI model using only the labeled training images in the training image set. It can be seen that the AI platform can obtain the initial AI model from the AI model storage module 106, or use the model training module 103 to train the training image set to obtain the initial AI model.
[0093] The initial AI model is an AI model obtained by training the constructed AI model using only the labeled training images in the training image set. The specific process is as follows: the AI platform determines the constructed AI model for the user to complete the user's task goal based on the user's task goal. The model training module 103 can communicate with both the user I / O module 101 and the AI model storage module 106. The model training module 103 selects a ready-made AI model from the AI model library stored in the AI model storage module 106 based on the user's task goal as the constructed AI model. Alternatively, the model training module 103 searches for the AI sub-model structure in the AI model library based on the user's task goal, the user's expected effect on the task goal, or some configuration parameters entered by the user, and specifies some AI model hyperparameters, such as the number of model layers, the number of neurons in each layer, etc., to construct the AI model and ultimately obtain a constructed AI model. It is worth noting that some of the AI model's hyperparameters can be hyperparameters determined by the AI platform based on its experience in constructing and training the AI model.
[0094] The model training module 103 obtains the training image set from the user I / O module 101. The model training module 103 determines some hyperparameters for training the constructed AI model based on the characteristics of the training image set and the structure of the constructed AI model. For example, the number of iterations, learning rate, batch size, etc. After setting the hyperparameters, the model training module 103 uses the annotated images in the acquired training image set to perform automatic training on the constructed AI model, and continuously updates the parameters within the constructed AI model during the training process to obtain the initial AI model. It is worth noting that some hyperparameters for training the constructed AI model can be hyperparameters determined by the AI platform based on the experience of model training.
[0095] The model training module 103 inputs the unlabeled images in the training image set into the initial AI model and outputs the inference results for the unlabeled images. The model training module 103 transmits the inference results to the hard example mining module 102. Based on the inference results, the hard example mining module 102 mines hard examples and hard example attributes in the unlabeled images. The hard example mining module 102 feeds back the determined hard examples and hard example attributes to the user I / O module 101. The user I / O module 101 obtains the hard examples annotated and confirmed by the user and feeds them back to the model training module 103. The model training module 103 continues to optimize and train the initial AI model based on the hard examples provided by the user I / O module 101, obtaining an optimized AI model, and optimizes the hard example mining module 102. The model training module 103 provides the optimized AI model to the inference module 104 for inference processing. It should be noted that if the initial AI model is the initial AI model stored in the AI model storage module 106, the training images in the training image set can be all unlabeled images. If the initial AI model is a constructed AI model, the training images in the training image set include some unlabeled images and some labeled images.
[0096] The reasoning module 104 uses the optimized AI model to reason about the unlabeled images in the unlabeled image set, and outputs the reasoning results of the unlabeled images in the unlabeled image set. The difficult example mining module 102 obtains the reasoning results from the reasoning module 104, and based on the reasoning results, determines the difficult examples and difficult example attributes in the unlabeled image set. The difficult example mining module 102 feeds back the determined difficult examples and difficult example attributes to the user I / O module 101. The user I / O module 101 obtains the difficult examples annotated and confirmed by the user, and feeds back the difficult examples annotated and confirmed by the user to the model training module 103. The model training module 103 continues to train the optimized AI model based on the difficult examples provided by the user I / O module 101 to obtain a more optimized AI model. The model training module 103 transmits the more optimized AI model to the AI model storage module 106 for storage, and transmits the more optimized AI model to the reasoning module 104 for reasoning processing. It should be noted here that when the inference module 104 infers the unlabeled images in the unlabeled image set to obtain difficult examples and then optimizes the optimized AI model, it is actually the same as the optimization process of the initial AI model using the difficult examples in the training images. At this time, the difficult examples in the unlabeled images are used as training images.
[0097] Optionally, the model training module 103 may also be configured to determine the AI model selected by the user on the GUI as the initial AI model, or to determine the AI model input by the user on the GUI as the initial AI model.
[0098] Optionally, the initial AI model may also include an AI model trained by using images in the training image set to train the AI model in the AI model storage module 106 .
[0099] The reasoning module 104 is configured to perform reasoning on the unlabeled images in the unlabeled image set based on the AI model to obtain reasoning results. The reasoning module 104 can communicate with the difficult example mining module 102, the user I / O module 101, and the AI model storage module 105. The reasoning module 104 obtains the unlabeled images in the unlabeled image set from the user I / O module 101, performs reasoning processing on the unlabeled images in the unlabeled image set, and obtains reasoning results for the unlabeled images in the unlabeled image set. The reasoning module 104 transmits the reasoning results to the difficult example mining module 102, so that the difficult example mining module 102 mines difficult examples and difficult example attributes in the unlabeled image set based on the reasoning results.
[0100] The data preprocessing module 105 is used to perform preprocessing operations on the training images in the training image set and the unlabeled images in the unlabeled image set received by the user I / O module 101. The data preprocessing module 105 can read the training image set or the unlabeled image set received by the user I / O module 101 from the data storage module 107, and then preprocess the unlabeled images in the unlabeled image set or the training images in the training image set. Preprocessing the training images in the training image set or the unlabeled images in the unlabeled image set uploaded by the user can make the training images in the training image set or the unlabeled images in the unlabeled image set consistent in size, and can also remove inappropriate data in the training images in the training image set or the unlabeled images in the unlabeled image set. The preprocessed training image set can be suitable for training the constructed AI model or training the initial AI model, and can also make the training effect better. The unlabeled images in the preprocessed unlabeled image set can be suitable for input into the AI model for inference processing. After the data preprocessing module 105 completes preprocessing of the training images in the training image set or the unlabeled images in the unlabeled image set, it stores the preprocessed training image set or unlabeled image set in the data storage module 107. Alternatively, the preprocessed training image set is sent to the model training module 103, and the preprocessed unlabeled image set is sent to the inference module 104. It should be understood that in another embodiment, the data storage module 107 may also be a part of the data preprocessing module 105, even if the data preprocessing module 105 has the function of storing images.
[0101] AI model storage module 106: used to store the initial AI model, optimized AI model, and AI sub-model structure, etc., and can also be used to store the AI model constructed based on the AI sub-model structure. The AI model storage module 106 can communicate with both the user I / O module 101 and the model training module 103. The AI model storage module 106 receives and stores the trained initial AI model and optimized AI model transmitted by the model training module 103. The AI model storage module 106 provides the constructed AI model or initial AI model to the model training module 103. The AI model storage module 106 stores the initial AI model uploaded by the user and received by the user I / O module 101. It should be understood that in another embodiment, the AI model storage module 106 can also be used as part of the model training module 103.
[0102] The data storage module 107 (e.g., it can be a data storage resource corresponding to the Object Storage Service (OBS) provided by a cloud service provider) is used to store the training image sets and unlabeled image sets uploaded by users, and is also used to store the data processed by the data preprocessing module 105.
[0103] It should be noted that the AI platform in this application can be a system that can interact with users. This system can be a software system or a hardware system, or a system that combines software and hardware. This is not limited in this application.
[0104] Due to the functions of the above-mentioned modules, the AI platform provided in the embodiment of the present application can mine difficult examples and difficult example attributes from unlabeled images, guide users to confirm and label difficult examples, and subsequently further train the difficult example mining algorithm and the initial AI model based on the difficult examples to obtain an optimized AI model and an optimized difficult example mining algorithm, so that the reasoning results of the AI model are more accurate, and the difficult examples mined by the difficult example mining algorithm are more accurate.
[0105] Figure 2 This is a schematic diagram of an application scenario of an AI platform 100 provided in an embodiment of the present application, such as Figure 2 As shown, in one embodiment, the AI platform 100 can be deployed entirely in a cloud environment. A cloud environment is an entity that uses basic resources to provide cloud services to users in a cloud computing model. The cloud environment includes a cloud data center and a cloud service platform. The cloud data center includes a large number of basic resources (including computing resources, storage resources, and network resources) owned by the cloud service provider. The computing resources included in the cloud data center can be a large number of computing devices (such as servers). The AI platform 100 can be deployed independently on a server or virtual machine in a cloud data center. The AI platform 100 can also be distributedly deployed on multiple servers in a cloud data center, or distributedly deployed on multiple virtual machines in a cloud data center, or distributedly deployed on servers and virtual machines in a cloud data center. As Figure 2As shown, the AI platform 100 is abstracted by the cloud service provider into an AI cloud service and provided to users on the cloud service platform. After the user purchases the cloud service on the cloud service platform (pre-charge is possible and then settlement is made based on the final resource usage), the cloud environment uses the AI platform 100 deployed in the cloud data center to provide the AI platform cloud service to the user. When using the AI platform cloud service, the user can use the application program interface (API) or GUI to determine the task to be completed by the AI model, upload the training image set and the unlabeled image set to the cloud environment, and the AI platform 100 in the cloud environment receives the user's task information, training image set and unlabeled image set, performs data preprocessing, AI model training, uses the trained AI model to reason on the unlabeled images in the unlabeled image set, performs hard example mining to determine hard examples and hard example attributes, and retrains the AI model based on the mined hard examples. The AI platform returns the mined hard examples and hard example attributes to the user through the API or GUI. The user further chooses whether to retrain the AI model based on the hard examples. The trained AI model can be downloaded or used online by the user to complete specific tasks.
[0106] In another embodiment of the present application, the AI platform 100 in a cloud environment is abstracted into an AI cloud service and provided to users. It can be divided into two parts, namely: basic AI cloud service and AI difficult example mining cloud service. Users can first purchase only the basic AI cloud service on the cloud service platform, and then purchase the AI difficult example mining cloud service when they need it. After purchase, the cloud service provider will provide the AI difficult example mining cloud service API, and ultimately the AI difficult example mining cloud service will be additionally billed according to the number of API calls.
[0107] The deployment of the AI platform 100 provided in this application is relatively flexible, such as Figure 3As shown, in another embodiment, the AI platform 100 provided by the present application can also be deployed in a distributed manner in different environments. The AI platform 100 provided by the present application can be logically divided into multiple parts, each of which has different functions. For example, in one embodiment, the AI platform 100 includes a user I / O module 101, a difficult example mining module 102, a model training module 103, an AI model storage module 105, and a data storage module 106. The various parts of the AI platform 100 can be deployed in any two or three environments of the terminal computing device, the edge environment, and the cloud environment. The terminal computing device includes: a terminal server, a smartphone, a laptop, a tablet computer, a personal desktop computer, a smart camera, etc. The edge environment is an environment that includes a collection of edge computing devices that are relatively close to the terminal computing device. The edge computing device includes: an edge server, an edge station with computing power, etc. The various parts of the AI platform 100 deployed in different environments or devices work together to provide users with functions such as determining and training the constructed AI model. For example, in one scenario, the user I / O module 101, data storage module 106, and data preprocessing module 107 of the AI platform 100 are deployed in a terminal computing device, while the hard example mining module 102, model training module 103, inference module 104, and AI model storage module 105 of the AI platform 100 are deployed in an edge computing device in an edge environment. A user sends a training image set and an unlabeled image set to the user I / O module 101 of the terminal computing device. The terminal computing device stores the training image set and the unlabeled image set in the data storage module 106. The data preprocessing module 102 preprocesses the training images in the training image set and the unlabeled images in the unlabeled image set, and also stores the preprocessed training images in the training image set and the unlabeled images in the unlabeled image set in the data storage module 106. The model training module 103 in the edge computing device determines an AI model to construct based on the user's task objectives, trains the constructed AI model and the training images in the training image set to obtain an initial AI model, and further optimizes the initial AI model based on the hard examples in the unlabeled images in the training image set and the training of the initial AI model. Optionally, the difficult example mining module 102 can also mine difficult examples and difficult example attributes included in the unlabeled image set based on the optimized AI model. The model training module 103 trains the optimized AI model based on the difficult examples provided by the user I / O module 101 to obtain a more optimized AI model. It should be understood that this application does not restrict which parts of the AI platform 100 are deployed in what environment. In actual application, adaptive deployment can be carried out according to the computing power of the terminal computing device, the resource usage of the edge environment and the cloud environment, or the specific application requirements.
[0108] The AI platform 100 can also be deployed separately on a computing device in any environment (such as separately deployed on an edge server in an edge environment). Figure 4 Schematic diagram of the hardware structure of a computing device 400 deployed with the AI platform 100. Figure 4 The computing device 400 shown includes a memory 401 , a processor 402 , a communication interface 403 , and a bus 404 . The memory 401 , the processor 402 , and the communication interface 403 are communicatively connected to each other via the bus 404 .
[0109] The memory 401 can be a read-only memory (ROM), a random access memory (RAM), a hard disk, a flash memory or any combination thereof. The memory 401 can store programs. When the program stored in the memory 401 is executed by the processor 402, the processor 402 and the communication interface 403 are used to execute the method of the AI platform 100 to train AI models for users, mine difficult examples and difficult example attributes, and further optimize the AI model based on difficult examples. The memory can also store image sets. For example, a portion of the storage resources in the memory 401 is divided into a data storage module 106 for storing data required by the AI platform 100, and a portion of the storage resources in the memory 401 is divided into an AI model storage module 105 for storing an AI model library.
[0110] Processor 402 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a graphics processing unit (GPU), or any combination thereof. Processor 402 may include one or more chips. Processor 402 may include an AI accelerator, such as a neural processing unit (NPU).
[0111] The communication interface 403 uses a transceiver module such as a transceiver to implement communication between the computing device 400 and other devices or a communication network. For example, data can be obtained through the communication interface 403.
[0112] Bus 404 may include a path for transmitting information between various components of computing device 400 (eg, memory 401 , processor 402 , communication interface 403 ).
[0113] With the development of AI technology, it has been widely applied in numerous fields. For example, AI technology is being applied in the fields of autonomous and assisted driving of vehicles, specifically for processes such as lane line recognition, traffic light recognition, automatic parking space identification, and sidewalk detection. These processes can be summarized as using AI models in the AI platform to perform image classification and / or object detection. For example, AI models are used to determine traffic light recognition and lane line recognition. Image classification is primarily used to determine the category to which an image belongs (i.e., input a frame of image and output the category to which the image belongs). Object detection can include two aspects: one is used to determine whether an object belonging to a specific category appears in the image, and the other is used to locate the object (i.e., determine the location of the object in the image). Before performing applications such as image classification and object detection, it is usually necessary to train an initial AI model based on labeled training images for a specific application scenario. This ensures that the trained AI model has the ability to perform image classification or object detection in that specific application scenario. Because obtaining labeled training images is a very time-consuming and labor-intensive process, the embodiments of this application will use image classification and object detection as examples to illustrate how to perform data annotation in the AI platform, which can improve data annotation efficiency.
[0114] The following combination Figure 5 The following describes the specific process of the data annotation method, taking the method executed by the AI platform as an example:
[0115] In step 501 , the AI platform determines a plurality of difficult examples and difficult example attributes of the plurality of difficult examples in a set of unlabeled images collected by a user.
[0116] The user is an entity that has registered an account on the AI platform. For example, the developer of the AI model. The unlabeled image set includes multiple unlabeled images. The difficulty attribute includes a difficulty coefficient, which is a number between 0 and 1 that reflects the difficulty of the example (for example, the difficulty of obtaining a correct result through classification or detection using the AI model). The larger the difficulty coefficient, the higher the difficulty level, and conversely, the smaller the difficulty coefficient, the lower the difficulty level.
[0117] In this embodiment, the user can place the images of the unlabeled image set in a folder and then open the image upload interface provided by the AI platform. The upload interface includes an input location for the image. The user can add the storage location of the unlabeled image set at the input location and upload multiple unlabeled images in the unlabeled image set to the AI platform. In this way, the AI platform can obtain the unlabeled image set. Figure 6As shown, the upload interface also displays an identifier (used to mark the image uploaded this time), an annotation type (used to indicate the purpose of the AI model trained using the image, such as target detection or image classification, etc.), creation time, image input location, image label set (such as people, cars, etc.), name (such as target, object, etc.), version name (version name of the AI model), whether it is team annotated (if it is "No", it means that only the user who uploaded the unannotated image will annotate it, if it is "Yes", it means that it can be annotated by team members including the user who uploaded the unannotated image), etc.
[0118] The AI platform can obtain the current AI model, which may be an AI model in a semi-trained state that has been trained with labeled images but still needs to be trained. Then, multiple images in the unlabeled image set are input into the AI model to obtain the inference results of the multiple unlabeled images. If the AI model is used for image classification here, the inference result of the image is the category to which the image belongs. For example, if the image is an image of an apple, the category to which the image belongs is apple. If the AI model is used for target detection, the inference result of the image is the position of the bounding box of the target included in the image and the category to which the target belongs, where the target can be an object included in the image, such as a car, a person, a cat, etc. If the AI model is used for both image classification and target detection, the inference result of the image is the category to which the image belongs, the position of the bounding box of the target included in the image, and the category to which the target belongs.
[0119] Based on the inference results, the AI platform determines the difficult examples in the unlabeled image set and the difficult example attributes of each difficult example. The difficult example attributes include the difficulty coefficient, which is used to describe the degree of difficulty. Specifically, the AI platform can use one or more of a temporal consistency algorithm, a data feature distribution-based algorithm, a data enhancement consistency-based algorithm, an uncertainty-based algorithm, a clustering-based algorithm, or an anomaly detection-based algorithm to determine the difficult examples in the unlabeled image set and the difficult example attributes of each difficult example. When the AI platform uses multiple algorithms to determine difficult examples and difficult example attributes, the weights of different algorithms are different, and the weights of different features are also different.
[0120] In step 502 , the AI platform displays at least one difficult example among the multiple difficult examples and corresponding difficult example attributes through a display interface.
[0121] In this embodiment, after the AI platform obtains difficult examples, it can provide multiple difficult examples to the user through a display interface. The user can control the display interface to display at least one difficult example among the multiple difficult examples and the corresponding difficult example attributes.
[0122] Optionally, for a hard example, the hard example attribute can be displayed at any position of the hard example. For example, in the case of object recognition, the hard example coefficient in the hard example attribute can be displayed in the upper right corner or lower right corner of the target box of the object.
[0123] In step 503 , the AI platform obtains the labeling result after the user confirms at least one difficult example based on the difficulty coefficients of multiple difficult examples in the display interface.
[0124] In this embodiment, the user can confirm the annotation of the difficult examples provided by the AI platform in the display interface (specifically including direct confirmation, confirmation after modification, etc.). The AI platform can obtain the annotation results of the annotation confirmation of at least one difficult example. Specifically, for difficult examples in different scenarios, the annotation results include different content. For example, in the image classification scenario, the annotation results of the difficult examples include the correct category to which the image belongs. In the target prediction scenario, the annotation results of the difficult examples include the location and category of the target in the image.
[0125] This way, because the AI platform can provide users with difficult examples, users only need to annotate and confirm the difficult examples, without having to annotate a large number of unannotated images. This can improve annotation efficiency, and in turn improve the efficiency of optimizing AI models. Furthermore, the platform provides users with the attributes of difficult examples, allowing them to better understand them and further improve annotation efficiency.
[0126] In a possible implementation, the hard example attribute also includes hard example reasons; wherein, the hard example reasons may include multiple reasons, which can be roughly divided into inconsistent data enhancement, inconsistent data feature distribution, and other reasons.
[0127] Data augmentation inconsistency is used to indicate that after certain processing is performed on an unlabeled image, the inference result is inconsistent with the unprocessed state. These processing results may include one or more of cropping, scaling, sharpening, flipping (horizontally and / or vertically), translating, adding Gaussian noise, and shearing. These are just examples, and other processing is also possible and is not limited in the embodiments of this application. For example, after a portion of an image is cropped, the inference result is different from that before the cropping. This is an image cropping augmentation inconsistency.
[0128] Inconsistent data feature distribution is used to indicate that the features of the unlabeled image are inconsistent with the features of the training image used to train the AI model. Specifically, it may include one or more inconsistencies between the resolution of the unlabeled image and the resolution, size, brightness, clarity, color saturation, and target frame features of the training image. The features of the target frame may include the number of target frames, the area of the target frame, the degree of stacking of the target frame, etc. This is just an example, and it can also be other, which is not limited in the embodiments of the present application. For example, the training image is an image taken during the day, and some images in the unlabeled image set are images taken at night. In this case, the data feature distribution (brightness feature) is inconsistent.
[0129] Other reasons may include inconsistent timing and inconsistent similar images. Inconsistent timing refers to inconsistent inference results for multiple consecutive similar images. For example, in a video, the first and third images of three adjacent images each contain three target boxes, while the second image contains only two target boxes. This indicates that the second image is a difficult example, and the difficulty is caused by inconsistent timing of consecutive images. Inconsistent similar images refer to significant differences in the inference results of unlabeled images compared to the inference results of images of the same category in the training images.
[0130] The reason for the difficulty in the difficult example attribute can be displayed on the side of the difficult example, such as the right side.
[0131] This way, by displaying the reason for a difficult example, users can clearly understand why the image is a difficult example, allowing them to take appropriate measures to reduce the number of difficult examples. For example, if a difficult example is difficult because the image brightness does not conform to the distribution, the user can add images with the same or similar brightness as the difficult example to the training image.
[0132] In one possible implementation, the AI platform displays suggestion information corresponding to the difficulty reason of at least one difficult example through a display interface.
[0133] Among them, the suggestion information is used to reflect the processing methods that can be used to reduce the causes of the corresponding difficult examples. For example, the cause of the difficulty of a certain difficult example is the inconsistency of consecutive images, and the corresponding suggestion information is "re-label, and add images similar to the difficult example." The cause of the difficulty of a certain difficult example is image cropping enhancement, which means that the AI model's generalization ability for the difficult example is not strong enough. The suggestion information is "re-label, and crop the difficult example" to increase the generalization ability of the AI model. The cause of the difficulty of a certain difficult example is that the brightness of the difficult example does not conform to the brightness distribution of the training image, which means that it is highly likely that the AI model has not learned images of the brightness of the difficult example. The suggestion information is "re-label, and add images of this brightness" to make the model more generalizable.
[0134] In this embodiment, the AI platform can also display suggestion information corresponding to the difficulty reason of at least one difficult case through the display interface, with the difficulty reason and the suggestion information displayed in correspondence. This provides users with more intuitive suggestion information, allowing users to take appropriate measures, thereby making the trained AI model more accurate.
[0135] Optionally, the reason for the difficult example and suggestion information can be displayed on the side of the difficult example, such as on the right side.
[0136] In one possible implementation, the hard case attribute also includes a hard case type. In object recognition scenarios, the hard case type can reflect the type of the target box. The hard case type can be represented by 1, 2, or 3. 1 indicates that the target box is a normal box, which means the AI platform analyzes that this box is likely to be correctly predicted. 2 indicates that the target box is a false positive, which means the AI platform analyzes that this box may have an error in position or label, and the user needs to pay special attention when making adjustments. 3 indicates that the target box is a missed positive, which means the AI platform analyzes that this box may have been missed by the model, and the user needs to pay special attention when making adjustments.
[0137] Optionally, some images in some open source data sets already have original annotations. In this embodiment, difficult examples can also be mined from images with original annotations. At this time, if the original annotations of the difficult examples are inaccurate, the difficult example type can also be represented by 0, which means that the original annotations are inaccurate.
[0138] Optionally, when displaying difficult example types, the AI platform may display different colors or lines for different difficult example types, for example, red for false positives, yellow for missed positives, and blue for normal positives.
[0139] In one possible implementation, when the AI platform displays difficult examples through a display interface, the user can filter certain difficult examples for display, or the user can control the display of difficult examples in a certain order. The process is as follows:
[0140] The AI platform obtains the filtering range and / or display order of the difficulty coefficient selected by the user on the display interface; based on the filtering range and / or display order of the difficulty coefficient, at least one difficult example and the corresponding difficult example attributes are displayed through the display interface.
[0141] In this embodiment, the difficult example display interface provided by the AI platform includes sorting options, difficult example coefficient selection options, etc. The sorting options are used by users to control the display order of difficult examples, and the difficult example coefficient selection options are used by users to filter certain difficult examples for display based on the difficult example coefficients. Users can select the display order in the option bar corresponding to the sorting options. The display order includes the order from high to low difficulty coefficients, the order from low to high difficulty coefficients, etc. The difficult example coefficient selection option is represented by a sliding bar, and users can drag the slider on the sliding bar to select the filtering range. The difficult example coefficient selection option can also be represented by a difficult example coefficient range input box. After the AI platform detects the user's selection, the AI platform can display the difficult examples and difficult example attributes corresponding to the filtering range and / or display order of the difficult example coefficient through the display interface.
[0142] In addition, when the user selects the order of difficulty coefficients from high to low, a prompt message will be displayed: "If you select the order of difficulty coefficients from high to low for labeling, it means that the data with model prediction errors will be labeled first. The amount of adjustment required by the user is relatively large, but most of this data was not recognized by the previous round of AI models. Labeling this data first can speed up the iteration speed of the AI model and quickly improve the accuracy of the AI model." When the user selects the order of difficulty coefficients from low to high, a prompt message will be displayed: "If you select the order of difficulty coefficients from low to high for labeling, most of this data was already recognized by the AI model in the previous round. The amount of adjustment required by the user is small, but the iteration speed of the AI model is slow."
[0143] In addition, the AI platform can also prompt "users to selectively label difficult examples or non-difficult examples based on their current energy (non-difficult examples refer to data that the AI model is likely to predict correctly, and the amount of changes in labeling these data is relatively small), or select difficult examples that currently need to be labeled based on the difficulty coefficient."
[0144] In one possible implementation, the AI platform may also display statistical information of multiple difficult examples, and the processing may be:
[0145] When at least one difficult example is multiple difficult examples, the AI platform displays statistical information corresponding to the unlabeled image set through a display interface, wherein the statistical information includes one or more of the distribution information of non-difficult examples and difficult examples in the unlabeled image set, the distribution information of difficult examples for each difficult example cause in the unlabeled image set, the distribution information of difficult examples for each difficult example coefficient range in the unlabeled image set, or the distribution information of difficult examples for each difficult example type in the unlabeled image set.
[0146] In this embodiment, when the AI platform displays multiple difficult examples in the unlabeled image set, the AI platform can determine the statistical information corresponding to the unlabeled image set. Specifically, the AI platform can determine the distribution information of non-difficult examples and difficult examples in the labeled image set, and can determine the distribution information of difficult examples for each difficult example reason in the unlabeled image set, and can also determine the distribution information of difficult examples for each difficult example coefficient range in the unlabeled image set. The difficult example coefficient range can be preset, such as the difficult example coefficient range including 0-0.3, 0.3-0.6, and 0.6-1. The AI platform can also determine the distribution information of difficult examples for each difficult example type in the unlabeled image set.
[0147] The AI platform can display the determined statistical information through the display interface. In this way, since the statistical information includes the distribution information of non-difficult examples and difficult examples in the unlabeled image set, the user can see the proportion of difficult examples in the unlabeled image set, and then determine whether the reasoning performance of the AI model meets the requirements. The smaller the proportion of difficult examples in the unlabeled image set, the better the reasoning performance of the AI model. Since the statistical information includes the distribution information of difficult examples for each difficult example cause in the unlabeled image set, when there are more images of certain difficult example causes, the user can focus on adding images of these difficult example causes to quickly optimize the AI model. Since the statistical information includes the distribution information of difficult examples for each difficult example coefficient range in the unlabeled image set, the user can understand the distribution of difficult examples. Since the statistical information includes the distribution information of difficult examples for each difficult example type in the unlabeled image set, the user can understand the distribution of difficult examples in the difficult example type in the unlabeled image set.
[0148] In a possible implementation, in step 502, the processing is as follows: the AI platform determines, based on a difficult example mining algorithm, a plurality of difficult examples and difficult example attributes of the plurality of difficult examples in the unlabeled image set collected by the user.
[0149] In this embodiment, the hard example mining algorithm includes one or more of a temporal consistency algorithm, a data feature distribution-based algorithm, a data augmentation consistency-based algorithm, an uncertainty-based algorithm, a clustering-based algorithm, or an anomaly detection-based algorithm. The AI platform can input the inference results of the unlabeled image into the hard example mining algorithm to obtain multiple hard examples and their hard example attributes from the unlabeled image set.
[0150] In one possible implementation, after obtaining the user's confirmation of the annotation results of the difficult example, the AI platform can update the difficult example mining algorithm to:
[0151] In the annotation results provided by the user, if the target difficult example is confirmed as a correct difficult example in at least one difficult example, the AI platform will increase the weight corresponding to the difficulty reason of the target difficult example; in the annotation results, if the target difficult example is confirmed as an incorrect difficult example, the AI platform will reduce the weight corresponding to the difficulty reason of the target difficult example; based on the updated weight and annotation results, the difficult example mining algorithm will be updated.
[0152] Among them, correct hard examples refer to those that users believe are hard examples among those mined by the hard example mining algorithm, and incorrect hard examples refer to those that users believe are not hard examples among those mined by the hard example mining algorithm.
[0153] In this embodiment, after obtaining the annotation results provided by the user, the AI platform can obtain the weight corresponding to the difficulty cause of at least one difficult example in the difficult example mining algorithm. If the target difficult example in at least one difficult example is confirmed to be a correct difficult example, the AI platform can increase the weight corresponding to the difficulty cause of the target difficult example, which is the weight of the difficulty cause in the difficult example mining algorithm. If the target difficult example in at least one difficult example is confirmed to be an incorrect difficult example, the AI platform can reduce the weight corresponding to the difficulty cause of the target difficult example, which is the weight of the difficulty cause in the difficult example mining algorithm.
[0154] The AI platform then updates the hard example mining algorithm using the user-provided annotations and the updated weights for each hard example cause. This adjustment of the weights for the hard example causes allows the updated hard example mining algorithm to mine more accurate hard examples, reducing the amount of annotations required.
[0155] In addition, in this embodiment, the difficult examples that have been annotated and confirmed by the user can also be used to update the AI model and optimize the AI model to make the reasoning results provided by the AI model more accurate.
[0156] In addition, in this embodiment, after the user has marked and confirmed the difficult examples, the marked and confirmed difficult examples will be synchronized to the marked image set. In addition, the AI platform can also convert the user's difficult example set to be confirmed into a marked difficult example set, or a marked non-difficult example set, or an unmarked difficult example set or an unmarked non-difficult example set based on the user's marking confirmation. For example, if an image has marking information, and one of the target boxes is confirmed by the user as a missed detection box or a false detection box, the AI platform will store the image in the marked difficult example set.
[0157] In addition, in this embodiment, the AI platform also provides a process for exporting difficult examples and difficult example attributes. Users can select the save path of difficult examples and difficult example attributes in the export difficult example interface, and export difficult examples and difficult example attributes. Figure 7 As shown, the interface for exporting difficult examples includes the save path, export range (which can include exporting the currently selected sample (i.e., the selected difficult example), and exporting all samples under the current screening conditions). In addition, the interface for exporting difficult examples also includes the option to turn on the difficult example attribute. If the option to turn on the difficult example attribute is selected, the difficult example attribute will also be exported when exporting the difficult example. If the option to turn on the difficult example attribute is not selected, the difficult example attribute will not be exported when exporting the difficult example. At the same time, corresponding to turning on the difficult example attribute, there is also a suggestion for turning on the difficult example attribute: it is recommended to turn on the difficult example attribute, with a difficult example coefficient of 0 to 1. The smaller the difficult example coefficient, the higher the accuracy, etc. In addition, the difficult example type is also given corresponding to the starting difficult example attribute.
[0158] In addition, in a possible implementation, in step 501, the current AI model may be an initial AI model, and the AI platform may provide the user with a labeling selection interface, which includes at least one labeling method that the user can select. The AI platform may receive the labeling method selected by the user and use the initial AI model corresponding to the labeling method to determine the inference result of the unlabeled image set. Specifically, Figure 8 As shown, the annotation selection interface provides multiple annotation methods, including active learning and pre-annotation. In the active learning mode, the AI platform first trains the constructed AI model using multiple annotated images provided by the user to obtain an initial AI model. Then, the unannotated image set is annotated based on the initial AI model to obtain the inference results of the unannotated image set. In the pre-annotation mode, the AI platform directly obtains the existing initial AI model, annotates the unannotated image set based on the initial AI model, and obtains the inference results of the unannotated image set. In addition, the annotation selection interface also displays the number of all images, the number of unannotated images, the number of annotated images, and the number of pending confirmation (pending confirmation refers to the number of difficult examples waiting for user confirmation). Among them, the multiple annotated images provided by the user can be annotated by the user through the AI platform, and the multiple images are selected by the AI platform from the unannotated images uploaded by the user using a data selection algorithm. Specifically, when the user selects the unannotated images that they have annotated, the user can determine whether the AI model they want to train is used in the image classification scenario, the target detection scenario, or the scenario of combining user image classification and target detection. If the AI model to be trained is used in an image classification scenario, the AI platform's image annotation interface for image classification scenarios will be launched. If the AI model to be trained is used in an object detection scenario, the AI platform's image annotation interface for object detection scenarios will be launched. This interface provides options such as selecting an image, selecting a bounding box, using the return key, zooming in, and zooming out. Users can select an image to open a frame. Within this image, objects are annotated using a bounding box and annotations are added. The annotations include the object's category within the bounding box and the box's location within the image (since bounding boxes are generally rectangular, the location can be indicated using the coordinates of the top left and bottom right corners). After the user annotates the object with a bounding box, the AI platform will obtain the object's bounding box. The image annotation interface will also display the annotation information bar for the image. The annotation information bar displays information about the user-annotated object, including the label, bounding box, and operation. The label indicates the object's category, the bounding box indicates the shape of the box used, and operations include delete and modify options. Users can modify the annotations added to the image through operations.
[0159] In order to better understand the embodiments of the present application, in the embodiments of the present application, an AI platform is provided to display multiple difficult examples through a display interface:
[0160] like Figure 9 As shown (0.8, 0.9, etc. in the figure represent the difficulty coefficient), the display interface displays the filtering conditions and the content corresponding to the filtering conditions. Specific filtering conditions may include data attributes, difficult example sets, sorting, and difficulty coefficients. When the filtering condition is a difficult example, the difficult example is displayed below the filtering condition, and the difficulty coefficient of each difficult example is displayed. The definition of the difficult example and the statistical information of the unlabeled image set are displayed on the right side of the filtering condition. The statistical information includes the distribution information of difficult examples and non-difficult examples in the unlabeled image set and the distribution information of difficult examples for each difficult example cause, and the difficult example causes and corresponding suggestion information. The difficult example causes and suggestion information can be represented in a table, but are not limited to this method. The distribution information of difficult examples for each difficult example cause is represented by a bar chart, but is not limited to this method. The distribution information of difficult examples and non-difficult examples in the unlabeled image set is represented by a pie chart, but is not limited to this method.
[0161] In addition, in order to better understand the embodiments of the present application, in the embodiments of the present application, an AI platform is provided to display a difficult example schematic diagram through a display interface:
[0162] like Figure 10 As shown in the figure, the display interface shows the difficult example that the user currently wants to mark and confirm. The corresponding difficult example displays the difficulty type, difficulty coefficient, difficulty reason, and recommended information. The display interface also displays the marking confirmation operation for each target box. For example, the marking box of the "person" is a square box, and the operation options are delete and modify.
[0163] And in Figure 10 In order to facilitate users to select difficult examples for annotation confirmation, multiple difficult examples are displayed below the currently marked and confirmed difficult example. Users can directly click on the difficult examples displayed below to switch difficult examples.
[0164] In addition, in this application, in order to better understand this embodiment, the following is also provided: Figure 11 The flow chart shown:
[0165] First, an AI model is obtained. This model can be a pre-trained model or trained based on user-annotated unlabeled data and annotated images. The AI model then infers the unlabeled image set and obtains inference results. The AI platform then uses a hard example mining algorithm to identify hard examples and their attributes within the unlabeled image set. The AI platform then provides these hard examples and their attributes to the user, who then confirms their annotations. Based on the user's confirmed annotations, the hard example mining algorithm and AI model are updated.
[0166] In the embodiments of the present application, the attributes of difficult examples are used to enable users to understand difficult example data and provide targeted suggestions. This reduces the labeling cost while also improving the reasoning performance of the AI model. Moreover, the present application forms a human-computer interaction system with the behavior of the AI platform and the user. The AI platform determines the reasoning results of the unlabeled image set, and the difficult example mining algorithm in the AI platform determines the difficult examples and difficult example attributes in the unlabeled image set and provides them to the user. The AI platform guides users to label and confirm difficult examples, and the AI platform uses the labeling results fed back by users to update the difficult example mining algorithm and AI model.
[0167] It should be noted that the above method for providing data annotation can be implemented by one or more modules on the AI platform 100. Specifically, the difficult example mining module is used to implement Figure 5 Step 502 in the user I / O module is used to implement Figure 5 The model training module is used to implement the update process of the hard example mining algorithm.
[0168] Another embodiment of the present application is as follows Figure 12 As shown, a method for updating the difficult example mining algorithm is provided, and the process of the method is as follows:
[0169] In step 1201 , the AI platform uses a hard example mining algorithm to determine multiple hard examples in the unlabeled image set.
[0170] In step 1202 , the AI platform displays at least one difficult example from among multiple difficult examples to the user through a display interface.
[0171] Step 1201 and step 1202 can be referred to the description in the previous text and will not be repeated here.
[0172] In step 1203 , the AI platform updates the difficult example mining algorithm based on the labeling result of at least one difficult example confirmed by the user through the display interface.
[0173] In this embodiment, after obtaining the annotation results of at least one difficult example confirmed by the user through the display interface, the AI platform can update the difficult example mining algorithm based on the annotation results, so that the accuracy of difficult example prediction becomes higher and higher.
[0174] Specifically, step 1201 can be performed by the difficult example mining module mentioned above, step 1202 can be performed by the user I / O module mentioned above, and step 1203 can be performed by the model training module.
[0175] In a possible implementation, the processing of step 1203 may be:
[0176] The AI platform obtains a weight for the difficulty cause of each difficult example in at least one difficult example; analyzes the annotation results, and if the target difficult example in at least one difficult example is confirmed as a correct difficult example by the user, increases the weight corresponding to the difficulty cause of the target difficult example; in the annotation results, if the target difficult example is confirmed as an incorrect difficult example by the user, reduces the weight corresponding to the difficulty cause of the target difficult example; and updates the difficult example mining algorithm based on the updated weight.
[0177] This processing process can be found in the description above and will not be repeated here.
[0178] In one possible implementation, the AI platform uses a difficult example mining algorithm to determine the causes of difficulty for multiple difficult examples, and provides users with suggestion information corresponding to the causes of difficulty through a display interface, wherein the suggestion information indicates possible processing methods for reducing the difficulty of examples with corresponding causes of difficulty.
[0179] In one possible implementation, the AI platform displays the difficult case attributes corresponding to at least one difficult case through a display interface, where the difficult case attributes include one or more of a difficult case coefficient, a difficult case reason, and a difficult case type.
[0180] How to display difficult examples and their attributes has been described in the previous article. Please refer to the description in the previous article and will not be repeated here.
[0181] Figure 13 This is a structural diagram of the data annotation device provided in the embodiment of the present application. The device can be implemented as part or all of the aforementioned AI platform through software, hardware, or a combination of both. The data annotation device provided in the embodiment of the present application can implement the embodiment of the present application. Figure 5 The process described above includes: a difficult example mining module 1310 and a user I / O module 1320, wherein:
[0182] A hard example mining module 1310 is configured to determine a plurality of hard examples in the unlabeled image set and hard example attributes of the plurality of hard examples, wherein the hard example attributes include hard example coefficients, and specifically can be used to implement the hard example mining function in step 501 and execute the implicit steps included in step 501;
[0183] User I / O module 1320 is used to:
[0184] Displaying at least one difficult example from the plurality of difficult examples and corresponding difficult example attributes to a user through a display interface;
[0185] Obtaining the annotation result after the user confirms the at least one difficult example in the display interface based on the difficulty coefficients of the multiple difficult examples can be specifically used to implement the user I / O function in step 502 and step 503 and execute the implicit steps included in step 502 and step 503.
[0186] In a possible implementation, the difficult example attribute further includes a difficult example reason;
[0187] The user I / O module 1320 is further configured to:
[0188] Displaying, through the display interface, a difficulty reason corresponding to the at least one difficult case and suggestion information corresponding to the difficulty reason;
[0189] The suggestion information indicates possible processing methods for reducing the difficulty of the corresponding difficult examples.
[0190] In a possible implementation, the hard case attribute further includes a hard case type, and the hard case type includes false positive, missed positive, or predicted normal.
[0191] In a possible implementation, the user I / O module 1320 is further configured to:
[0192] Different difficult example types are displayed using different display colors or different display lines through the display interface.
[0193] In one possible implementation, the hard example mining module 1310 is configured to:
[0194] According to a hard example mining algorithm, a plurality of hard examples in the unlabeled image set and hard example attributes of the plurality of hard examples are determined.
[0195] In one possible implementation, Figure 14 As shown, the apparatus further includes a model training module 1330, which is used to:
[0196] After obtaining a marking result of the at least one difficult example confirmed by the user in the display interface according to the difficulty coefficients of the multiple difficult examples, obtaining a weight for a difficulty reason of each difficult example in the at least one difficult example;
[0197] Analyzing the annotation result, if a target difficult example in the at least one difficult example is confirmed by the user as a correct difficult example, increasing a weight corresponding to a difficulty reason of the target difficult example;
[0198] If the target difficult example is confirmed by the user as an incorrect difficult example, reducing the weight corresponding to the difficult example reason of the target difficult example;
[0199] The difficult example mining algorithm is updated according to the updated weights.
[0200] In a possible implementation, the user I / O module 1320 is configured to:
[0201] Obtaining a screening range and / or display order of difficult example coefficients selected by a user on the display interface;
[0202] According to the screening range and / or display order of the difficulty coefficient, the at least one difficulty example and the corresponding difficulty example attribute are displayed through the display interface.
[0203] In a possible implementation, the user I / O module 1320 is further configured to:
[0204] When the at least one difficult example is one of the multiple difficult examples, statistical information corresponding to the unlabeled image set is displayed through the display interface, wherein the statistical information includes one or more of distribution information of non-difficult examples and difficult examples in the unlabeled image set, distribution information of difficult examples for each cause of difficult examples in the unlabeled image set, distribution information of difficult examples for each range of difficulty coefficients in the unlabeled image set, or distribution information of difficult examples for each type of difficult examples in the unlabeled image set.
[0205] Optional, Figure 13 and Figure 14 The device shown can also be used to perform other actions described in the aforementioned data labeling method embodiment, which will not be repeated here.
[0206] It should be noted that the above Figure 13 and Figure 14 The device shown can be Figure 1 Part of the AI platform shown.
[0207] Figure 15 This is a structural diagram of the data annotation device provided in the embodiment of the present application. The device can be implemented as part or all of the AI platform through software, hardware, or a combination of both. The data annotation device provided in the embodiment of the present application can realize the embodiment of the present application. Figure 12 The process described above includes: a difficult example mining module 1510, a user I / O module 1520, and a model training module 1530, wherein:
[0208] A hard example mining module 1510 is configured to determine multiple hard examples in the unlabeled image set using a hard example mining algorithm, and specifically to implement the hard example mining function in step 1201 and to execute the implicit steps included in step 1201;
[0209] A user I / O module 1520 is configured to display at least one difficult example from the plurality of difficult examples to a user via a display interface, and specifically to implement the user I / O function in step 1202 and to execute the implicit steps included in step 1202;
[0210] The model training module 1530 is used to update the difficult example mining algorithm based on the labeling results after the user confirms the at least one difficult example through the display interface. It can be specifically used to implement the model training function in step 1203 and execute the implicit steps included in step 1203.
[0211] In a possible implementation, the difficult example mining module 1510 is further configured to:
[0212] Determining the reasons why the multiple difficult examples are difficult using the difficult example mining algorithm;
[0213] The user I / O module 1520 is further configured to provide the user with suggestion information corresponding to the cause of the difficult case through the display interface, wherein the suggestion information indicates a possible processing method for reducing the difficulty of the case caused by the corresponding difficult case.
[0214] In one possible implementation, the model training module 1530 is configured to:
[0215] Obtaining a weight for a difficulty reason of each difficult example in the at least one difficult example;
[0216] Analyzing the annotation result, if a target difficult example in the at least one difficult example is confirmed by the user as a correct difficult example, increasing a weight corresponding to a difficulty reason of the target difficult example;
[0217] If the target difficult example is confirmed by the user as an incorrect difficult example, reducing the weight corresponding to the difficult example reason of the target difficult example;
[0218] The difficult example mining algorithm is updated according to the updated weights.
[0219] In a possible implementation, the user I / O module 1520 is configured to:
[0220] The difficult example attributes corresponding to the at least one difficult example are displayed through the display interface, and the difficult example attributes include one or more of a difficult example coefficient, a difficult example reason, and a difficult example type.
[0221] Optional, Figure 15 The device shown can also be used to perform other actions described in the aforementioned method embodiment for updating the difficult example mining algorithm, which will not be described in detail here.
[0222] It should be noted that the above Figure 15 The device shown can be Figure 1 Part of the AI platform shown.
[0223] This application also provides a Figure 4In the computing device 400 shown, the processor 402 in the computing device 400 reads the program and image set stored in the memory 401 to execute the method executed by the aforementioned AI platform.
[0224] Since the various modules in the AI platform 100 provided by this application can be distributedly deployed on multiple computers in the same environment or in different environments, this application also provides a Figure 16 The computing device shown includes multiple computers 1600, each of which includes a memory 1601, a processor 1602, a communication interface 1603, and a bus 1604. The memory 1601, the processor 1602, and the communication interface 1603 are communicatively connected to each other via the bus 1604.
[0225] Memory 1601 can be a read-only memory, a static storage device, a dynamic storage device, or a random access memory. Memory 1601 can store programs. When the program stored in memory 1601 is executed by processor 502, processor 1602 and communication interface 1603 are used to execute part of the method for the AI platform to obtain an AI model. The memory can also store image sets. For example, a portion of the storage resources in memory 1601 is divided into an image set storage module for storing the image set required by the AI platform, and a portion of the storage resources in memory 1601 is divided into an AI model storage module for storing the AI model library.
[0226] The processor 1602 may be a general-purpose central processing unit, a microprocessor, an application-specific integrated circuit, a graphics processor, or one or more integrated circuits.
[0227] The communication interface 1603 uses a transceiver module such as, but not limited to, a transceiver to implement communication between the computer 1600 and other devices or a communication network. For example, an image set can be obtained through the communication interface 1603.
[0228] The bus 504 may include a path for transmitting information between various components of the computer 1600 (eg, the memory 1601 , the processor 1602 , and the communication interface 1603 ).
[0229] A communication path is established between each of the aforementioned computers 1600 via a communication network. Each computer 1600 runs any one or more of the user I / O module 101, the hard example mining module 102, the model training module 103, the inference module 104, the AI model storage module 105, the data storage module 106, and the data preprocessing module 107. Any computer 1600 can be a computer in a cloud data center (e.g., a server), a computer in an edge data center, or a terminal computing device.
[0230] The descriptions of the processes corresponding to the above figures have different focuses. For parts that are not described in detail in a certain process, please refer to the relevant descriptions of other processes.
[0231] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, or a combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product providing the AI platform includes one or more computer instructions for the AI platform, which, when loaded and executed on a computer, generate all or part of the embodiments of the present application. Figure 5 The process or function described.
[0232] The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, twisted pair, or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium stores computer program instructions for providing an AI platform. The computer-readable storage medium may be any medium that a computer can access or a data storage device such as a server or data center that includes one or more media integrations. The medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., an optical disk), or a semiconductor medium (e.g., an SSD).
Claims
1. A data annotation method, characterized in that: The method comprises: Determining a plurality of difficult examples in the unlabeled image set and difficulty attributes of the plurality of difficult examples, wherein the difficulty attributes include a difficulty coefficient and a difficulty reason, and the difficulty reason includes one or more of inconsistent data augmentation, inconsistent data feature distribution, or inconsistent inference results of similar images; Displaying at least one difficult example from the plurality of difficult examples and corresponding difficult example attributes to a user through a display interface; Obtain a marking result after the user confirms the at least one difficult example according to the difficult example attributes of the multiple difficult examples in the display interface.
2. The method according to claim 1, characterized in that The method further comprises: Displaying suggestion information corresponding to the difficulty reason of the at least one difficult case through the display interface; The suggestion information indicates possible processing methods for reducing the difficulty of the corresponding difficult examples.
3. The method according to claim 1 or 2, characterized in that The hard case attribute further includes a hard case type, and the hard case type includes false detection, missed detection, or predicted normal.
4. The method according to claim 3, characterized in that The displaying of at least one difficult example among the multiple difficult examples and the corresponding difficult example attributes through a display interface further includes: Different difficult example types are displayed using different display colors or different display lines through the display interface.
5. The method according to claim 1 or 2, characterized in that The determining of a plurality of difficult examples in the unlabeled image set and the difficult example attributes of the plurality of difficult examples includes: According to a hard example mining algorithm, a plurality of hard examples in the unlabeled image set and hard example attributes of the plurality of hard examples are determined.
6. The method according to claim 5, characterized in that After obtaining the labeling result of the user confirming the at least one difficult example according to the difficulty coefficients of the multiple difficult examples in the display interface, the method further includes: Obtaining a weight for a difficulty reason of each difficult example in the at least one difficult example; Analyzing the annotation result, if a target difficult example in the at least one difficult example is confirmed by the user as a correct difficult example, increasing a weight corresponding to a difficulty reason of the target difficult example; If the target difficult example is confirmed by the user as an incorrect difficult example, reducing the weight corresponding to the difficult example reason of the target difficult example; The difficult example mining algorithm is updated according to the updated weights.
7. The method according to claim 1 or 2, characterized in that The displaying at least one difficult example from the plurality of difficult examples and the corresponding difficult example attributes to the user through a display interface includes: Obtaining a screening range and / or display order of difficult example coefficients selected by a user on the display interface; According to the screening range and / or display order of the difficulty coefficient, the at least one difficulty example and the corresponding difficulty example attribute are displayed through the display interface.
8. The method according to claim 1 or 2, characterized in that The method further comprises: When the at least one difficult example is one of the multiple difficult examples, statistical information corresponding to the unlabeled image set is displayed through the display interface, wherein the statistical information includes one or more of distribution information of non-difficult examples and difficult examples in the unlabeled image set, distribution information of difficult examples for each cause of difficult examples in the unlabeled image set, distribution information of difficult examples for each range of difficulty coefficients in the unlabeled image set, or distribution information of difficult examples for each type of difficult examples in the unlabeled image set.
9. A method for updating a difficult example mining algorithm, characterized in that: The method comprises: Determine, using a hard example mining algorithm, a plurality of hard examples in the unlabeled image set and reasons for the hard examples, wherein the reasons for the hard examples include one or more of inconsistent data augmentation, inconsistent data feature distribution, or inconsistent inference results of similar images; displaying, to a user through a display interface, at least one difficult example among the plurality of difficult examples and a reason for the difficulty of the at least one difficult example; The difficult example mining algorithm is updated according to the labeling result after the user confirms the at least one difficult example through the display interface.
10. The method according to claim 9, characterized in that The method further comprises: Suggestion information corresponding to the difficulty cause of the at least one difficult example is provided to the user through the display interface, wherein the suggestion information indicates a processing method that can be performed to reduce the difficulty of the example with the corresponding difficulty cause.
11. The method according to claim 10, characterized in that The updating of the difficult example mining algorithm according to the labeling result of the at least one difficult example confirmed by the user through the display interface includes: Obtaining a weight for a difficulty reason of each difficult example in the at least one difficult example; Analyzing the annotation result, if a target difficult example in the at least one difficult example is confirmed by the user as a correct difficult example, increasing a weight corresponding to a difficulty reason of the target difficult example; In the annotation result, if the target difficult example is confirmed by the user as an incorrect difficult example, reducing the weight corresponding to the difficult example reason of the target difficult example; The difficult example mining algorithm is updated according to the updated weights.
12. A data annotation device, characterized in that: The device comprises: a hard example mining module, configured to determine a plurality of hard examples in the unlabeled image set and hard example attributes of the plurality of hard examples, wherein the hard example attributes include a hard example coefficient and a hard example reason, and the hard example reason includes one or more of inconsistent data augmentation, inconsistent data feature distribution, or inconsistent inference results of similar images; User input / output I / O modules for: Displaying at least one difficult example from the plurality of difficult examples and corresponding difficult example attributes to a user through a display interface; Obtain a marking result after the user confirms the at least one difficult example according to the difficult example attributes of the multiple difficult examples in the display interface.
13. The device according to claim 12, characterized in that The user I / O module is further used for: Displaying suggestion information corresponding to the difficulty reason of the at least one difficult case through the display interface; The suggestion information indicates possible processing methods for reducing the difficulty of the corresponding difficult examples.
14. The device according to claim 12 or 13, characterized in that The hard case attribute further includes a hard case type, and the hard case type includes false detection, missed detection, or predicted normal.
15. The device according to claim 14, characterized in that The user I / O module is further used for: Different difficult example types are displayed using different display colors or different display lines through the display interface.
16. The device according to claim 12 or 13, characterized in that The difficult example mining module is specifically used to: According to a hard example mining algorithm, a plurality of hard examples in the unlabeled image set and hard example attributes of the plurality of hard examples are determined.
17. The device according to claim 16, characterized in that The device also includes a model training module, which is used to: After obtaining a marking result of the at least one difficult example confirmed by the user in the display interface according to the difficulty coefficients of the multiple difficult examples, obtaining a weight for a difficulty reason of each difficult example in the at least one difficult example; Analyzing the annotation result, if a target difficult example in the at least one difficult example is confirmed by the user as a correct difficult example, increasing a weight corresponding to a difficulty reason of the target difficult example; If the target difficult example is confirmed by the user as an incorrect difficult example, reducing the weight corresponding to the difficult example reason of the target difficult example; The difficult example mining algorithm is updated according to the updated weights.
18. The device according to claim 12 or 13, characterized in that The user I / O module is used to: Obtaining a screening range and / or display order of difficult example coefficients selected by a user on the display interface; According to the screening range and / or display order of the difficulty coefficient, the at least one difficulty example and the corresponding difficulty example attribute are displayed through the display interface.
19. The device according to claim 12 or 13, characterized in that The user I / O module is further used for: When the at least one difficult example is one of the multiple difficult examples, statistical information corresponding to the unlabeled image set is displayed through the display interface, wherein the statistical information includes one or more of distribution information of non-difficult examples and difficult examples in the unlabeled image set, distribution information of difficult examples for each cause of difficult examples in the unlabeled image set, distribution information of difficult examples for each range of difficulty coefficients in the unlabeled image set, or distribution information of difficult examples for each type of difficult examples in the unlabeled image set.
20. A device for updating a difficult example mining algorithm, characterized in that: The device comprises: a hard example mining module, configured to determine, using a hard example mining algorithm, multiple hard examples in the unlabeled image set and reasons for the hard examples, wherein the reasons for the hard examples include one or more of inconsistent data augmentation, inconsistent data feature distribution, or inconsistent inference results of similar images; A user input / output (I / O) module is configured to display at least one difficult example among the multiple difficult examples and a reason for the difficulty of the at least one difficult example to a user through a display interface; A model training module is used to update the difficult example mining algorithm according to the labeling result after the user confirms the at least one difficult example through the display interface.
21. The device according to claim 20, characterized in that The user I / O module is further used for: Suggestion information corresponding to the difficulty cause of the at least one difficult example is provided to the user through the display interface, wherein the suggestion information indicates a processing method that can be performed to reduce the difficulty of the example with the corresponding difficulty cause.
22. The device according to claim 21, characterized in that The model training module is used to: Obtaining a weight for a difficulty reason of each difficult example in the at least one difficult example; Analyzing the annotation result, if a target difficult example in the at least one difficult example is confirmed by the user as a correct difficult example, increasing a weight corresponding to a difficulty reason of the target difficult example; If the target difficult example is confirmed by the user as an incorrect difficult example, reducing the weight corresponding to the difficult example reason of the target difficult example; The difficult example mining algorithm is updated according to the updated weights.
23. A computing device, characterized in that The computing device includes a memory and a processor, the memory being configured to store a set of computer instructions; The processor executes a set of computer instructions stored in the memory to perform the method according to any one of claims 1 to 11.
24. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program code. When the computer program code is executed by a computing device, the computing device performs the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Vehicle classifier training method
CN106372658A
Difficult sample mining and model training method, device and electronic equipment
CN110610197A
Method for providing AI model, AI platform, computing device and storage medium
CN112529026A