Image labeling method and apparatus

By filtering and processing the prediction results of deep learning models, constructing a multi-scale image pyramid and calculating the loss function, the problem of high cost and low accuracy of image annotation is solved, and an efficient image annotation method is realized.

CN114529756BActive Publication Date: 2026-04-21传申弘安智能(深圳)有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
传申弘安智能(深圳)有限公司
Filing Date
2022-01-24
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Current image annotation technologies are costly and inaccurate, especially in vertical fields that require specialized knowledge, where manual annotation is uncertain and costly.

Method used

The predicted labels of sample images are filtered by receiving the prediction results of the target deep learning model, a multi-scale image pyramid is constructed and different data processing is performed, the loss function is calculated to iteratively update the model, and the model is trained using its own labeled data.

Benefits of technology

While reducing annotation costs, the accuracy of image annotation is improved. Through multi-scale processing and iterative updates of the loss function, the detection accuracy of the model is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114529756B_ABST
    Figure CN114529756B_ABST
Patent Text Reader

Abstract

This invention discloses an image annotation method and apparatus. The method includes: receiving prediction results from a target deep learning model for unlabeled sample images; selecting predicted labels for several sample images based on the prediction results as training annotation images; acquiring a multi-scale image pyramid of the training annotation images and copying it into two copies; performing a first data processing on one copy of the multi-scale image pyramid to obtain a first image, and performing a second data processing (different from the first data processing) or no processing on the other copy of the multi-scale image pyramid to obtain a second image; inputting the first image and the second image into the target deep learning model to obtain corresponding first and second predicted labels, calculating corresponding loss functions, and iteratively updating the target deep learning model. This invention can fully utilize the annotation data output by the original target deep learning model, reducing annotation costs while improving the accuracy of the annotation data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of machine learning technology, and in particular to an image annotation method and apparatus. Background Technology

[0002] The first step in using deep learning models to solve real-world problems is obtaining labeled data for the relevant application scenario. Generally, training a well-performing model requires tens of thousands of labeled data points, resulting in a massive amount of data. Furthermore, when the labeling task involves specialized knowledge in a vertical domain, on-the-job training for relevant personnel is also required, leading to a sharp increase in both human and time costs.

[0003] The accuracy of annotation is also crucial. Manual annotation is inherently unpredictable and prone to chance, requiring different quality control methods for different scenarios and the training of more specialized quality control personnel, resulting in high overall costs. Therefore, there is an urgent need for automated image annotation methods that can reduce annotation costs while obtaining high-precision annotation data. Summary of the Invention

[0004] This invention aims to at least solve one of the technical problems existing in the prior art. To this end, this invention proposes an image annotation method that can improve the accuracy of annotated data while reducing annotation costs.

[0005] The present invention also proposes an image annotation apparatus having the above-described image annotation method.

[0006] The present invention also proposes a computer-readable storage medium having the above-described image annotation method.

[0007] An image annotation method according to a first aspect of the present invention includes the following steps: receiving prediction results of a target deep learning model for unlabeled sample images; selecting prediction labels for a number of sample images based on the prediction results as training annotation images; obtaining a multi-scale image pyramid of the training annotation images and copying it into two copies; performing a first data processing on one copy of the multi-scale image pyramid to obtain a first image, and performing a second data processing different from the first data processing or not performing any processing on the other copy of the multi-scale image pyramid to obtain a second image; inputting the first image and the second image into the target deep learning model to obtain corresponding first prediction labels and second prediction labels; calculating a corresponding loss function based on the first prediction labels and the second prediction labels; and iteratively updating the target deep learning model.

[0008] The image annotation method according to the embodiments of the present invention has at least the following beneficial effects: it can make full use of the annotation data output by the original target deep learning model, select a number of training annotation images, perform two different processing on each training annotation image, input them as a set of samples into the target deep learning model, calculate the loss function based on the two predicted labels, and iterate the target deep learning model, thereby reducing the annotation cost and effectively improving the accuracy of the annotation data.

[0009] According to some embodiments of the present invention, a method for selecting predicted labels for a number of sample images based on the prediction results as training labeled images includes: receiving the prediction results of unlabeled sample images; selecting prediction boxes with a confidence level above a preset threshold from the prediction results as labels for the sample images; and using the labeled sample images as the training labeled images.

[0010] According to some embodiments of the present invention, the first data processing is strong data augmentation, and the second data processing is weak data enhancement.

[0011] According to some embodiments of the present invention, the step of calculating the corresponding loss function based on the first predicted label and the second predicted label includes: taking the second predicted label as the true label of the first multi-scale image pyramid after performing the first data processing, comparing it with the first predicted label, and calculating the corresponding loss function.

[0012] According to some embodiments of the present invention, the iterative update method for the target deep learning model includes: inputting the first image and the second image into the second multi-scale thinning branch of the target deep learning model, calculating the loss function of the second multi-scale thinning branch, merging it into the corresponding loss function of the main branch of the target deep learning model, and iteratively updating the target deep learning model.

[0013] According to some embodiments of the present invention, the second multi-scale refinement branch of the target deep learning model shares weights with the feature extraction network of the main branch of the target deep learning model.

[0014] According to some embodiments of the present invention, the iterative update method for the target deep learning model includes: inputting the first image and the second image into the target deep learning model, calculating the corresponding loss function, and iteratively updating the target deep learning model.

[0015] An image annotation apparatus according to a second aspect of the present invention includes: a selection module, configured to receive prediction results of a target deep learning model for unlabeled sample images, and select prediction labels for a plurality of sample images based on the prediction results as training annotation images; an image processing module, configured to acquire a multi-scale image pyramid of the training annotation images, and copy it into two copies; perform a first data processing on one copy of the multi-scale image pyramid to obtain a first image, and perform a second data processing different from the first data processing or perform no processing on the other copy of the multi-scale image pyramid to obtain a second image; and a training module, configured to input the first image and the second image into the target deep learning model to obtain corresponding first prediction labels and second prediction labels, calculate a corresponding loss function based on the first prediction labels and the second prediction labels, and iteratively update the target deep learning model.

[0016] The image annotation apparatus according to embodiments of the present invention has at least the following beneficial effects: it can make full use of the annotation data output by the original target deep learning model, select a number of training annotation images, perform two different processing on each training annotation image, input them as a set of samples into the target deep learning model, calculate the loss function based on the two predicted labels obtained, and iterate the target deep learning model, thereby reducing the annotation cost and effectively improving the accuracy of the annotation data.

[0017] A computer-readable storage medium according to a third aspect of the present invention stores a computer program thereon, which, when executed by a processor, implements the method according to a first aspect of the present invention.

[0018] The computer-readable storage medium according to embodiments of the present invention has at least the same beneficial effects as the method of the first aspect of the present invention.

[0019] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0020] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0021] Figure 1 This is a schematic diagram of the main flow of the method according to an embodiment of the present invention;

[0022] Figure 2 This is a schematic diagram of data interaction in the method of an embodiment of the present invention;

[0023] Figure 3 Two examples of target deep learning models;

[0024] Figure 4 A schematic block diagram illustrating the training process of a target deep learning model example using the method of an embodiment of the present invention;

[0025] Figure 5 This is a schematic block diagram of the system modules according to an embodiment of the present invention.

[0026] Figure label:

[0027] Select module 100, image processing module 200, and training module 300. Detailed Implementation

[0028] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0029] In the description of this invention, "several" means one or more, "multiple" means two or more, "greater than," "less than," and "exceeding" are understood to exclude the stated number, while "above," "below," and "within" are understood to include the stated number. The use of "first" and "second" in the description is merely for distinguishing technical features and should not be construed as indicating or implying relative importance, the number of indicated technical features, or the order of the indicated technical features. In the description of this invention, step numbers are merely for convenience of description or citation; the number of each step does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this invention.

[0030] Reference Figure 1 The method of an embodiment of the present invention includes:

[0031] (1) Receive the prediction results of the target deep learning model for unlabeled sample images, and select the prediction labels of several sample images based on the prediction results as training labeled images.

[0032] (2) Obtain the multi-scale image pyramid of the training labeled image and copy it into two copies; perform the first data processing on one copy of the multi-scale image pyramid to obtain the first image, and perform a second data processing different from the first data processing or do not perform any processing on the other copy of the multi-scale image pyramid to obtain the second image.

[0033] (3) Input the first image and the second image into the target deep learning model to obtain the corresponding first prediction label and second prediction label; calculate the corresponding loss function based on the first prediction label and the second prediction label, and iteratively update the target deep learning model.

[0034] The following will be based on Figure 2 Taking an example, this paper describes an implementation of applying this method to a deep learning model trained on a small number of samples, outlining the entire system's data processing flow. The method of this embodiment of the invention... Figure 2 This corresponds to semi-supervised iterative training. It should be understood that... Figure 2 This is merely an example; the method described in this embodiment of the invention is not limited thereto and can also be applied to other target deep learning models to improve the accuracy of image annotation while reducing annotation costs.

[0035] First, users upload image data through the dataset upload interface, and the data is cleaned. The data cleaning process includes removing damaged images, removing duplicate images, removing images with unsupported formats, and assessing data quality (quantity, resolution, etc.). Then, a subset of images is selected from the cleaned image data using a clustering method as initial images for annotation. This data is provided to users for annotation through an interactive interface; users can also use annotation tools to annotate the initial images. The annotated initial images (equivalent to sample annotation data) are then input, and a few-shot deep learning model (…) is used to… Figure 2 The model is trained using a few-shot model to obtain a coarse-labeled model. This coarse-labeled model can then sample and return partial data through an interactive interface for users to confirm the labeling effect. If the required accuracy is achieved, training stops and the labeling results are output. If the accuracy obtained from training the few-shot deep learning model does not meet the preset requirements, semi-supervised loop training is added to improve the model accuracy.

[0036] In this embodiment, the training method for few-shot deep learning can be implemented by applying a multi-scale refinement branch. This multi-scale refinement branch training method is applicable to, for example, Figure 3 The single-stage detection model and the two-stage detection model are shown. Figure 3 As shown, in a single-stage detection model, training data is input into a feature extractor to obtain a feature map, and the location and category of the target are directly found based on the feature map. In a two-stage detection model, in the first stage, training data is input into a feature extractor to obtain a feature map; in the second stage, candidate regions are determined based on the feature map, and the target region containing the target is identified from the candidate regions. Then, classification and target localization are performed based on the target region. Generally, the two-stage detection model provides more accurate target localization than the single-stage detection model.

[0037] The following section will use a two-stage detection model as an example to illustrate the application of multi-scale refinement branches. Figure 2 The illustrated embodiment describes the specific training process for few-shot deep learning. This training process includes the following steps S110 to S140.

[0038] Step S110: Input the sample annotation image, crop out the positive sample target from the sample annotation image, perform multi-scale scaling on the cropped positive sample target to generate a multi-scale image pyramid, which is used as the input of the first multi-scale refinement branch.

[0039] Step S120: Input the original sample labeled image into the backbone (also called the main branch), and input its corresponding multi-scale image pyramid into the first multi-scale refinement branch. After passing through the second feature extraction network, the corresponding image features are obtained. The second feature extraction network shares weights with the first feature extraction network.

[0040] Step S130: The original image features in the main branch follow the normal training process. The labeled sample images are input into the backbone network, and after passing through the first feature extraction network, the corresponding loss function is calculated.

[0041] Step S140: Input the multi-scale image pyramid into the first multi-scale refinement branch, calculate the corresponding loss function of the branch, and merge it into the loss function of the main branch to iteratively update the detection network.

[0042] Figure 4 The diagram illustrates a small-sample training process for an example two-stage detection model. In step S130 above, the labeled sample image is input into the backbone network. After passing through the first feature extraction network, it enters the ROI (region of interest) classification and regression network to obtain the final prediction result. The corresponding function of the backbone network is calculated, for example... Figure 4 The background classification loss, RPN bounding box regression loss, class classification loss, and ROI bounding box regression loss shown are used to iteratively update the backbone network. In this embodiment, the first feature extraction network can be an FPN network or other networks.

[0043] In step S140 above, since the image features obtained from the first multi-scale thinning branch are positive sample image features, it is only necessary to calculate the corresponding category classification loss and background classification loss for this branch. (Refer to...) Figure 4 The class classification loss of the first multi-scale refinement branch is merged into the class classification loss of the backbone; the background classification loss of the first multi-scale refinement branch is merged into the background class classification loss of the backbone. The backbone network is then iteratively updated using the updated background and class classification losses, as well as the RPN bounding box regression loss and ROI bounding box regression loss of the backbone network. During the iterative update process, the weights of the first feature extraction network are synchronized to the second feature extraction network through weight sharing.

[0044] If the detection model is a single-stage detection model, then the loss function calculated differs from that of the single-stage detection model when performing steps S110 to S140 above. For the single-stage detection model, the first multi-scale refinement branch only needs to calculate, for example, the classification loss, and then merge it into the backbone network.

[0045] The detection model can also be any other type of neural network structure, including various well-known structures. The loss function in this detection model can also be other types of loss functions. In this case, it is only necessary to calculate the corresponding loss function in steps S130 and S140 above.

[0046] Clearly, training methods for few-shot deep learning can also be implemented without applying multi-scale refinement branches. That is, both the original labeled image and the corresponding multi-scale image pyramid are input into the backbone network, and the corresponding loss function is calculated to iterate the detection network. At this point, the detection network only contains the backbone network (refer to...). Figure 4 The middle box in the middle does not have multi-scale refinement branches.

[0047] The enhancement effect of the first multi-scale refinement branch effectively improves the model's ability to recognize sample features and enhances its detection accuracy. This model can typically achieve 80% detection accuracy with only a few dozen data points.

[0048] When the annotation performance of a few-shot learning model fails to achieve the preset accuracy, in situations such as... Figure 2 The illustrated embodiments also enable semi-supervised loop training, using the method of this invention to further improve model accuracy. (Refer to...) Figure 4 The second multi-scale refinement branch in the model includes the following training steps:

[0049] Step 1: Use the detection network model trained in the previous round (equivalent to the target deep learning model) to predict all unlabeled data, select the prediction boxes with a confidence level above a certain threshold as the labels for the image, and use them as the sample labeled images for the input of this round of training.

[0050] In other words, the labeled images of the sample input for this round of training only include the predicted bounding boxes with a confidence level above a certain threshold.

[0051] Step 2: Duplicate the multi-scale image pyramid into two copies. Perform the first data processing on one copy to obtain the first image. Leave the other copy unprocessed or perform the second data processing to obtain the second image. Use the two image data, namely the first image and the second image, as a set of input samples and input them into the second multi-scale refinement branch for prediction through the third feature extraction network. The third feature extraction network shares weights with the first and second feature extraction networks.

[0052] In this embodiment, the first data processing is strong data enhancement, and the second data processing is weak data enhancement. Figure 4 The diagram illustrates how strong and weak data upscaling are performed on two replicated multi-scale image pyramids to obtain corresponding strongly enhanced and weakly enhanced images.

[0053] In this embodiment, if the other image is left unprocessed, it means that the multi-scale image pyramid is directly input.

[0054] Strong data enhancement combines various data enhancement methods. It can include methods that alter and do not alter the image data structure and characteristics, or it can be a combination of methods that only alter the image data structure and characteristics. In other words, strong data enhancement processes the input image using at least one method that changes the image data structure and characteristics, such as Gaussian blurring or adding noise. Weak data enhancement, on the other hand, uses methods like flipping and translation that do not change the image data structure and characteristics. In other words, strong enhancement can be considered as weak enhancement combined with methods that change the data structure and characteristics, or a combination thereof.

[0055] Step 3: For a set of input samples, the second image (equivalent to...) Figure 4 The first predicted label of the weakly augmented image is used as the pseudo-label, that is, it is set as the label of the first image (equivalent to the first image). Figure 4 The true labels of the strongly enhanced images are used to calculate the corresponding loss function of the second multi-scale thinning branch. This loss function is then merged into the loss function of the backbone network, and the backbone network is iteratively updated to optimize the network.

[0056] by Figure 4 Taking the two-stage detection model in the example, for instance, the category classification loss and background classification loss are calculated in the second multi-scale refinement branch, and then merged into the category classification loss and background classification loss in the backbone network. The backbone network is then iteratively updated to optimize the network.

[0057] If the target deep learning model is a single-stage detection model, then for example, the classification loss is calculated in the second multi-scale refinement branch, merged into the classification loss in the backbone network, and the backbone network is iteratively updated to optimize the network.

[0058] The target deep learning model in this embodiment is not limited to the single-stage or two-stage detection model described above; it can be any other type of neural network structure, including various well-known structures. The loss function in this detection model can also be other types of loss functions. In this case, it is only necessary to calculate the corresponding loss function in steps S130 and S140 above. The loss function in the target deep learning model in this embodiment can also be other types of loss functions.

[0059] Step 4: Repeat steps 1-3 for iterative training until the model meets the accuracy requirements or the set maximum number of iterations.

[0060] This training method can reduce the impact of noisy labels on network accuracy. Furthermore, through different data augmentations, the network learns more target patterns, becomes more robust to complex environments, and can better learn the representative features of targets, thereby improving model accuracy.

[0061] In other embodiments of the present invention, the processed first and second images can be directly input into the target deep learning model, and the corresponding loss function can be calculated to iteratively update the target deep learning model. That is, instead of constructing a multi-scale refinement branch for the target deep learning model, the images are directly input into the main branch of the target deep learning model to calculate the corresponding loss and iteratively update the target deep learning model. Figure 4 Taking the two-stage detection model as an example, the loss functions that need to be calculated include, for example, category classification loss, background classification loss, RPN bounding box regression loss, and ROI bounding box regression loss. Taking the single-stage detection model as an example, the loss function that needs to be calculated includes, for example, classification loss. It should be understood that the two-stage and single-stage detection models in this paper are only examples of the target deep learning model and are not intended to limit the target deep learning model. The above loss functions are also only illustrative examples, and the loss functions of the embodiments of this invention are not limited thereto.

[0062] Reference Figure 5 The internal modules of the device in this embodiment of the invention include: a selection module 100, an image processing module 200, and a training module 300.

[0063] Module 100 is used to receive the prediction results of the target deep learning model for unlabeled sample images. For each sample image, the prediction result includes multiple predicted labels. Then, several sample images and their corresponding predicted labels are selected from the prediction results as training labeled images. Specifically, the predicted labels can be selected based on their confidence level. That is, only predicted bounding boxes with a confidence level above a certain threshold in the image are used as labeled predicted bounding boxes; if there are no predicted bounding boxes with a confidence level above a certain threshold in an image, it is not included as a training labeled image.

[0064] The image processing module 200 receives the training labeled image selected by the selection module 100, obtains a multi-scale image pyramid of the training labeled image, and copies it into two copies. Different processing is applied to these two multi-scale image pyramids to obtain a first image and a second image. Specifically, a first data processing is performed on one multi-scale image pyramid to obtain the first image, and a second data processing, different from the first data processing, is performed on the other multi-scale image pyramid to obtain the second image. Alternatively, the first data processing is performed on one multi-scale image pyramid to obtain the first image, and the other multi-scale image pyramid is left unprocessed and directly used as the second image.

[0065] The first data processing described above specifically involves strong data enhancement. The second data processing specifically involves weak data enhancement.

[0066] The training module 300 is used to receive the first image and the second image input from the image processing module 200, input the first image and the second image into the target deep learning model, obtain the corresponding first prediction label and the second prediction label, calculate the corresponding loss function based on the first prediction label and the second prediction label, and iteratively update the target deep learning model.

[0067] During the iteration process, the target deep learning model can be retrained by selecting new training labeled images based on the prediction results of the previous round through module 100.

[0068] The device in this embodiment of the invention can reduce the impact of noise labels on network accuracy, and through different data augmentations, the network learns more target patterns, has higher robustness to complex environments, can better learn the representative features of the target, and improve model accuracy.

[0069] Although specific embodiments are described herein, those skilled in the art will recognize that many other modifications or alternative embodiments are also within the scope of this disclosure. For example, any of the functions and / or processing capabilities described in connection with a particular device or component can be performed by any other device or component. Furthermore, while various exemplary embodiments and architectures have been described according to embodiments of this disclosure, those skilled in the art will recognize that many other modifications to the exemplary embodiments and architectures described herein are also within the scope of this disclosure.

[0070] The foregoing description, with reference to block diagrams and flowcharts of systems, methods, systems, and / or computer program products according to exemplary embodiments, has described certain aspects of this disclosure. It should be understood that one or more blocks in the block diagrams and flowcharts, as well as combinations of blocks in the block diagrams and flowcharts, can be implemented by executing computer-executable program instructions, respectively. Similarly, according to some embodiments, some blocks in the block diagrams and flowcharts may not need to be executed in the order shown, or may not all need to be executed. Furthermore, additional components and / or operations beyond those shown in the blocks in the block diagrams and flowcharts may exist in some embodiments.

[0071] Therefore, blocks in block diagrams and flowcharts support combinations of means for performing a specified function, combinations of elements or steps for performing a specified function, and program instruction means for performing a specified function. It should also be understood that each block in a block diagram and flowchart, and combinations of blocks in block diagrams and flowcharts, can be implemented by a dedicated hardware computer system or a combination of dedicated hardware and computer instructions that performs a specific function, element, or step.

[0072] The program modules, applications, etc., described herein may include one or more software components, including, for example, software objects, methods, data structures, etc. Each such software component may include computer-executable instructions that, in response to execution, cause at least a portion of the functionality described herein (e.g., one or more operations of the exemplary methods described herein) to be performed.

[0073] Software components can be coded using any of a variety of programming languages. An exemplary programming language could be a low-level programming language, such as assembly language associated with a specific hardware architecture and / or operating system platform. Software components including assembly language instructions may need to be converted into executable machine code by an assembler before being executed by the hardware architecture and / or platform. Another exemplary programming language could be a higher-level programming language that is portable across multiple architectures. Software components including higher-level programming languages ​​may need to be converted into an intermediate representation by an interpreter or compiler before execution. Other examples of programming languages ​​include, but are not limited to, macro languages, shell or command languages, job control languages, scripting languages, database query or search languages, or report writing languages. In one or more exemplary embodiments, a software component containing instructions from one of the above-described programming language examples can be executed directly by the operating system or other software components without first being converted into another form.

[0074] Software components can be stored as files or other data storage structures. Software components of similar type or related function can be stored together in a specific directory, folder, or library. Software components can be static (e.g., pre-defined or fixed) or dynamic (e.g., created or modified at runtime).

[0075] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. An image annotation method, characterized in that, Includes the following steps: Image data is acquired, and a portion of the image data is selected as unlabeled sample images using a clustering method. Use the annotation tool to annotate the unannotated sample images to obtain annotated sample images; The labeled sample images are input into the target deep learning model for small-sample training to obtain a coarse-labeled model; wherein, the target deep learning model includes: a first multi-scale thinning branch, a backbone network and a second multi-scale thinning branch, the first multi-scale thinning branch extracts features through a second feature extraction network, the backbone network extracts features through a first feature extraction network, and the second multi-scale thinning branch extracts features through a third feature extraction network. If the coarse labeling model meets the preset requirements, training will stop and the labeling results will be output. If the accuracy of the coarse-labeled model does not meet the preset requirements, then additional semi-supervised loop training is added; The few-sample training includes: Input a sample labeled image, crop out positive sample targets from the sample labeled image, perform multi-scale scaling on the cropped positive sample targets to generate a multi-scale image pyramid, which serves as the input for the first multi-scale thinning branch; The labeled sample image is input into the backbone network, and its corresponding multi-scale image pyramid is input into the first multi-scale refinement branch. After passing through the second feature extraction network, the corresponding image features are obtained; wherein, the weights of the second feature extraction network and the first feature extraction network are shared. The labeled sample images are input into the backbone network, and after passing through the first feature extraction network, the corresponding loss function is calculated. The multi-scale image pyramid is input into the first multi-scale thinning branch, the corresponding loss function of the first multi-scale thinning branch is calculated, and then merged into the loss function of the backbone network. The semi-supervised loop training method includes: Receive the prediction results of the target deep learning model for unlabeled sample images, and select the predicted labels of several sample images based on the prediction results as training labeled images; Obtain the multi-scale image pyramid of the training labeled image and copy it into two copies; perform a first data processing on one copy of the multi-scale image pyramid to obtain a first image, and perform a second data processing different from the first data processing or perform no processing on the other copy of the multi-scale image pyramid to obtain a second image. The first image and the second image are input into the target deep learning model to obtain the corresponding first predicted label and the second predicted label. Based on the first predicted label and the second predicted label, the corresponding loss function is calculated, and the target deep learning model is iteratively updated. The method for selecting predicted labels for several sample images based on the prediction results and using them as training labeled images includes: Receive the prediction results from unlabeled sample images; From the prediction results, prediction boxes with a confidence level above a preset threshold are selected as annotations for the sample images; The labeled sample image is used as the training labeled image; The step of inputting the first image and the second image into the target deep learning model to obtain corresponding first predicted labels and second predicted labels, calculating the corresponding loss function based on the first predicted labels and the second predicted labels, and iteratively updating the target deep learning model includes: The first image and the second image are used as a set of input samples and fed into the second multi-scale thinning branch. The third feature extraction network is used for prediction to obtain the first and second predicted labels. The third feature extraction network shares weights with the first and second feature extraction networks. The second predicted label is used as the true label of the first image after the first data processing, and compared with the first predicted label. The loss function corresponding to the second multi-scale thinning branch is calculated, and the loss function is merged into the loss function of the backbone network. The backbone network is then iteratively updated.

2. The image annotation method according to claim 1, characterized in that, The first data processing is strong data augmentation, and the second data processing is weak data augmentation.

3. An image annotation apparatus, using the method of any one of claims 1 to 2, characterized in that, include: The acquisition module is used to acquire image data and select a portion of the image data as unlabeled sample images using a clustering method. The annotation module is used to annotate unannotated sample images using annotation tools to obtain annotated sample images; The few-sample training module is used to input the labeled sample images into the target deep learning model for few-sample training to obtain a coarse-labeled model; wherein, the target deep learning model includes: a first multi-scale thinning branch, a backbone network and a second multi-scale thinning branch, the first multi-scale thinning branch extracts features through a second feature extraction network, the backbone network extracts features through a first feature extraction network, and the second multi-scale thinning branch extracts features through a third feature extraction network. The training stop module is used to stop training and output the annotation results if the coarse labeling model reaches the preset requirements. An additional module is provided to add semi-supervised loop training if the accuracy of the coarse-calibrated model does not meet the preset requirements. The small sample training module is also used for: Input a sample labeled image, crop out positive sample targets from the sample labeled image, perform multi-scale scaling on the cropped positive sample targets to generate a multi-scale image pyramid, which serves as the input for the first multi-scale thinning branch; The labeled sample image is input into the backbone network, and its corresponding multi-scale image pyramid is input into the first multi-scale refinement branch. After passing through the second feature extraction network, the corresponding image features are obtained; wherein, the weights of the second feature extraction network and the first feature extraction network are shared. The labeled sample images are input into the backbone network, and after passing through the first feature extraction network, the corresponding loss function is calculated. The multi-scale image pyramid is input into the first multi-scale thinning branch, the corresponding loss function of the first multi-scale thinning branch is calculated, and then merged into the loss function of the backbone network. The append module is followed by: The selection module is used to receive the prediction results of the target deep learning model for unlabeled sample images, and select the predicted labels of several sample images based on the prediction results as training labeled images. The image processing module is used to acquire the multi-scale image pyramid of the training labeled image, and copy it into two copies; perform a first data processing on one copy of the multi-scale image pyramid to obtain a first image, and perform a second data processing different from the first data processing or no processing on the other copy of the multi-scale image pyramid to obtain a second image. The training module is used to input the first image and the second image into the target deep learning model to obtain the corresponding first predicted label and the second predicted label, calculate the corresponding loss function based on the first predicted label and the second predicted label, and iteratively update the target deep learning model. The selection module is also used for: Receive the prediction results from unlabeled sample images; From the prediction results, prediction boxes with a confidence level above a preset threshold are selected as annotations for the sample images; The labeled sample image is used as the training labeled image; The training module is also used for: The first image and the second image are used as a set of input samples and fed into the second multi-scale thinning branch. The third feature extraction network is used for prediction to obtain the first and second predicted labels. The third feature extraction network shares weights with the first and second feature extraction networks. The second predicted label is used as the true label of the first image after the first data processing, and compared with the first predicted label. The loss function corresponding to the second multi-scale thinning branch is calculated, and the loss function is merged into the loss function of the backbone network. The backbone network is then iteratively updated.

4. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method of any one of claims 1 to 2.

Citation Information

Patent Citations

  • Image labeling method and device and computer readable storage medium

    CN110059696A

  • Image pre-labeling method and device, electronic equipment and medium

    CN112418287A