Pre-annotation Method, System, Device and Storage Medium Based on Multi-Expert Model

By using the combined detection and threshold judgment of multiple expert models in the pre-labeling method, the problem that the labeling results in the prior art depend on the accuracy of the detection model is solved, and a more efficient and accurate pre-labeling process is achieved.

CN118644662BActive Publication Date: 2025-06-13BEIJING SINOITS TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410767342.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-14
Publication Date
2025-06-13
Estimated Expiration
2044-06-14

AI Technical Summary

Technical Problem

The labeling results of existing pre-labeling methods depend on the accuracy of the detection model, resulting in missed detection and mislabeling that require manual supplementation or correction, especially in difficult samples in dense areas, the model capability is limited, resulting in pre-labeling time-consuming and inefficient.

Method used

The pre-labeling method based on the multi-expert model is adopted, and the final pre-labeling box is determined by detecting the first target detection expert model and the second target detection expert model.

Benefits of technology

It significantly improves the accuracy of labeling, reduces the labeling time, improves the efficiency of labeling, and reduces the need for manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118644662B_ABST
    Figure CN118644662B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing technology, and specifically discloses a pre-annotation method, system, device and storage medium based on a multi-expert model. The method includes: using a first object detection expert model to detect a to-be-annotated image to obtain first prediction values of a plurality of first detection frames; determining first detection frames with first prediction values less than a first threshold as second detection frames, and judging whether there are overlapping and dense detection frames among all the second detection frames to obtain a first judgment result; when the first judgment result is negative, using a second object detection expert model to detect all the second detection frames to obtain second prediction values of a plurality of second detection frames; determining first detection frames with first prediction values not less than the first threshold and second detection frames with second prediction values not less than a second threshold as pre-annotation frames of the to-be-annotated image. The present invention can significantly improve the accuracy of annotation, reduce the annotation time, and improve the annotation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] The labeling result of the existing pre-labeling method depends on the accuracy of the detection model. If there are undetected objects in the detection model, manual labeling is required to supplement them. If there are misdetected objects, manual removal is needed. For difficult samples in dense areas, the model has limited capabilities. When there are a particularly large number of bounding boxes, the pre-labeling in this way is actually more time-consuming.

[0003] Therefore, there is an urgent need to provide a technical solution to solve the above problems. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a pre-labeling method, system, device, and storage medium based on a multi-expert model.

[0005] In a first aspect, the present invention provides a pre-labeling method based on a multi-expert model. The technical solution of this method is as follows:

[0006] Use a first object detection expert model to detect the image to be labeled and obtain the first prediction values of multiple first detection bounding boxes;

[0007] Determine the first detection bounding boxes with the first prediction values less than the first threshold as the second detection bounding boxes, and determine whether there are overlapping and dense detection bounding boxes among all the second detection bounding boxes to obtain a first judgment result;

[0008] When the first judgment result is negative, use a second object detection expert model to detect all the second detection bounding boxes and obtain the second prediction values of multiple second detection bounding boxes;

[0009] Determine the first detection bounding boxes with the first prediction values not less than the first threshold and the second detection bounding boxes with the second prediction values not less than the second threshold as the pre-labeling bounding boxes of the image to be labeled.

[0010] The beneficial effects of a pre-labeling method based on a multi-expert model of the present invention are as follows:

[0011] The method of the present invention can significantly improve the labeling accuracy, reduce the labeling time, and improve the labeling efficiency.

[0012] On the basis of the above solution, a pre-labeling method based on a multi-expert model of the present invention can also be improved as follows.

[0013] In an optional manner, the second object detection expert model includes: a small object expert model and an object classification expert model; the step of using the second object detection expert model to detect all the second detection bounding boxes and obtain the second prediction values of multiple second detection bounding boxes includes:

[0014] Using the small target expert model, each second detection box of the small target is detected respectively to obtain the second prediction value of each second detection box of the small target;

[0015] Using the target classification expert model, each second detection box of the large target is detected respectively to obtain the second prediction value of each second detection box of the large target.

[0016] In an alternative manner, it further includes:

[0017] When there are second detection boxes of large targets with minority classes among the second detection boxes of large targets whose second prediction values are less than the second threshold, the minority class expert model is used to detect the second detection boxes of large targets with minority classes, and the second detection boxes of large targets with minority classes not less than the third threshold are determined as the pre-annotation boxes of the image to be annotated.

[0018] In an alternative manner, it further includes:

[0019] When the first judgment result is yes, the second detection boxes with cross-density are used as the first difficult detection boxes, and the difficult sample region expert model is used to detect all the first difficult detection boxes. The first difficult detection boxes not less than the fourth threshold and the second detection boxes of small targets whose second prediction values are less than the second threshold are used as the second difficult detection boxes;

[0020] The difficult sample expert model is used to detect all the second difficult detection boxes, and the second difficult detection boxes not less than the fifth threshold are determined as the pre-annotation boxes of the image to be annotated.

[0021] In an alternative manner, it further includes:

[0022] The multi-modal large model is used to detect the second difficult detection boxes less than the fifth threshold, and the second difficult detection boxes of large targets are input into the classification expert model for detection to determine whether the second difficult detection boxes of large targets are determined as the pre-annotation boxes of the image to be annotated.

[0023] In a second aspect, the present invention provides a pre-annotation system based on a multi-expert model, and the technical solution of the system is as follows:

[0024] It includes: a first processing module, a first judgment module, a second processing module, and a first operation module;

[0025] The first processing module is used for: using the first target detection expert model to detect the image to be annotated to obtain the first prediction values of multiple first detection boxes;

[0026] The first judgment module is used to: determine the first detection boxes with the first prediction value less than the first threshold as the second detection boxes, and judge whether there are overlapping and dense detection boxes among all the second detection boxes to obtain a first judgment result;

[0027] The second processing module is used to: when the first judgment result is negative, use the second target detection expert model to detect all the second detection boxes to obtain the second prediction values of the multiple second detection boxes;

[0028] The first operation module is used to: determine the first detection boxes with the first prediction value not less than the first threshold and the second detection boxes with the second prediction value not less than the second threshold as the pre-annotation boxes of the image to be annotated.

[0029] The beneficial effects of a pre-annotation system based on multiple expert models of the present invention are as follows:

[0030] The system of the present invention can significantly improve the accuracy of annotation, reduce the annotation time, and improve the annotation efficiency.

[0031] On the basis of the above solution, a pre-annotation system based on multiple expert models of the present invention can also be improved as follows.

[0032] In an optional manner, the second target detection expert model includes: a small target expert model and a target classification expert model; the second processing module is specifically used to:

[0033] Use the small target expert model to detect each small target second detection box respectively to obtain the second prediction value of each small target second detection box;

[0034] Use the target classification expert model to detect each large target second detection box respectively to obtain the second prediction value of each large target second detection box.

[0035] In an optional manner, it further includes: a second operation module;

[0036] The second operation module is used to: when there are large target second detection boxes of a minority class among the large target second detection boxes with the second prediction value less than the second threshold, use the minority class expert model to detect the large target second detection boxes of the minority class, and determine the large target second detection boxes of the minority class not less than the third threshold as the pre-annotation boxes of the image to be annotated.

[0037] In a third aspect, the technical solution of an electronic device of the present invention is as follows:

[0038] It includes a memory, a processor, and a program stored on the memory and running on the processor. When the processor executes the program, the steps of the pre-annotation method based on the multi-expert model of the present invention are implemented.

[0039] In a fourth aspect, a technical solution of a computer-readable storage medium provided by the present invention is as follows:

[0040] Instructions are stored in the computer-readable storage medium. When the computer-readable storage medium reads the instructions, the computer-readable storage medium is caused to execute the steps of the pre-annotation method based on the multi-expert model of the present invention.

[0041] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other objects, features, and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically exemplified below. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The drawings are only used to illustrate the embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0043] Figure 1 is a schematic flowchart of an embodiment of a pre-annotation method based on a multi-expert model of the present invention;

[0044] Figure 2 is a schematic diagram of the principle of a small target expert model;

[0045] Figure 3 is a schematic diagram of the principle of a difficult sample expert model;

[0046] Figure 4 is a schematic diagram of the principle of a multi-modal large model;

[0047] Figure 5 is a schematic structural diagram of an embodiment of a pre-annotation system based on a multi-expert model of the present invention;

[0048] Figure 6 is a schematic structural diagram of an embodiment of an electronic device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] Hereinafter, the exemplary embodiments of the present invention will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein.

[0050] Figure 1The flowchart of an embodiment of a pre-annotation method based on a multi-expert model provided by the present invention is shown. The pre-annotation method based on the multi-expert model can be executed by an electronic device such as a terminal device or a server. Among them, the terminal device can be any fixed or mobile terminal such as a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The server can be a single server or a server cluster composed of multiple servers. Any electronic device can implement the pre-annotation method based on the multi-expert model by a processor calling computer-readable instructions stored in a memory. As Figure 1 shown, the method includes the following steps:

[0051] S1. Use a first object detection expert model to detect the image to be annotated, and obtain first prediction values of multiple first detection frames.

[0052] Among them, the first object detection expert model refers to a model using the transformer series, such as CO-DETR, MoCaE, DINO, etc. The image to be annotated is the image that needs to be annotated in this embodiment, such as an image of the toll station intersection on a highway. The first detection frame is the detection frame output by the first object detection expert model (CO-DETR model). The first prediction value is the score of the first detection frame. For example, if the score of a certain first detection frame for a car is 0.75 (equivalent to the probability that the image in the first detection frame is a car is 0.75), then the first prediction value of the target (car) in this first detection frame is 0.75.

[0053] S2. Determine the first detection frames with first prediction values less than the first threshold as second detection frames, and determine whether there are overlapping and dense detection frames among all the second detection frames to obtain a first judgment result.

[0054] Among them, the first threshold can be set according to the actual situation (such as 0.6), and there is no limit here. The overlapping and dense detection frames are multiple detection frames with overlapping between the detection frames, that is: among three or more detection frames, for any two detection frames, there is a third frame that intersects both of these two frames, and these detection frames are overlapping and dense detection frames.

[0055] S3. When the first judgment result is negative, use a second object detection expert model to detect all the second detection frames, and obtain second prediction values of multiple second detection frames.

[0056] Among them, the second target detection expert model includes: a small target expert model and a target classification expert model. The small target expert model refers to a model that can optimize the detection results of small targets, such as yolov9-e-dual. The target classification expert model is cspresnet-50.

[0057] It should be noted that the size of the target is defined according to the area of the target instance (the number of pixels occupied in the image). The definition is as follows:

[0058] ① Small objects: The area of the target is less than 32*32 pixels.

[0059] ② Medium objects: The area of the target is between 32*32 and 96*96 pixels.

[0060] ③ Large objects: The area of the target is greater than 96*96 pixels.

[0061] S4. Determine the pre-annotation boxes of the to-be-annotated image as the first detection boxes with the first prediction value not less than the first threshold and the second detection boxes with the second prediction value not less than the second threshold.

[0062] Among them, the second threshold and the first threshold can be the same or different, and there is no limit here. The pre-annotation box is equivalent to the label of the to-be-annotated image, and its format is COCO, VOC or YOLO, which is used for subsequent training of the model.

[0063] In an optional manner, the step of using the second target detection expert model to detect all the second detection boxes to obtain the second prediction values of the multiple second detection boxes includes:

[0064] Use the small target expert model to detect each small target second detection box respectively to obtain the second prediction value of each small target second detection box.

[0065] Among them, the second detection boxes are divided into: small target second detection boxes and large target second detection boxes.

[0066] It should be noted that the principle of the small target expert model is as Figure 2 shown. Figure 2 There is a pedestrian in the upper left corner, but the model predicts it as a motorperson. The second prediction value of the small target second detection box is 0.34.

[0067] Use the target classification expert model to detect each large target second detection box respectively to obtain the second prediction value of each large target second detection box.

[0068] In an optional manner, it further includes:

[0069] When there are large target second detection boxes of minority classes among the large target second detection boxes where the second predicted value is less than the second threshold, use the minority class expert model to detect the large target second detection boxes of minority classes, and determine the large target second detection boxes of minority classes that are not less than the third threshold as the pre-annotation boxes of the to-be-annotated image.

[0070] Among them, the minority class expert model is the CO-DETR model, and this CO-DETR is only trained for minority classes (such as obstacles) during training, which is different from the first target detection expert model. The third threshold can be set according to the actual situation and is not limited here.

[0071] In an alternative manner, it further includes:

[0072] When the first judgment result is yes, regard the second detection boxes with cross-density as the first difficult detection boxes, and use the difficult sample region expert model to detect all the first difficult detection boxes, and regard the first difficult detection boxes that are not less than the fourth threshold and the small target second detection boxes where the second predicted value is less than the second threshold as the second difficult detection boxes.

[0073] Among them, the difficult sample region expert model uses the yolov5s model, and this model does not need to obtain accurate detection boxes, but only needs to detect small target dense regions or difficult sample regions. The fourth threshold can be set according to the actual situation and is not limited here.

[0074] Use the difficult sample expert model to detect all the second difficult detection boxes, and determine the second difficult detection boxes that are not less than the fifth threshold as the pre-annotation boxes of the to-be-annotated image.

[0075] Among them, the difficult sample expert model is yolov9-e-dual. The fifth threshold can be set according to the actual situation and is not limited here.

[0076] It should be noted that the principle of the difficult sample expert model is as Figure 3 shown. In Figure 3 , it is easy to misdetect the areas where cyclists and tricycles are dense and overlapping. Here, the difficult sample expert model is used. First, the model detects this area, and after cropping this area, it is detected separately.

[0077] In an alternative manner, it further includes:

[0078] Use the multi-modal large model to detect the second difficult detection boxes that are less than the fifth threshold, and input the large target second difficult detection boxes into the classification expert model for detection to determine whether to determine the large target second difficult detection boxes as the pre-annotation boxes of the to-be-annotated image.

[0079] Among them, the multi-modal large model is the multi-modal large model of Tongyi Qianwen. This model uses prompt words to determine whether it is the target, or whether the target is fuzzy and cannot be judged. For the pre-annotated boxes determined in this embodiment, they are marked and stored. For those detected by the model but determined not to be pre-annotated boxes, corresponding smearing is performed in the image, so that it can be confirmed that there is no target in this area, thus not affecting the subsequent training of the model.

[0080] It should be noted that the principle of the multi-modal large model is as Figure 4 shown. Through the joint judgment of the multi-modal large model of Tongyi Qianwen and the prediction results of other experts, it is determined whether to smear this area.

[0081] The technical solution of this embodiment can significantly improve the accuracy of annotation, reduce the annotation time, and improve the annotation efficiency.

[0082] Figure 5 FIG. shows a schematic structural diagram of an embodiment of a pre-annotation system 200 based on a multi-expert model provided by the present invention. As Figure 5 shown, the system 200 includes: a first processing module 210, a first judgment module 220, a second processing module 230, and a first operation module 240;

[0083] The first processing module 210 is configured to: use a first object detection expert model to detect the image to be annotated, and obtain first prediction values of a plurality of first detection boxes;

[0084] The first judgment module 220 is configured to: determine the first detection boxes with first prediction values less than a first threshold as second detection boxes, and judge whether there are cross-dense detection boxes among all the second detection boxes, and obtain a first judgment result;

[0085] The second processing module 230 is configured to: when the first judgment result is negative, use a second object detection expert model to detect all the second detection boxes, and obtain second prediction values of a plurality of second detection boxes;

[0086] The first operation module 240 is configured to: determine the first detection boxes with first prediction values not less than the first threshold and the second detection boxes with second prediction values not less than a second threshold as the pre-annotation boxes of the image to be annotated.

[0087] In an alternative manner, the second object detection expert model includes: a small object expert model and an object classification expert model; the second processing module 230 is specifically configured to:

[0088] Use the small object expert model to detect each small object second detection box respectively, and obtain second prediction values of each small object second detection box;

[0089] Using the target classification expert model, each second detection box of the large target is detected respectively to obtain the second prediction value of each second detection box of the large target.

[0090] In an alternative manner, it further includes: a second operation module;

[0091] The second operation module is configured to: when there are second detection boxes of large targets of minority classes among the second detection boxes of large targets whose second prediction values are less than the second threshold, use the minority class expert model to detect the second detection boxes of large targets of minority classes, and determine the second detection boxes of large targets of minority classes that are not less than the third threshold as the pre-annotation boxes of the to-be-annotated image.

[0092] In an alternative manner, it further includes: a third operation module; the third operation module is configured to:

[0093] When the first judgment result is yes, take the second detection boxes with cross-density as the first difficult detection boxes, and use the difficult sample region expert model to detect all the first difficult detection boxes, and take the first difficult detection boxes that are not less than the fourth threshold and the second detection boxes of small targets whose second prediction values are less than the second threshold as the second difficult detection boxes;

[0094] Use the difficult sample expert model to detect all the second difficult detection boxes, and determine the second difficult detection boxes that are not less than the fifth threshold as the pre-annotation boxes of the to-be-annotated image.

[0095] In an alternative manner, it further includes: a fourth operation module;

[0096] The fourth operation module is configured to: use the multi-modal large model to detect the second difficult detection boxes that are less than the fifth threshold, and input the second difficult detection boxes of large targets into the classification expert model for detection to determine whether to determine the second difficult detection boxes of large targets as the pre-annotation boxes of the to-be-annotated image.

[0097] For the parameters and the steps of each module in the above-mentioned pre-annotation system 200 based on multi-expert models in this embodiment to implement corresponding functions, reference can be made to the parameters and steps in the embodiment of the pre-annotation method based on multi-expert models in the above text, which will not be elaborated here.

[0098] As Figure 6As shown in the figure, an electronic device 300 according to an embodiment of the present invention includes a processor 320, the processor 320 is coupled to a memory 310, and at least one computer program 330 is stored in the memory 310. The at least one computer program 330 is loaded and executed by the processor 320 so that the electronic device 300 implements any one of the above-mentioned pre-annotation methods based on a multi-expert model. Specifically:

[0099] The electronic device 300 may vary greatly due to configuration or performance differences, and may include one or more processors 320 (Central Processing Units, CPUs) and one or more memories 310. Among them, at least one computer program 330 is stored in the one or more memories 310, and the at least one computer program 330 is loaded and executed by the one or more processors 320 so that the electronic device 300 implements any one of the pre-annotation methods based on a multi-expert model provided in the above embodiments. Of course, the electronic device 300 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The electronic device 300 may also include other components for implementing device functions, which will not be elaborated here.

[0100] A computer-readable storage medium according to an embodiment of the present invention stores at least one computer program, and the at least one computer program is loaded and executed by a processor so that a computer implements any one of the above-mentioned pre-annotation methods based on a multi-expert model.

[0101] Optionally, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0102] In an exemplary embodiment, a computer program product or a computer program is also provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions so that the electronic device executes any one of the above-mentioned pre-annotation methods based on a multi-expert model.

[0103] It should be noted that the terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to limit a specific order or sequence. In appropriate cases, the order of use of similar objects can be interchanged so that the embodiments of the present application described herein can be implemented in an order other than the illustrated or described order.

[0104] Those skilled in the art know that the present invention can be implemented as a system, a method or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, that is, it can be completely hardware, or completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, which is generally referred to as "circuit", "module" or "system" herein. In addition, in some embodiments, the present invention can also be implemented in the form of a computer program product in one or more computer-readable media, which contain computer-readable program code.

[0105] Any combination of one or more computer-readable media can be adopted. The computer-readable media can be computer-readable signal media or computer-readable storage media. The computer-readable storage media can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable storage media can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.

[0106] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A pre-labeling method based on a multi-expert model, characterized in that: include: Using the first object detection expert model, the image to be labeled is detected to obtain first prediction values ​​of multiple first detection boxes; Determine a first detection frame whose first prediction value is less than a first threshold as a second detection frame, and determine whether there are any intersection-intensive detection frames among all second detection frames to obtain a first determination result; When the first judgment result is no, using the second object detection expert model to detect all second detection frames to obtain second prediction values ​​of the plurality of second detection frames; Determine a first detection frame whose first prediction value is not less than a first threshold and a second detection frame whose second prediction value is not less than a second threshold as pre-annotated frames of the image to be annotated; The second target detection expert model includes: a small target expert model and a target classification expert model; the step of detecting all second detection frames using the second target detection expert model to obtain second prediction values ​​of multiple second detection frames includes: Using the small target expert model, each small target second detection frame is detected respectively to obtain a second prediction value of each small target second detection frame; Using the target classification expert model, each large target second detection frame is detected respectively to obtain a second prediction value of each large target second detection frame; Also includes: When there is a large target second detection frame of a minority category in the large target second detection frame whose second prediction value is less than the second threshold, using the minority category expert model to detect the large target second detection frame of the minority category, and determining the large target second detection frame of the minority category that is not less than the third threshold as the pre-annotated frame of the image to be annotated; Also includes: When the first judgment result is yes, the second detection frame with dense intersection is used as the first difficult detection frame, and the difficult sample area expert model is used to detect all the first difficult detection frames, and the first difficult detection frame not less than the fourth threshold and the small target second detection frame whose second prediction value is less than the second threshold are used as the second difficult detection frame; All second difficult detection frames are detected by using the difficult sample expert model, and second difficult detection frames whose value is not less than a fifth threshold are determined as pre-labeled frames of the image to be labeled.

2. The pre-labeling method based on a multi-expert model according to claim 1, characterized in that: Also includes: The multimodal large model is used to detect the second difficult detection box that is smaller than the fifth threshold, and the second difficult detection box of the large target is input into the classification expert model for detection to determine whether the second difficult detection box of the large target is determined as the pre-labeling box of the image to be labeled.

3. A pre-labeling system based on a multi-expert model, characterized in that: include: A first processing module, a first judging module, a second processing module and a first operating module; The first processing module is used to: detect the image to be labeled using the first target detection expert model to obtain first prediction values ​​of multiple first detection boxes; The first judgment module is used to: determine the first detection frame whose first prediction value is less than the first threshold as the second detection frame, and judge whether there are cross-dense detection frames among all the second detection frames to obtain a first judgment result; The second processing module is used to: when the first judgment result is no, use the second target detection expert model to detect all second detection frames to obtain second prediction values ​​of multiple second detection frames; The first operation module is used to: determine a first detection frame whose first prediction value is not less than a first threshold and a second detection frame whose second prediction value is not less than a second threshold as pre-annotated frames of the image to be annotated; The second target detection expert model includes: a small target expert model and a target classification expert model; the second processing module is specifically used for: Using the small target expert model, each small target second detection frame is detected respectively to obtain a second prediction value of each small target second detection frame; Using the target classification expert model, each large target second detection frame is detected respectively to obtain a second prediction value of each large target second detection frame; Also includes: a second operation module; The second operation module is used to: when there is a large target second detection frame of a minority category in the large target second detection frame whose second prediction value is less than the second threshold, use the minority category expert model to detect the large target second detection frame of the minority category, and determine the large target second detection frame of the minority category that is not less than the third threshold as the pre-annotated frame of the image to be annotated; Also includes: a third operation module; the third operation module is used to: When the first judgment result is yes, the second detection frame with dense intersection is used as the first difficult detection frame, and the difficult sample area expert model is used to detect all the first difficult detection frames, and the first difficult detection frame not less than the fourth threshold and the small target second detection frame whose second prediction value is less than the second threshold are used as the second difficult detection frame; All second difficult detection frames are detected by using the difficult sample expert model, and second difficult detection frames whose value is not less than a fifth threshold are determined as pre-labeled frames of the image to be labeled.

4. An electronic device, characterized in that: The electronic device includes a processor, the processor is coupled to a memory, at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor so that the electronic device implements the pre-labeling method based on a multi-expert model as described in claim 1 or 2.

5. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by the processor so that the computer-readable storage medium implements the pre-labeling method based on a multi-expert model as described in claim 1 or 2.

Citation Information

Patent Citations

  • Target detection method and device, computer equipment and storage medium

    CN116109816A