Annotation support device and annotation support method

The annotation support device enhances AI model training data creation by highlighting unreliable annotation results with caution flags, addressing inefficiencies and errors in manual correction processes.

JP7723542B2Active Publication Date: 2025-08-14DENSO TEN LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021145893
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-08
Publication Date
2025-08-14
Estimated Expiration
2041-09-08

AI Technical Summary

Technical Problem

Manual annotation of AI model inference results is inefficient and prone to errors due to the need for thorough checking of accurate inferences, leading to neglect of low reliability scores and prolonged processing times.

Method used

An annotation support device that highlights annotation results with reliability scores close to a threshold or partial object detection, using caution flags to focus user attention on potentially erroneous areas, thereby reducing annotation steps and errors.

Benefits of technology

Simultaneously reduces the number of annotation steps and suppresses errors by concentrating user effort on critical areas, improving efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007723542000001
    Figure 0007723542000001
  • Figure 0007723542000002
    Figure 0007723542000002
  • Figure 0007723542000003
    Figure 0007723542000003
Patent Text Reader

Abstract

To provide a technique that can reduce annotation man-hours while reducing annotation errors.SOLUTION: An annotation support device comprises an acquisition part and a display part. The acquisition part acquires image data to which an annotation result and a reliability score for the annotation result are given. The display part noticeably displays the annotation result if the annotation result meets at least either one of a first condition that the reliability score is in the vicinity of a threshold or a second condition that only part of an object has been detected.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technology for supporting annotation for creating training data used in training an artificial intelligence model. [Background technology]

[0002] Object detection is performed by using an AI model to recognize objects contained in images. Annotation to create training data used for learning AI models generally requires a large amount of work. This is because a large number of images must be annotated to improve the object detection accuracy of the AI model.

[0003] To reduce the amount of work required for annotation, there is a method of annotating data by manually correcting the inference results of a trained AI model. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Publication No. 2019-78554 Summary of the Invention [Problem to be solved by the invention]

[0005] However, when manually correcting the inference results of a trained AI model, the worker must thoroughly check the image to ensure there are no errors in the inference results. In other words, the worker must concentrate on checking even the parts where the AI model is almost certain to have made accurate inferences. As a result, the worker is unable to concentrate on the parts that they should be concentrating on (such as parts with a low reliability score in the AI model's inference results), which could lead to correction errors. In addition, because the image must be thoroughly checked, the process takes a long time.

[0006] The object detection method disclosed in Patent Document 1 detects an object when the reliability derived by statistically processing scores related to the object exceeds a threshold. In other words, Patent Document 1 does not employ a method of annotating by manually correcting the inference results of a trained artificial intelligence model.

[0007] In view of the above-mentioned problems, the present invention aims to provide a technology that can simultaneously reduce the number of annotation steps and suppress annotation errors. [Means for solving the problem]

[0008] An exemplary annotation support device of the present invention includes an acquisition unit that acquires image data to which an annotation result and a reliability score for the annotation result have been assigned, and a display unit that prominently displays the annotation result when the annotation result satisfies at least one of a first condition that the reliability score is close to a threshold value and a second condition that only a portion of an object has been detected. [Effects of the Invention]

[0009] According to the exemplary embodiment of the present invention, it is possible to reduce the number of annotation steps and suppress annotation errors at the same time. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram illustrating an example of a schematic configuration of a server device according to an embodiment of the present invention; [Figure 2] A flowchart showing an example of the operation of the server device according to the present embodiment. [Figure 3] A diagram showing an example of creating image data to which annotation results and reliability scores for the annotation results have been assigned. [Figure 4] A diagram showing an example of a highlighted image. [Figure 5] Figure 10 shows another example of a highlighted image. [Figure 6] A diagram showing an example of a corrected image. [Figure 7] A diagram showing an example of a normal display DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the drawings.

[0012] <1. Configuration of Server Device and Terminal Device> 1 is a diagram showing an example of a schematic configuration of a server device 1 according to this embodiment. The server device 1 is an example of an annotation support device.

[0013] In this embodiment, annotation supported by the server device 1 refers to the act of adding information such as tags, annotations, and bounding boxes to data such as images used in machine learning.

[0014] The server device 1 creates training data used for training an artificial intelligence model. The server device 1 includes a control unit 11, a recording unit 12, a communication unit 13, a display unit 14, and an input unit 15. The control unit 11, the recording unit 12, and the communication unit 13 constitute an information processing device, and the display unit 14 constitutes a display device.

[0015] The server device 1 is not limited to a server device set up in a single location, but may be a distributed server device in which components are installed in multiple locations. The server device 1 may also be configured as a cloud server.

[0016] The control unit 11 is a computer having at least one processor. Specifically, the control unit 11 is a computer having a CPU (Central Processing Unit), RAM (Random Access Memory), and ROM (Read Only Memory), which are not shown. The control unit 11 processes, transmits, and receives information based on a program recorded in the recording unit 12, and controls the entire server device 1.

[0017] The communication unit 13 performs wireless communication with an external device via a network (not shown). The communication unit 13 is capable of wireless communication with multiple external devices. The communication unit 13 is an example of an acquisition unit that acquires image data to which annotation results and reliability scores for the annotation results have been assigned.

[0018] The annotation results are displayed on the display unit 14. As the display unit 14, for example, a thin display such as a liquid crystal display or an organic EL (Electro Luminescence) display can be used.

[0019] The input unit 15 inputs instruction information (correction instruction information) related to correction of the annotation result. The input unit 15 is a device operated by a user. For example, a keyboard, a pointing device, a touch panel, a joystick, etc. can be used as the input unit 15.

[0020] <2. Operation of the server device> Fig. 2 is a flowchart showing an example of the operation of the server device 1. For example, when a user performs a predetermined operation (operation for starting operation) on the input unit 15, the server device 1 starts the operation of the flowchart shown in Fig. 2. The server device 1 performs the operation of the flowchart shown in Fig. 2, thereby implementing the annotation support method according to this embodiment. The following description will be given with reference to Figs. 2 to 7.

[0021] See Figure 2. The communication unit 13 of the server device 1 acquires image data sent from an external device (for example, the trained artificial intelligence model 2 described below) (step S10). The image acquired by the communication unit 13 of the server device 1 is image data to which annotation results and reliability scores for the annotation results have been assigned. Because image data to which annotation results and reliability scores for the annotation results have been assigned is acquired, the amount of annotation work can be reduced.

[0022] FIG. 3 is a diagram showing an example of image data to which annotation results and reliability scores for the annotation results have been assigned.

[0023] The trained artificial intelligence model 2 receives input of image P1 data. Image P1 is an image showing person H1 and person H2. Person H1 is facing forward, and person H2 is facing backward. Therefore, image P1 shows the face F1 of person H1, but not the face of person H2.

[0024] The trained AI model 2 detects whole people and human faces as separate classes. The trained AI model 2 calculates a confidence score for each detection candidate. If the confidence score exceeds a threshold, the trained AI model 2 adopts the detection candidate as an annotation result. In the following explanation, the threshold is set to 0.25 as an example.

[0025] The trained artificial intelligence model 2 generates an image P2 by annotating the image P1 and outputs the data of the image P2.

[0026] The data of image P2 is image data to which the annotation results and the confidence scores for the annotation results are added as meta-information. The annotation results are visualized by bounding boxes that enclose the outlines of the detected objects.

[0027] Image P2 includes a bounding box BB1 that is an annotation result in which face F1 of person H1 is detected, and a bounding box BB2 that is an annotation result in which person H2 is detected. For example, by changing the colors of the bounding boxes that are annotation results in which faces are detected and the bounding boxes that are annotation results in which people are detected, it becomes easier to see which objects are detected.

[0028] Referring again to Fig. 2, the control unit 11 of the server device 1 determines whether the image data acquired by the communication unit 13 includes an annotation result that satisfies at least one of the first condition and the second condition (step S20).

[0029] The first condition is that the reliability score is close to a threshold value. In the following description, as an example, the first condition is that the reliability score is within a range of the threshold value ±0.05. Note that by including cases where the reliability score is smaller than the threshold value in the first condition, it is possible to increase the number of annotation results that are visible to the user. However, cases where the reliability score is smaller than the threshold value may not be included in the first condition.

[0030] The second condition is that only a part of an object is detected. For example, if a face is detected but the entire person including the face is not detected, the second condition is met.

[0031] If the image data acquired by the communication unit 13 does not include an annotation result that satisfies at least one of the first and second conditions (NO in step S20), the control unit 11 of the server device 1 causes the display unit 14 to normally display an image based on the image data acquired by the communication unit 13 (step S30). The normal display will be described later. When step S30 ends, the process proceeds to step S60.

[0032] On the other hand, if the image data acquired by the communication unit 13 includes annotation results that satisfy at least one of the first and second conditions (YES in step S20), the control unit 11 of the server device 1 assigns a caution flag to the annotation results that satisfy at least one of the first and second conditions (step S40). By introducing the caution flag as in this embodiment, the display process in step S50 can be performed smoothly.

[0033] In the following description, it is assumed that data of image P2 shown in Fig. 3 is acquired in step S10. It is also assumed that the reliability score of the annotation result corresponding to bounding box BB1 of image P2 shown in Fig. 3 is 0.4. It is also assumed that the reliability score of the annotation result corresponding to bounding box BB2 of image P2 shown in Fig. 3 is 0.26.

[0034] The annotation result corresponding to the bounding box BB1 of image P2 shown in Fig. 3 satisfies the second condition. Furthermore, the annotation result corresponding to the bounding box BB2 of image P2 shown in Fig. 3 satisfies the first condition. Therefore, a caution flag is assigned to each of the annotation result corresponding to the bounding box BB1 of image P2 shown in Fig. 3 and the annotation result corresponding to the bounding box BB2 of image P2 shown in Fig. 3.

[0035] In step S50 following step S40, the control unit 11 of the server device 1 causes the annotation results that satisfy at least one of the first condition and the second condition to be prominently displayed on the display unit 14. In other words, the control unit 11 of the server device 1 causes the annotation results to which a caution flag has been assigned to be prominently displayed on the display unit 14.

[0036] An example of the display executed in step S50 can be seen in Figure 4. In the display example in Figure 4, the internal areas of the bounding boxes BB1 and BB2 corresponding to the annotation results to which caution flags have been assigned are highlighted, thereby making the annotation results to which caution flags have been assigned stand out.

[0037] An example of the display executed in step S50 is shown in Figure 5. In the display example in Figure 5, the bounding boxes BB1 and BB2 corresponding to the annotation results to which caution flags have been assigned are made thicker, thereby making the annotation results to which caution flags have been assigned stand out.

[0038] Alternatively, the bounding box corresponding to the annotation result to which the caution flag has been assigned may be displayed in a blinking manner. Alternatively, the portion other than the bounding box corresponding to the annotation result to which the caution flag has been assigned may be made less noticeable by, for example, lowering the brightness, so that the annotation result to which the caution flag has been assigned stands out relatively.

[0039] Naturally, annotation results to which caution flags have been added may be prominently displayed using methods other than the above-described example. In the following description, it is assumed that image P2 shown in FIG. 4 is displayed in step S50.

[0040] In step S60 following step S50, the control unit 11 of the server device 1 corrects the annotation result based on the correction instruction information input to the input unit 15. This makes it possible to manually correct the annotation result.

[0041] Fig. 6 is a diagram showing an example of image P2 after correction. In image P2 shown in Fig. 6, as a result of the correction of the annotation result, person H1 is entirely surrounded by bounding box BB3, and person H2 is entirely surrounded by bounding box BB2.

[0042] FIG. 7 is a diagram showing an example of a normal display. Image P3 shown in FIG. 7 is an image showing person H3. Person H3 is facing forward. Image P3 shown in FIG. 7 includes bounding box BB4, which is an annotation result in which face F2 of person H3 is detected, and bounding box BB5, which is an annotation result in which person H3 is detected. Also, assume that the reliability scores of the annotation results corresponding to bounding boxes BB4 and BB5 of image P3 shown in FIG. 7 are 0.5. Unlike image P2 shown in FIGS. 4 and 5, image P3 shown in FIG. 7 does not have any prominently displayed parts.

[0043] In step S60, the user can check the prominently displayed portion with a high level of concentration, thereby reducing annotation errors.

[0044] In step S70 following step S60, if a caution flag has been assigned to the annotation result, the control unit 11 of the server device 1 deletes the caution flag, thereby preventing unnecessary meta-information from remaining in the training data.

[0045] In step S80 following step S70, the control unit 11 of the server device 1 causes the recording unit 12 to record the corrected image data obtained by executing the processes up to this point as training data. The training data can be transmitted to an external device of the server device 1 using the communication unit 13. When the process of step S80 ends, the operation of the flowchart shown in FIG. 2 ends.

[0046] <3. Modifications> The above-described embodiments should be considered to be illustrative in all respects and not restrictive. The technical scope of the present invention is indicated by the claims, not by the description of the above-described embodiments, and should be understood to include all modifications that fall within the meaning and scope of the claims.

[0047] <Variation 1> In the above-described embodiment, the trained artificial intelligence model 2 inputs image data to which no annotation results have been added. However, the present invention is not limited to this. The trained artificial intelligence model 2 may input image data to which annotation results have been added in advance. Examples of image data to which annotation results have been added in advance include images captured using face recognition (face detection) on a smartphone.

[0048] <Variation 2> If the annotation result satisfies the second condition, the display unit 14 may display all of the object's detection candidates when prominently displaying the annotation result. For example, one way to display all of the object's detection candidates is to also display annotation results whose reliability scores are significantly lower than a threshold. This makes it easier to manually correct annotation results that satisfy the second condition.

[0049] <Variation 3> In the above-described embodiment, a human face and a person (the whole body of a person) are detected, but only a human face may be detected. When only a human face is detected, the processing related to the second condition may be omitted. For example, when detecting a driver from an image captured by a camera capturing an image of the interior of a vehicle, only the upper half of the driver's body is captured in the image, and there is no point in detecting the person (the whole body of a person), so it is desirable not to detect the person (the whole body of a person).

[0050] <Variation 4> In the above-described embodiment, a single trained artificial intelligence model 2 was used. However, multiple types of trained artificial intelligence models 2 with different learning contents and performance may be used to create image data to which annotation results and reliability scores for the annotation results are assigned. When multiple types of trained artificial intelligence models 2 are used to create image data to which annotation results and reliability scores for the annotation results are assigned, each annotation result may be determined by majority voting among the multiple types of trained artificial intelligence models 2. Furthermore, when there is an even number of types, each annotation result may be weighted, taking into consideration the possibility that each annotation result cannot be determined by majority voting. The more reliable a trained artificial intelligence model is, the greater the weighting. For example, if there are two types of trained artificial intelligence models 2 and one trained artificial intelligence model (model A) is more reliable than the other trained artificial intelligence model (model B), the weight of the output of model A may be set to a first value, and the weight of the output of model B may be set to a second value (< first value), and the annotation result of the larger of the output of model A after weighting and the output of model B after weighting may be used.

[0051] By using multiple types of trained artificial intelligence models 2 in this way, the accuracy of annotation results for image data in its pre-correction state is improved, and the burden of manual correction is reduced. [Explanation of symbols]

[0052] 1. Server device 2. Trained AI model 11 Control section 12 Recording section 13 Communications Department 14 Display section 15 Input section BB1~BB4 bounding boxes F1, F2 face H1~H3 people P1~P3 images

Claims

1. An annotation support device comprising a control unit, The control unit Obtaining image data to which an annotation result and a reliability score for the annotation result have been assigned; When the annotation result satisfies at least one of a first condition that the reliability score is close to a threshold and a second condition that only a part of the object has been detected, the annotation result is prominently displayed on a display unit; acquiring correction instruction information regarding correction of the annotation result; The threshold is used by the trained artificial intelligence model to determine whether or not to adopt the detected candidate as the annotation result, annotation support device.

2. The control unit The annotation support device according to claim 1 , further comprising: a warning flag attached to the annotation result when the annotation result satisfies at least one of the first condition and the second condition.

3. The control unit correcting the annotation result based on the correction instruction information; The annotation support device according to claim 2 , wherein the attention-requiring flag is deleted after correction is completed.

4. The control unit The annotation according to any one of claims 1 to 3, wherein, when the annotation result satisfies the second condition, the control unit displays all of the detection candidates of the object when causing the display unit to prominently display the annotation result. Support equipment.

5. an image data acquisition step in which an information processing device acquires image data to which an annotation result and a reliability score for the annotation result have been assigned; a display step of prominently displaying the annotation result on a display device when the annotation result satisfies at least one of a first condition that the confidence score is close to a threshold and a second condition that only a portion of an object has been detected; a correction instruction information acquisition step of acquiring correction instruction information regarding correction of the annotation result; Equipped with An annotation assistance method, wherein the threshold is used to determine whether or not the trained artificial intelligence model will adopt the detection candidate as the annotation result.

Citation Information

Patent Citations

  • Image processing apparatus and method, and program

    JP2011205387A

  • Object detector

    JP2019078554A

  • Annotation device and method

    JP2021089491A

  • Information processing apparatus, information processing method, and program

    JP2021099582A

  • Inspection system

    WO2020178913A1