Information processing method, information processing device, and information processing program

JP2026126473APending Publication Date: 2026-08-05PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
Filing Date
2023-06-06
Publication Date
2026-08-05

AI Technical Summary

Benefits of technology

【0008】 本開示によれば、アノテータのスキルを効率良く向上させることができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026126473000001_ABST
    Figure 2026126473000001_ABST
Patent Text Reader

Abstract

Efficiently improve the annotator's skills. [Solution] The information processing device acquires the annotation work history by the annotator, the annotation is applied to the original training data to generate training data for a machine learning model, calculates an evaluation value to evaluate the annotation based on the work history, generates a training menu according to the evaluation value to improve annotation skills, and outputs the training menu to the annotator's display.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a technique for annotating images.

Background Art

[0002] Patent Document 1 discloses a technique for receiving an annotation for creating learning data and evaluating the annotation based on the contribution of the received annotation - added learning data to learning.

[0003] Patent Document 2 discloses a technique for retaining the history of annotations input by a user for video content and calculating the contribution points of the user to the annotation by weighting based on the type and frequency of the annotation.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] Thus, the prior art only discloses evaluating annotations and does not consider improving the skills of annotators. Therefore, the prior art cannot efficiently improve the skills of annotators.

[0006] This disclosure is made to solve such problems and aims to provide a technique for efficiently improving the skills of annotators.

Means for Solving the Problems

[0007] An information processing method in one aspect of the present disclosure is an information processing method in a computer, comprising: acquiring the work history of annotation by an annotator; the annotation being applied to raw training data in order to generate training data for a machine learning model; calculating an evaluation value for evaluating the annotation based on the work history; generating a training menu corresponding to the evaluation value, for improving the annotation skills; and outputting the training menu to the display of the annotator. [Effects of the Invention]

[0008] According to this disclosure, annotators' skills can be improved efficiently. [Brief explanation of the drawing]

[0009] [Figure 1] This is a block diagram showing an example of the overall configuration of an information processing system in an embodiment. [Figure 2] A flowchart showing an example of processing in the information processing device according to the embodiment. [Figure 3] This is a flowchart that continues from Figure 2. [Figure 4] This figure shows the display screen for the first example of the accuracy training menu. [Figure 5] This figure shows the display screen for the second example of the accuracy training menu. [Figure 6] This figure shows the display screen for the third example of the accuracy training menu. [Figure 7] This figure shows the display screen for the fourth example of the accuracy training menu. [Figure 8] This figure shows the display screen for the fifth example of the accuracy training menu. [Figure 9] This is a diagram showing the display screen for the sixth example of the accuracy training menu. [Figure 10] This figure shows the display screen for the first example of a time efficiency training menu. [Figure 11]It is a diagram showing a display screen of a second example of a time efficiency training menu. [Figure 12] It is a diagram showing a display screen of a third example of a time efficiency training menu. [Figure 13] It is a diagram explaining an accuracy annotation test. [Figure 14] It is a diagram explaining a time efficiency annotation test. [Figure 15] It is a diagram explaining how a threshold value is set. [Figure 16] It is an explanatory diagram of a fourth example of an accuracy evaluation value. [Figure 17] It is an explanatory diagram of a fifth example of an accuracy evaluation value. [Figure 18] It is an explanatory diagram of a seventh example of an accuracy evaluation value.

Mode for Carrying Out the Invention

[0010] (Findings Leading to an Aspect of the Present Disclosure) Rather than directly using a highly versatile machine learning model trained with various data, there is a need to fine-tune the machine learning model using in-house data and obtain a high-performance machine learning model specialized for that site. For example, this is the case when object recognition is performed using a machine learning model on a manufacturing line in a factory.

[0011] When performing fine-tuning, it is necessary to prepare a large amount of training data with annotations indicating the correct answers added to the in-house data collected at the site. Such training data is generated by annotators adding annotations indicating the correct answers to a large amount of in-house data. To generate a high-performance machine learning model, it is necessary to prepare a large amount of training data with accurately added annotations. For this purpose, it is essential to train annotators who can quickly add highly reliable annotations.

[0012] In the prior art, there has only been an evaluation of annotations so as to increase the motivation for annotations or an evaluation of annotations according to the contribution degree of the annotations. Therefore, the prior art cannot efficiently improve the skills of annotators.

[0013] This disclosure has been made to solve such problems.

[0014] (1) An information processing method in one aspect of this disclosure is an information processing method in a computer, which acquires the work history of annotations by an annotator. The annotation is given to original learning data in order to generate learning data of a machine learning model, calculates an evaluation value for evaluating the annotation based on the work history, generates a training menu corresponding to the evaluation value, which is a training menu for improving the skill of the annotation, and outputs the training menu to the display of the annotator.

[0015] According to this configuration, a display screen of a training menu corresponding to the evaluation value for the annotation is displayed on the display of the annot. Therefore, the annotator can overcome the annotations that the annotator is not good at through the training menu. As a result, the skills of the annotator can be efficiently improved. In addition, it is possible to prevent a large amount of learning data to which an immature annotator has given undesirable annotations from being mass-produced.

[0016] [[ID=]](2) In the information processing method described in (1) above, the evaluation value is calculated for each of a plurality of evaluation criteria, and the training menu may include training menu contents corresponding to the evaluation criteria for which the calculated evaluation value does not satisfy the reference conditions.

[0017] According to this configuration, since the training menu contents corresponding to the evaluation criteria for which the evaluation value does not satisfy the reference conditions are presented to the annotator, the skill of the annotation for the evaluation criteria that the annotator is not good at can be efficiently improved.

[0018] (3) In the information processing method described in (2) above, the plurality of evaluation criteria may include an accuracy evaluation criterion for evaluating the accuracy of the annotation and a time efficiency evaluation criterion for evaluating the time spent on the annotation work.

[0019] This configuration allows for efficient improvement of annotator skills in terms of accuracy and time efficiency.

[0020] (4) In the information processing method described in (2) or (3) above, the annotation is classified into a plurality of annotation items, the evaluation value is calculated for each of the plurality of evaluation criteria for each annotation item, and the training menu may include training menu content corresponding to the evaluation criteria and annotation items for which the calculated evaluation value does not meet the criteria conditions.

[0021] With this configuration, annotators are presented with training menus tailored to evaluation criteria and annotation items whose evaluation values ​​do not meet the standard conditions. This allows annotators to efficiently improve their annotation skills for evaluation criteria and annotation items they find difficult, or for evaluation criteria and annotation items for which they are required to perform the work.

[0022] (5) In the information processing method described in (4) above, the plurality of annotation items may be at least one of the following types: one that assigns a class label to an object included in an image, one that assigns a bounding box to the object, and one that assigns a time segment label to a time series image.

[0023] This configuration allows for efficient improvement of the annotator's skills in at least one of the following: assigning class labels, assigning bounding boxes, and assigning time segment labels.

[0024] (6) In the information processing method described in any one of (1) to (5) above, the training menu may consist of a display screen that includes at least the bad annotation sample, which is a good annotation sample showing the annotation performed by a good annotator or a correct annotation sample to which the correct annotation has been assigned, and a bad annotation sample showing the annotation performed by a bad annotator.

[0025] This configuration allows the annotator to specifically understand at least the bad annotations, rather than just the good ones.

[0026] (7) In the information processing method described in (6) above, the good annotation sample and the bad annotation sample may be images showing the results of the annotation by the good annotator and the bad annotator.

[0027] This configuration allows the annotator to specifically understand at least the results of bad annotation, rather than just good annotation results.

[0028] (8) In the information processing method described in (6) above, the good annotation sample and the bad annotation sample may be a video showing the process by which the annotation is applied by the good annotator and the bad annotator.

[0029] This configuration allows the annotator to specifically understand the steps required to apply annotations accurately.

[0030] (9) In the information processing method described in any one of (1) to (8) above, the training menu may include an annotation test having test content corresponding to the evaluation value, wherein the annotation test is assigned to an annotator to perform the task of adding annotations to a test image.

[0031] This configuration ensures that annotators can significantly improve their annotation skills, which are often a weak point for them.

[0032] (10) In the information processing method described in (9) above, the annotation test may include calculating a test score for the annotation test from the results of the annotation test, determining whether the annotation test passed or failed from the test score, and resuming the annotation work if it is determined that the annotator has passed the annotation test.

[0033] With this configuration, annotation work is performed by an annotator that has overcome its weaknesses in annotation, resulting in training data with accurate annotations.

[0034] (11) In the information processing method described in (10) above, the training menu may include outputting an annotation sample to the display if the annotator is determined to have failed the annotation test.

[0035] With this configuration, annotation samples are presented to annotators who have difficulty with certain annotations, ensuring that they can reliably understand the annotations they struggle with.

[0036] (12) In the information processing method described in any one of (1) to (11) above, the method may further include displaying a threshold setting screen for evaluating the evaluation value on the display.

[0037] This configuration allows for arbitrary setting of thresholds, enabling the training of annotators capable of providing annotations that meet the requirements of annotation requesters. Furthermore, by appropriately setting thresholds, it is possible to prevent the mass production of training data with low-quality annotations. Additionally, by appropriately setting thresholds, it is possible to prevent the training of annotators that provide excessively accurate annotations over a considerable amount of time.

[0038] (13) In the information processing method described in (12) above, the threshold includes a first threshold for evaluating the accuracy of the annotation, the setting screen includes a first setting screen for setting the first threshold, and the first setting screen may include an adjustment unit for adjusting the first threshold and a display field for displaying annotation samples corresponding to the first threshold.

[0039] This configuration allows users to adjust the threshold while viewing annotation samples corresponding to the first threshold, making it easy to set the first threshold to meet the required annotation criteria.

[0040] (14) In the information processing method described in (12) or (13) above, the threshold includes a second threshold for evaluating the accuracy of the annotation, the setting screen includes a second setting screen for setting the second threshold, and the second setting screen may include an adjustment unit for adjusting the second threshold and a display field for displaying the annotation work time per sample according to the second threshold.

[0041] This configuration makes it easier to set the second threshold because you can adjust the threshold while checking the processing time per sample according to the second threshold.

[0042] (15) In the information processing method described in any one of (12) to (14) above, the annotation includes a plurality of annotation items, and the display of the setting screen may include receiving an instruction to select one annotation item from the plurality of annotation items, and displaying a setting screen for setting the threshold for the one annotation item.

[0043] This configuration makes it possible to set thresholds for each annotation item.

[0044] (16) In the information processing method described in any one of (1) to (15) above, the calculation of the evaluation value may include dividing the annotations that the annotator has applied to a person included in the image into an upper body region corresponding to the upper body of the person and a lower body region corresponding to the lower body of the person; calculating the Interaction Over Union (IOU) of the upper body region and the IOU of the lower body region; and calculating the evaluation value based on a weighted average of the IOU of the upper body and the IOU of the lower body such that the weight value of the IOU of the upper body region is higher than the weight value of the IOU of the lower body region.

[0045] Since the upper body contains many of a person's features, such as the face, accuracy in annotations of people tends to be prioritized for the upper body more than for the lower body. With this configuration, the evaluation value is calculated based on a weighted average of both IOUs, so that the IOU for the upper body has a higher weight than the IOU for the lower body. Therefore, annotations that accurately annotate the upper body can be given a higher evaluation.

[0046] (17) In the information processing method described in any one of (1) to (16) above, the calculation of the evaluation value may include calculating the correct inclusion ratio, which is the ratio of annotations added by the annotator to the correct annotations, and calculating the evaluation value based on the IOU (Interaction Over Union) of the annotations and the correct inclusion ratio.

[0047] For anomaly detection training data, there is a tendency to require annotations that include all anomalous parts. This is because machine learning models trained on training data where only a portion of the anomalous parts are annotated are more likely to miss anomalies. In this configuration, annotations are evaluated based on the correct coverage rate, making it possible to obtain training data that trains a machine learning model to perform anomaly detection accurately.

[0048] (18) In the information processing method described in any one of (1) to (17) above, the evaluation value is calculated based on the annotation results that the annotator has applied to the evaluation data included in the dataset on which the annotator performs the annotation, and the evaluation data may be the original training data having the correct annotation for display.

[0049] This configuration allows annotators to be evaluated through their actual annotation work, without requiring them to perform dedicated annotation tasks to calculate evaluation values.

[0050] (19) An information processing device in another aspect of the present disclosure is an information processing device including a processor, the processor which acquires an annotation work history by an annotator, the annotations are applied to raw training data to generate training data for a machine learning model, calculates an evaluation value to evaluate the annotations based on the work history, generates a training menu corresponding to the evaluation value to improve the annotation skills, and outputs the training menu to the display of the annotator.

[0051] This configuration provides an information processing device that efficiently enhances the annotator's skills.

[0052] (20) An information processing program in yet another aspect of the present disclosure causes a computer to perform the following processes: acquire an annotation work history by an annotator; the annotations are applied to raw training data to generate training data for a machine learning model; calculate an evaluation value for evaluating the annotations based on the work history; generate a training menu corresponding to the evaluation value for improving the annotation skills; and output the training menu to the annotator's display.

[0053] This configuration allows us to provide an information processing program that efficiently enhances the annotator's skills.

[0054] This disclosure can also be implemented as an information processing system operated by such an information processing program. Furthermore, it goes without saying that such a computer program can be distributed via computer-readable, non-temporary recording media such as CD-ROMs or via communication networks such as the Internet.

[0055] The embodiments described below are all specific examples of this disclosure. The numerical values, shapes, components, steps, and order of steps shown in the following embodiments are examples only and are not intended to limit this disclosure. Furthermore, among the components in the following embodiments, those not described in the independent claim representing the highest-level concept will be described as optional components. In addition, the contents of each embodiment can be combined.

[0056] (Embodiment) Figure 1 is a block diagram showing an example of the overall configuration of an information processing system in an embodiment. The information processing system includes an information processing device 1 and a terminal device 2. The information processing device 1 and the terminal device 2 are connected to each other so as to be able to communicate with each other via a network NT. An example of a network NT is a wide-area communication network including the Internet and a mobile phone communication network.

[0057] Information processing device 1 is composed of a computer, such as a cloud server. However, this is just one example, and information processing device 1 may also be composed of an edge computer.

[0058] Terminal device 2 consists of a portable computer such as a tablet or smartphone, or a stationary computer. Terminal device 2 includes a central processing unit (CPU), memory, display, control unit, and communication circuit. Terminal device 2 is the terminal used by the annotator. In Figure 1, one terminal device 2 is shown for illustrative purposes, but this is just an example, and there may be multiple terminal devices 2.

[0059] The information processing device 1 includes a processor 10, a memory 20, and a communication unit 30. The processor 10 is composed of a central processing unit (CPU) and includes an acquisition unit 11, an evaluation unit 12, a generation unit 13, an output unit 14, and a setting unit 15. The acquisition unit 11 to the setting unit 15 are realized, for example, by the processor 10 executing an information processing program. However, this is just an example, and the acquisition unit 11 to the setting unit 15 may be composed of dedicated hardware circuits. Furthermore, the acquisition unit 11 to the setting unit 15 may be distributed across multiple computers, or some of the functions may be implemented in a terminal device 2.

[0060] The acquisition unit 11 acquires the annotation work history performed by the annotator. The work history includes the training data to which annotations have been added, the time required for annotation of each training data, and the annotator identifier, the training data identifier, and the identifier of the dataset to which the training data belongs. The original training data to which the annotator performs annotations includes evaluation data for which the correct annotation has already been determined. The evaluation of the annotations is performed based on the annotations that the annotator has added to this evaluation data. Note that the evaluation data presented to the annotator does not have the correct annotation added, so the annotator cannot recognize the correct annotation.

[0061] An annotator is a person who adds annotations to the original training data. In this embodiment, the annotator adds annotations to the original training data using the terminal device 2. Therefore, the acquisition unit 11 acquires the annotation results transmitted from the terminal device 2 as work history using the communication unit 30, and stores the acquired work history in the work history database of the memory 20. Original training data is data before annotations are added. For example, still images and moving images containing the object to be recognized are used as original training data. The annotation results include the annotated training data, the work time required for annotation of each training data, the annotator's identifier, the training data identifier, and the identifier of the dataset to which the training data belongs.

[0062] Annotations include class labels and bounding boxes attached to objects in an image. Annotations can also be time segment labels indicating time segments that meet specific conditions within a video. The terms used to distinguish these annotations are called annotation items.

[0063] An object is an object that a machine learning model recognizes. A time segment that satisfies a predetermined condition refers to a period within the entire duration of a video where a particular object is in a specific state. For example, a period in which a person is walking or a period in which a fire is occurring would be considered a time segment that satisfies a predetermined condition. A class label is textual data that indicates the type of object, such as dog or cat. A bounding box is a frame added to an object in an image to indicate its location. The shape of the frame is, for example, a rectangle.

[0064] Any machine learning model can be used as long as it is a supervised learning model. For example, machine learning models include deep neural networks, convolutional neural networks, random forests, decision trees, and support vector machines. The machine learning model may be pre-trained using a general-purpose dataset. In this case, the machine learning model can be further trained using training data created by an annotator trained through the training menu of this disclosure.

[0065] The evaluation unit 12 calculates an evaluation value for evaluating the annotation based on the work history. The evaluation value is calculated for each of the multiple evaluation criteria. The multiple evaluation criteria include an accuracy evaluation criterion for evaluating the accuracy of the annotation and a time efficiency evaluation criterion for evaluating the time spent on the annotation work. Hereinafter, the evaluation value for the accuracy evaluation criterion will be called the accuracy evaluation value, and the evaluation value for the time efficiency evaluation criterion will be called the time evaluation value.

[0066] The accuracy score is represented by the error margin defined for each annotation item. The smaller the error margin, the higher the accuracy score. Therefore, a smaller accuracy score indicates a higher rating.

[0067] If there are multiple objects to be annotated for a single training dataset, the error may include an object-level error calculated for each object and a data-level error calculated for the entire training dataset. The data-level error is, for example, expressed as the average of the object-level errors for a single training dataset.

[0068] The error indicates the discrepancy between the correct annotation and the annotation applied by the annotator. For example, if the annotation is a class label or time segment label, the error is "0" if the applied class label is correct and "1" if it is incorrect. For example, if the annotation is a bounding box, the error is expressed as 1-IOU (Interaction over Union). IOU is the ratio of the area represented by the logical AND of the correct bounding box and the applied bounding box to the area represented by the logical OR of the correct bounding box and the applied bounding box. Details of the error will be described later.

[0069] The time evaluation score represents the time required to annotate each training data point. The time efficiency is evaluated based on this time; a shorter time is better. Therefore, a smaller time evaluation score indicates a higher evaluation.

[0070] However, this is just one example, and multiple evaluation criteria may include criteria for evaluating quality variability and criteria for evaluating monetary cost. In this case, the evaluation unit 12 should decrease the evaluation value and increase the annotator's evaluation as quality variability decreases. Also, the evaluation unit 12 should decrease the evaluation value and lower the annotator's evaluation as monetary cost decreases.

[0071] The evaluation unit 12 may calculate accuracy evaluation values ​​and time evaluation values ​​for each annotation item. Furthermore, the evaluation unit 12 may calculate accuracy evaluation values ​​and time evaluation values ​​for each combination of annotation item and dataset. A dataset refers to a group of training data grouped by type. For example, a group of images taken at a certain site constitutes one dataset. Examples of sites include factories and construction sites. Therefore, datasets are classified as a dataset for factory α, a dataset for factory β, a dataset for construction site γ, and a dataset for construction site δ.

[0072] The generation unit 13 generates a training menu corresponding to the evaluation value, which is a training menu for improving annotation skills. The training menu may include training menu content corresponding to evaluation criteria where the evaluation value calculated by the evaluation unit 12 does not meet the standard conditions. For example, if the accuracy evaluation value does not meet the standard conditions, a training menu to improve the accuracy of annotation is generated. For example, if the time evaluation value does not meet the standard conditions, a training menu to shorten the time spent on annotation is generated. For example, if both the accuracy evaluation value and the time evaluation value do not meet the standard conditions, training menus for both are generated.

[0073] If an evaluation value is calculated for each annotation item, the generation unit 13 should generate a training menu to improve annotation skills for annotation items whose evaluation values ​​do not meet the standard conditions. If an evaluation value is calculated for each pair of annotation items and datasets, the generation unit 13 should generate a training menu to improve annotation skills for pairs whose evaluation values ​​do not meet the standard conditions.

[0074] A threshold is used as the evaluation criterion. As mentioned above, a smaller evaluation value indicates a higher evaluation of the annotation. Therefore, the generation unit 13 should determine that the criteria are met if the evaluation value is less than the threshold, and that the criteria are not met if the evaluation value is greater than the threshold. As the threshold, the average value of the evaluation values ​​corresponding to the bottom X% of all annotators that performed the annotation can be used. X can be any appropriate value such as 10, 20, or 30. For example, if evaluation values ​​have been calculated for training data D1 to Dn, and the evaluation values ​​corresponding to the bottom X% for each of the training data D1 to Dn are V1 to Vn, then the threshold will be the average value of the evaluation values ​​V1 to Vn.

[0075] The training menu may consist of a display screen that includes at least one poor annotation sample, which is either a good annotation sample showing annotations made by a good annotator or a correct annotation sample with the correct annotation, and a poor annotation sample showing annotations made by a poor annotator. A good annotator may be a skilled annotator or an annotator whose evaluation value falls in the top Y% of all annotators. Y can be any appropriate value, such as 10, 20, or 30. A poor annotator may be an annotator whose evaluation value falls in the bottom X% of all annotators. Hereinafter, good annotation samples and poor annotation samples will be collectively referred to as annotation samples.

[0076] The good annotation samples and bad annotation samples may be images that show the results of annotation by a good annotator and a bad annotator, respectively.

[0077] The good annotation samples and the bad annotation samples may be videos showing the process of annotation being applied by the good annotator and the bad annotator, respectively.

[0078] The training menu may include annotation tests that have test content corresponding to evaluation values, and may also include annotation tests that assign an annotator the task of adding annotations to test images.

[0079] In the annotation test, the generation unit 13 may perform the following processes: calculating the test score of the annotation test from the results of the annotation test; determining whether the annotation test passed or failed from the test score; and resuming the annotation work if it is determined that the annotator has passed the annotation test.

[0080] The output unit 14 outputs the training menu generated by the generation unit 13 to the annotator's display. For example, the output unit 14 generates display data for displaying the training menu screen on the display, and transmits the generated display data to the terminal device 2 using the communication unit 30.

[0081] The setting unit 15 displays the threshold setting screen described above on the annotator's display. In this case, the setting unit 15 only needs to transmit the display data of the setting screen to the terminal device 2 using the communication unit 30.

[0082] The setting unit 15 includes a process for receiving a selection instruction for one annotation item from among multiple annotation items, and a process for displaying a setting screen for setting a threshold for the one annotation item. The setting unit 15 stores the threshold set through the setting screen in the memory 20.

[0083] The settings screen includes a first settings screen for setting a first threshold for evaluating accuracy. The first settings screen includes an adjustment section for adjusting the first threshold and a display area for displaying annotation samples corresponding to the first threshold.

[0084] The settings screen may include a second settings screen for setting a second threshold for evaluating the time evaluation value. The second settings screen includes an adjustment unit for adjusting the second threshold and a display field for displaying the annotation work time per sample according to the second threshold.

[0085] The memory 20 consists of a rewritable, non-volatile storage device such as a solid-state drive or a hard disk drive, and stores the work history database and thresholds. The work history database stores the work history acquired by the acquisition unit 11.

[0086] The communication unit 30 is a communication circuit for connecting the information processing device 1 to the network NT. The communication unit 30 transmits display data for the training menu to the terminal device 2 and displays data for the settings screen to the terminal device 2. The communication unit 30 receives the annotation results transmitted from the terminal device 2.

[0087] Figure 2 is a flowchart showing an example of processing by the information processing device 1 in the embodiment. Figure 3 is a flowchart continuing from Figure 2. In step S1, the evaluation unit 12 determines whether the annotator to be evaluated (hereinafter referred to as the target annotator) is a new annotator. A new annotator is an annotator whose annotation is being evaluated for the first time by the information processing device 1. For example, the evaluation unit 12 can determine whether it is a new annotator based on the annotator information transmitted from the terminal device 2. The memory 20 stores annotator information of annotators whose annotation has already been evaluated. Therefore, if the annotator information transmitted from the terminal device 2 is not stored in the memory 20 as evaluated annotator information, the evaluation unit 12 determines that the target annotator is a new annotator, and if the transmitted annotator information is stored in the memory 20 as evaluated annotator information, the evaluation unit 12 determines that the target annotator is not a new annotator.

[0088] If the target annotator is determined to be a new annotator (YES in step S1), the process proceeds to step S2. If the target annotator is determined not to be a new annotator (NO in step S1), the process proceeds to step S3.

[0089] Next, in step S2, the generation unit 13 generates a training menu for the new annotator, and the output unit 14 transmits the display data of the generated training menu to the terminal device 2 via the communication unit 30. As a result, the display of the terminal device 2 displays the display screen for the training menu for the new annotator. The training menu for the new annotator includes information to help the new annotator master the basics of annotation. For example, the training menu for the new annotator may include the content of training menus that have been presented frequently, determined based on the viewing history of training menus performed by existing annotators. This allows for efficient training of the new annotator.

[0090] Next, in step S3, the acquisition unit 11 transmits an annotation work request to the terminal device 2 using the communication unit 30. The work request includes a dataset containing the original training data to be annotated. Hereinafter, the dataset included in this work request will be referred to as the target dataset. The work request also includes work instructions for multiple annotation items. The original training data included in the dataset is displayed sequentially on the display of the terminal device 2, and the target annotator adds annotations to each piece of original training data. The results of the annotations added by the target annotator to each piece of original training data are transmitted from the terminal device 2 to the information processing device 1.

[0091] Next, in step S4, the acquisition unit 11 acquires the annotation results transmitted from the terminal device 2 as work history. The acquired work history is stored in the work history database.

[0092] The following processes, consisting of steps S5 and S6, and steps S7 and S8, are performed in parallel.

[0093] In step S5, the evaluation unit 12 reads the work history of the target annotator for the target dataset from the work history database and calculates the average accuracy evaluation value for each annotation item from the read work history. For example, if the annotation items include the assignment of class labels and the assignment of bounding boxes, the average accuracy evaluation value is calculated for both the assignment of class labels and the assignment of bounding boxes.

[0094] Next, in step S6, the generation unit 13 extracts annotation items that do not meet the criteria. In this case, the generation unit 13 only needs to extract annotation items whose average accuracy evaluation value is greater than the first threshold as annotation items that do not meet the criteria.

[0095] In step S7, the evaluation unit 12 reads the work history of the target annotator for the target dataset from the work history database and calculates the average time evaluation value for each annotation item from the read work history. For example, if the annotation items include the assignment of class labels and the assignment of bounding boxes, the average time evaluation value is calculated for both the assignment of class labels and the assignment of bounding boxes.

[0096] Next, in step S8, the generation unit 13 extracts annotation items that do not meet the criteria. In this case, the generation unit 13 only needs to extract annotation items whose average time evaluation value is greater than the second threshold as annotation items that do not meet the criteria.

[0097] Next, in step S9, the generation unit 13 determines from the extraction results of step S6 or step S8 whether or not there are any annotation items that do not meet the criteria. If it is determined that there are no annotation items that do not meet the criteria (NO in step S9), the process returns to step S3. In this case, a request for annotation work on the next target dataset is sent to the terminal device 2, and the annotator performs the annotation work on the next target dataset. Note that if NO is obtained in step S9, the process may be terminated.

[0098] If it is determined that there are annotation items that meet the criteria (YES in step S9), the process proceeds to step S10.

[0099] Next, in step S10, the generation unit 13 generates a training menu to improve the annotator's skills for annotation items that do not meet the standard conditions.

[0100] Next, in step S11, the output unit 14 transmits the generated growth menu display data to the terminal device 2 using the communication unit 30.

[0101] Next, in step S12, the output unit 14 sends a request to execute the annotation test to the terminal device 2 using the communication unit 30. As a result, the target annotator executes the annotation test using the terminal device 2.

[0102] Next, in step S13, the generation unit 13 obtains the annotation test answers from the target annotator using the communication unit 30.

[0103] Next, in step S14, the generation unit 13 calculates a test score from the annotation test answers and determines whether the target annotator passed the annotation test based on the calculated test score. If the target annotator passes the annotation test, the process proceeds to step S15. If the target annotator fails the annotation test, the process returns to step S10. In this case, the target annotator will review the same training menu again.

[0104] Next, in step S15, the generation unit 13 resets the evaluation values ​​(accuracy evaluation value or time evaluation value) for annotation items that were determined not to meet the criteria in step S9. When step S15 is completed, the process returns to step S3. In this case, the target annotator again adds annotations to the target dataset.

[0105] Figure 4 shows the display screen G11 of the first example of the accuracy training menu. The accuracy training menu is a training menu that is generated when it is determined that the accuracy evaluation value of the target annotator does not meet the standard conditions.

[0106] Display screen G11 includes an NG sample field 401, a correct answer field 402, and a test button 403. The NG sample field 401 displays the NG annotation samples 410 for each of data A, B, and C. The correct answer field 402 displays the correct annotation samples 420 for each of data A, B, and C. Data A, B, and C are three evaluation data sets randomly selected from the evaluation data included in the target dataset to which annotations related to the target annotation items have been assigned. The target annotation items are annotation items that the target annotator has determined not to meet the criteria. The three NG annotation samples 410 displayed in the NG sample field 401 are NG annotation samples randomly selected from the NG annotation samples for each of data A, B, and C.

[0107] The correct annotation sample 420 displayed in the correct answer field 402 is an annotation sample in which the correct annotation has been applied to data A, B, and C.

[0108] The defective annotation sample 410 displayed in the NG sample column 401 may include defective annotation samples from the target annotator.

[0109] Test button 403 is pressed when the target annotator performs an annotation test. The target annotator compares the faulty annotation sample 410 and the correct annotation sample 420, checks for any annotation issues, and then presses test button 403.

[0110] If the source training data included in the target dataset consists of video images, then the incorrect annotation sample 410 and the correct annotation sample 420 will consist of video images. In this case, the target annotator can press the test button 403 only after it has finished playing all the video images displayed in the NG sample column 401 and the correct answer column 402. The content of the video images will show what time segment labels were assigned to which period.

[0111] By viewing display screen G11, the annotator can compare the faulty annotation sample 410 with the correct annotation sample 420 to identify the reason why the criteria were not met. This improves the accuracy of the annotator's annotations. Furthermore, since display screen G11 also displays faulty annotation samples 410 from other annotators, the annotator can proactively check for potential future annotation errors. This helps prevent future annotation errors.

[0112] Display screen G11 shows annotation samples for three data points A to C, but this is just an example; annotation samples for four or more data points may be displayed, or for two or fewer data points. Furthermore, the three data points A to C may be any three original training data points randomly selected from the original training data to which the target annotation has been applied, regardless of the dataset.

[0113] Figure 5 shows the display screen G12 of the second example of the accuracy training menu. Display screen G12 differs from display screen G11 in that for each incorrect annotation sample 410, the target annotator is prompted to input the annotation again, and after the re-annotation input, the correct annotation sample 420 is displayed. Components identical to those in display screen G11 are denoted by the same reference numerals in display screen G12 and their explanations are omitted.

[0114] The target annotator inputs the operation of pressing the re-annotation button 501 corresponding to one data from data A to C. The generation unit 13 then displays the re-annotation screen G13 for data 1 on the display. The re-annotation screen G13 includes a work area 502. The work area 502 displays the original training data corresponding to data 1. In the example in Figure 5, the original training data for data C is displayed.

[0115] The target annotator inputs re-annotation to the original training data and inputs the operation of pressing the judgment button 503. Then, the generation unit 13 redisplays the display screen G12 on the display. The redisplayed display screen G12 displays the correct annotation sample 420 corresponding to the data from step 1 in the correct answer column 402, and displays the pass / fail judgment result for the re-annotation in the NG sample column 401. The generation unit 13 determines that the re-annotation is a pass if the accuracy evaluation value of the re-annotation is smaller than the threshold, and a fail if the accuracy evaluation value of the re-annotation is larger than the threshold. If the threshold is, for example, the accuracy evaluation value corresponding to the bottom X%, then if the accuracy evaluation value of the re-annotation is smaller than the accuracy evaluation value corresponding to the bottom X%, it is determined to be a pass.

[0116] In the example in Figure 5, the re-annotation of data A passed, so "OK" is displayed as the result for data A. On the other hand, the re-annotation of data B did not pass, so "NG" is displayed as the result for data B. Therefore, the target annotator needs to re-annotate data B until it passes.

[0117] The annotator displays the re-annotation screen G13 for all of data A through C and enters the re-annotations. If the re-annotation of all data A through C is successful, the annotator can proceed to the annotation test by pressing the test button 403.

[0118] Display screen G12 offers the following benefits in addition to those of display screen G11: The target annotator can learn good annotation techniques by actually inputting re-annotations and having their work judged as pass or fail. Also, since the correct annotation sample 420 is not displayed until the first re-annotation is performed, the target annotator can easily recognize the difference between their own annotations and the correct annotations.

[0119] Figure 6 shows the display screen G14 of the third example of the accuracy training menu. Display screen G14 differs from display screen G12 in that it displays the score of the re-annotation, rather than whether the re-annotation passed or failed. The NG sample column 401 of display screen G14 includes a score display column 601 that displays the re-annotation score for each of the data A to C.

[0120] The score is a numerical representation of the accuracy evaluation score on a 5-point scale from 1 to 5. For example, the score is determined based on the accuracy evaluation score ranking: 5 points for the top 80% of all annotators, 4 points for the top 80% to 60% of all annotators, and so on. The generation unit 13 should, for example, determine that the re-annotation is successful if the score is 3 points or higher. The score display field 601 displays the passing criteria, such as "3 points or higher is a pass."

[0121] Display screen G14 offers the following benefits in addition to those of display screens G11 and G12. Display screen G14 displays the evaluation of re-annotation on a 5-point scale, and this score is determined based on the ranking, allowing for a relative evaluation of re-annotation. As a result, the annotator can more easily understand the quality of their own re-annotation.

[0122] Figure 7 shows the display screen G21 of the fourth example of the accuracy training menu. Display screen G21 differs from display screens G11, G12, and G14 in that it displays good annotation samples 430 and poor annotation samples 410 for each of the data A to C. Also, display screen G21 differs from display screens G11, G12, and G14 in that it displays the good annotation sample 430 instead of the correct annotation sample 420. In Figure 7, display screen G21 for data A is shown as an example.

[0123] Display screen G21 includes an OK sample field 701, an NG sample field 702, a gauge 703, and a forward button 704. Display screen G21 places the OK sample field 701 on the left and the NG sample field 702 on the right. Since the correct annotation samples are not displayed on display screen G21, data A to C are not limited to evaluation data, but can be three original training data randomly selected from the target dataset that has annotations related to the target annotation items. Alternatively, data A to C may be three original training data randomly selected from the original training data that has annotations related to the target annotation items, regardless of the dataset.

[0124] The OK sample column 701 displays three excellent annotation samples 430 randomly selected from the excellent annotation samples in Data A, arranged horizontally from left to right in order of highest accuracy evaluation value. The NG sample column 702 displays three poor annotation samples randomly selected from the poor annotation samples in Data A, arranged horizontally from left to right in order of highest accuracy evaluation value. The NG sample column 702 may also include poor annotation samples from the target annotator.

[0125] Gauge 703 is a horizontally elongated, leftward-pointing arrow-shaped image, indicating that annotation samples positioned to the left are of higher quality.

[0126] The "Next" button 704 is pressed when the target annotator wants to display good annotation samples for the next data. When the "Next" button 704 is pressed, the generation unit 13 displays the data B display screen G21 on the display.

[0127] If the target dataset is video, the poorly annotated sample 410 and the goodly annotated sample 430 will consist of video. In this case, the target annotator can press the next button 704 only after they have finished playing all the video displayed in the OK sample column 701 and the NG sample column 702.

[0128] Note that the display screen G21 for data C will show a test button 403 (see Figure 4) instead of the "Proceed" button 704. By pressing the test button 403, the target annotator can undergo annotation testing.

[0129] By viewing display screen G21, the annotator can compare the good annotation sample 430 with the bad annotation sample 410 to understand what level of accuracy is required for annotation, thereby improving the accuracy of their annotations. Furthermore, display screen G21 can display bad annotation samples from other annotators, allowing the annotator to proactively identify potential annotation errors. This helps prevent future annotation mistakes. The OK sample column 701 and NG sample column 702 display good and bad annotation samples in descending order of accuracy evaluation, allowing the annotator to intuitively understand what kind of annotation meets the standard conditions.

[0130] Figure 8 shows the display screen G22, which is the fifth example of the accuracy training menu. Display screen G22 differs from display screen G21 in that it always includes the defective annotation sample 705 of the target annotator in the NG sample column 702. Otherwise, the configuration of display screen G22 is the same as display screen G21, so a detailed explanation is omitted.

[0131] Display screen G22 offers the following additional benefits compared to display screen G21: Display screen G21 clearly shows the 705 defective annotation samples of the target annotator, allowing the target annotator to check the level of its own annotation accuracy.

[0132] Figure 9 shows the display screen G23 of the sixth example of the accuracy training menu. Display screen G23 differs from display screen G22 in that it re-annotates the defective annotation sample 705 of the target annotator with the target annotator.

[0133] On display screen G23, the NG sample column 702 displays the defective annotation sample 705 of the target annotator. The target annotator inputs the operation of pressing the re-annotation button 706. Then, the generation unit 13 displays the re-annotation screen G13 on the display. On the re-annotation screen G13, the original training data of the defective annotation sample 705 is displayed in the work column 502. In the work column 502, the target annotator who has entered the re-annotation inputs the operation of pressing the judgment button 503. Then, the generation unit 13 determines whether the re-annotation is successful or not. If the re-annotation is successful, the generation unit 13 displays display screen G25 on the display. Display screen G25 is displayed when the re-annotation is successful. In this example, since the re-annotation is successful, on display screen G25, the OK sample column 701 has the original training data to which the re-annotation was applied added as a good annotation sample 411. Furthermore, since the re-annotation is successful, the display screen G25 shows the "Next" button 704. When the target annotator inputs the operation to press the "Next" button 704, the generation unit 13 displays the display screen G23 on the display, which shows the next data, data B, including the excellent annotation sample 430.

[0134] On the other hand, if the re-annotation fails, display screen G23 is shown on the screen. In this case, the annotator inputs the operation of pressing the re-annotation button 706 again to perform re-annotation. In other words, display screen G23 is configured so that the next data cannot be viewed unless the re-annotation is passed.

[0135] Figure 10 shows the display screen G31 of the first example of a time-efficiency training menu. A time-efficiency training menu is a training menu that is generated when it is determined that the time evaluation value of the target annotator does not meet the standard conditions.

[0136] Display screen G31 includes an OK sample field 801, an NG sample field 802, and a test button 804. The NG sample field 802 displays a bad annotation sample 810 for each of the data A, B, and C. The bad annotation sample 810 is a bad annotation sample randomly selected from the bad annotation samples for each of the data A, B, and C. The NG sample field 802 may also include bad annotation samples of the target annotator.

[0137] The OK sample column 801 displays a good annotation sample 820 for each of the data sets A, B, and C. The good annotation sample 820 is a good annotation sample randomly selected from the good annotation samples for each of the data sets A, B, and C.

[0138] On display screen G31, data A to C are not limited to evaluation data; three data points randomly selected from the target dataset that has annotations related to the target annotation items can be used. Alternatively, data A to C may be three original training data points randomly selected from the original training data that has annotations related to the target annotation items, regardless of the dataset.

[0139] The poorly annotated sample 810 and the goodly annotated sample 820 are videos. These videos show the process by which the poorly annotated and goodly annotated models annotate the original training data. This allows the target annotator to see what procedure it should follow to improve its time evaluation score.

[0140] The annotator can play back annotation samples by pressing the play button 803 located on the annotation sample. Once the annotator has finished playing back all annotation samples, they can press the test button 804. Pressing the test button 804 allows the annotator to undergo an annotation test.

[0141] By viewing display screen G31, the annotator can compare the poorly annotated sample 810 with the goodly annotated sample 820 to identify the reasons why the standard conditions were not met, thereby improving the accuracy of their annotations. Furthermore, by comparing the goodly annotated sample 820 with the poorly annotated sample 810, the annotator can understand the appropriate procedure and speed for annotation, thus reducing the time required for annotation. Additionally, since display screen G31 shows poorly annotated samples 810 from other annotators, the annotator can proactively identify potential annotation errors. This helps prevent future annotation mistakes.

[0142] Figure 11 shows the display screen G32 of the second example of the time efficiency training menu. Display screen G32 differs from display screen G31 in that it displays 820 good annotation samples and 810 bad annotation samples for each of the data A to C. In Figure 11, display screen G32 for data A is shown as an example. The selection criteria for data A to C are the same as in display screen G31.

[0143] Display screen G32 includes an OK sample field 801, an NG sample field 802, a gauge 805, and a forward button 807. Display screen G32 places the OK sample field 801 on the left and the NG sample field 802 on the right.

[0144] The OK sample column 801 displays three excellent annotation samples 820 randomly selected from the excellent annotation samples in Data A, arranged horizontally from left to right in descending order of time evaluation value. The NG sample column 802 displays three poor annotation samples randomly selected from the poor annotation samples in Data A, arranged horizontally from left to right in descending order of time evaluation value. The NG sample column 802 may also include poor annotation samples from the target annotator.

[0145] Gauge 805 is a horizontally elongated, left-pointing arrow-shaped image, indicating that annotation samples positioned to the left are of higher quality.

[0146] The "Next" button 807 is pressed when the target annotator wants to display good annotation samples for the next data. When the "Next" button 807 is pressed, the generation unit 13 displays the data B display screen G32 on the display.

[0147] Display screen G32 offers the following benefits in addition to display screen G31: The OK sample column 801 and NG sample column 802 display the excellent annotation sample 820 and the poor annotation sample 810, respectively, in descending order of time evaluation value, allowing the target annotator to intuitively understand what kind of annotation will meet the standard conditions.

[0148] Figure 12 shows the display screen G33 of the third example of the time efficiency training menu. Display screen G33 differs from display screen G32 in that it always includes the defective annotation sample 806 of the target annotator in the NG sample column 802. Otherwise, the configuration of display screen G33 is the same as display screen G32, so a detailed explanation is omitted.

[0149] Display screen G33 offers the following additional benefits in addition to display screen G32: Display screen G33 clearly shows the 806 defective annotation samples of the target annotator, allowing the target annotator to check the level of its own annotation time efficiency.

[0150] Next, we will explain annotation testing. Figure 13 is a diagram illustrating the accuracy annotation test. Accuracy annotation testing is performed on annotation items in which the accuracy evaluation value of the target annotator is determined not to meet the standard conditions.

[0151] Test screen G41 is a screen for the annotator undergoing annotation testing to input annotations. The generation unit 13 displays test screen G41 when the target annotator inputs the operation of pressing the aforementioned test buttons 403 and 804.

[0152] Test screen G41 includes a test image 901 and a "Next" button 902. Test image 901 is randomly selected from evaluation data included in the target dataset, for example, if the target annotator is determined not to meet the criteria. In this example, the target annotator performs annotation that adds bounding boxes 903. This is because the accuracy evaluation value of the annotation that adds bounding boxes to the target dataset was determined not to meet the criteria. If the accuracy evaluation value of the annotation that adds class labels is determined not to meet the criteria, the target annotator is subjected to an annotation test that adds class labels.

[0153] Once the annotator has finished annotating the test image 901, they input the operation of pressing the next button 902. The generation unit 13 then displays the next test screen G41 on the display. After the annotation work is completed for all test screens G41, the generation unit 13 determines whether the annotation test passed or failed.

[0154] The generation unit 13 calculates an accuracy evaluation value for each test image 901 by comparing the correct annotation with the annotation of the target annotator. The generation unit 13 calculates the average of the accuracy evaluation values ​​calculated for each test image 901. This average value is an example of a test score. The generation unit 13 determines that the annotation test is a pass if the calculated average of the accuracy evaluation values ​​is smaller than the first threshold for the test. As the first threshold for the test, for example, the average of the accuracy evaluation values ​​corresponding to the bottom X% for each test image 901 is adopted.

[0155] If the generation unit 13 determines that the annotation test has passed, it displays a pass screen G43 on the display. The pass screen G43 includes a message indicating that the test has passed and a back button 904. If the operation of pressing the back button 904 is input, the generation unit 13 returns to step S3 in Figure 2. As a result, the target annotator resumes the work of adding annotations to the target dataset.

[0156] On the other hand, if the generation unit 13 determines that the annotation test has failed, it displays a failure screen G44 on the display. The failure screen G44 includes a message indicating that it has failed, a test result image 905, and a back button 906. The test result image 905 is a test image 901 that has been annotated by the target annotator, and is a test image 901 that does not meet the accuracy evaluation criteria. The target annotator looks at the bounding box 903 included in the test result image 905 to confirm the shortcomings of the annotations it has added.

[0157] Once this verification is complete, the target annotator presses the back button 906. When the operation of pressing the back button 906 is input, the generation unit 13 displays the test screen G41 on the display again and restarts the annotation test.

[0158] Thus, annotators that are not good at accurate annotation cannot resume annotation work until they pass an accuracy annotation test. Therefore, annotation work is prevented from being performed by annotators whose level of annotation accuracy does not meet the standard, and the generation of poor quality training data is prevented.

[0159] Figure 14 illustrates the time efficiency annotation test. The time efficiency annotation test is performed on annotation items where the time evaluation value of the target annotator is determined not to meet the criteria.

[0160] Test screen G51 is a screen for the annotator undergoing annotation testing to input annotations. The generation unit 13 displays test screen G41 when the target annotator inputs the operation of pressing the aforementioned test buttons 403 and 804.

[0161] Test screen G51 differs from test screen G41 in that it has a judgment button 908 instead of a forward button 902. The target annotator adds annotations to the test image 901, similar to test screen G41. Here, annotations that add bounding boxes 903 are performed. When test screen G51 is displayed, the generation unit 13 starts measuring the time required for the annotation work.

[0162] Once the annotator has finished annotating test image 901, they input the operation of pressing the judgment button 908. The generation unit 13 then calculates the accuracy evaluation value of the annotation applied to test image 901. If the calculated accuracy evaluation value is smaller than the first threshold, that is, if the annotation result does not fall under the category of a poorly annotated sample, the test screen G51 with the next test image 901 is displayed on the display. On the other hand, if the calculated accuracy evaluation value is larger than the first threshold, that is, if the annotation result falls under the category of an unnecessary annotation sample, the generation unit 13 displays the test screen G51 with the same test image 901 again on the display. In this way, the test screen G51 is configured so that annotation of the next test image 901 cannot be performed until the annotation that meets the accuracy criteria is performed. As a result, if the annotator performs a quick but sloppy annotation, they will not be able to proceed to the next test image 901, making it difficult to pass the annotation test. If the system proceeds to the next test image 901, the generation unit 13 stops measuring the time required for annotation on the previous test image 901 and obtains the time required for the previous test image 901. This time is then used as the time evaluation value for the previous test image 901.

[0163] Once the process of adding annotations to all test screens G51 is complete, the generation unit 13 determines whether the annotations are correct or incorrect.

[0164] The generation unit 13 calculates the average value of the time evaluation values ​​for each test image 901, and determines that the annotation test is successful if the calculated average value of the time evaluation values ​​is smaller than the second threshold for testing. As the second threshold for testing, for example, the average value of the time evaluation values ​​corresponding to the bottom X% for each test image 901 is adopted.

[0165] If the generation unit 13 determines that the annotation test has passed, it displays the pass screen G53 on the display. The pass screen G53 is the same as the pass screen G43.

[0166] On the other hand, if the generation unit 13 determines that the annotation test has failed, it displays a failure screen G54 on the display. The failure screen G54 differs from the failure screen G44 in that it includes a test results field 907. The test results field 907 displays the average working time per test image 901 of the target annotator in the annotation test, i.e., the average time evaluation value. Furthermore, the test results field 907 displays the passing time. The passing time is the second threshold for testing as described above.

[0167] Annotators that struggle with efficient annotation cannot resume annotation work until they pass an annotation time efficiency test. This prevents annotation work from being performed by annotators whose time efficiency level does not meet the standard, thus enabling the efficient generation of large amounts of training data.

[0168] Next, we will explain how to set a threshold. Figure 15 illustrates how a threshold is set. First, the setting unit 15 displays a selection screen 1001 on the display for selecting an annotation item. The annotator inputs the operation to select the desired annotation item from the selection screen 1001. In this example, the selectable annotation items are "Enclose area with rectangle," "Classify area," and "Select time period." "Enclose area with rectangle" is an annotation item that adds a bounding box. "Classify area" is an annotation item that adds a class label. "Select time period" is an annotation item that adds a time period label. The annotator selects the annotation item for which they wish to set a threshold from among these annotation items.

[0169] When an operation to select an annotation item is input, the setting unit 15 displays the first setting screen 1002 and the second setting screen 1003 on the display.

[0170] The first setting screen 1002 is a screen for setting the first threshold, that is, the threshold for evaluating the accuracy evaluation value. The first setting screen 1002 includes an adjustment unit 1011, a gauge 1012, and a sample field 1013. The adjustment unit 1011 consists of a GUI component of an adjustment knob. The gauge 1012 indicates the adjustment range of the adjustment unit 1011 in 10 steps from 1 to 10. The adjustment unit 1011 is configured to slide vertically along the gauge 1012, and the first threshold is configured to be adjustable in 10 steps from 1 to 10. For example, step "10" corresponds to an accuracy evaluation value corresponding to the top 10% of all annotators, step "9" corresponds to an accuracy evaluation value corresponding to the top 20% of all annotators, and so on, with accuracy evaluation values ​​corresponding to ranks assigned to steps "1" to "10".

[0171] The sample field 1013 displays multiple annotation samples that meet the accuracy evaluation value corresponding to the stage indicated by the adjustment unit 1011. This allows the annotator to check the level of annotation according to the stage. Six annotation samples are displayed here, but this is just one example.

[0172] The second settings screen 1003 is a screen for setting the second threshold, that is, the threshold for evaluating the time evaluation value. The second settings screen 1003 differs from the first settings screen 1002 in that it has a reference time field 1014 instead of a sample field 1013.

[0173] In the second settings screen 1003, level "10" corresponds to the time evaluation value of the top 10% of all annotators, level "9" corresponds to the time evaluation value of the top 20% of all annotators, and so on, with each level from "1" to "10" being assigned a time evaluation value corresponding to its rank.

[0174] The reference time column 1014 displays the work time per sample corresponding to the stage indicated by the adjustment unit 1011. The work time per sample is the annotation work time that can be spent on one original training data sheet, and is a time evaluation value assigned to each stage.

[0175] It should be noted that accuracy and time efficiency are in a trade-off relationship. Therefore, the setting unit 15 may automatically determine the threshold of the other setting screen 1003 if a threshold is determined in either the first setting screen 1002 or the second setting screen 1003. For example, if the first threshold is set to level "9" in the first setting screen 1002, the setting unit 15 may automatically set the second threshold to level "1", and if the first threshold is set to level "8" in the first setting screen 1002, the setting unit 15 may automatically set the second threshold to level "2".

[0176] Next, we will explain the detailed method for calculating the accuracy evaluation value.

[0177] As mentioned above, the accuracy score is defined by the error between the correct annotation and the annotations added by the annotator.

[0178] The first example of an accuracy evaluation value applies when an annotation item is assigned a class label or a time segment label. In this case, as mentioned above, if the class label assigned by the annotator is correct, the error is "0", so the accuracy evaluation value is "0". On the other hand, if the class label assigned by the annotator is incorrect, the error is "1", so the accuracy evaluation value is "1".

[0179] The second example of an accuracy metric applies when an annotation item assigns multiple class labels or multiple time segment labels to the original training data. In this case, the error, i.e., the accuracy metric, is defined as follows:

[0180] Accuracy score = (Number of missing class labels + Number of incorrectly assigned class labels) / Total number of classes (1) The total number of classes is the total number of class labels that should be assigned to the original training data. For example, if there are 3 class labels that should be assigned to a given set of original training data, the total number of classes is 3. The number of missing class labels is the number of class labels that the annotator has not assigned relative to the total number of classes. For example, if only 2 class labels have been assigned to an original training data set with a total of 3 classes, the number of missing class labels is 1. The number of incorrectly assigned class labels is, for example, the number of class labels that the annotator has incorrectly assigned to a given set of original training data. For example, if the class label "cat" is assigned to a dog object, the number of incorrectly assigned class labels is 1.

[0181] The third example of an accuracy rating applies when an annotation item has a bounding box. In this case, as mentioned above, the accuracy rating is defined as 1-IOU.

[0182] The fourth example of the accuracy evaluation value is applied when the annotation item is the assignment of a bounding box and the original training data includes a person object. In person estimation, such as person matching or attribute estimation, the upper body provides more useful information than the lower body. Therefore, the evaluation unit 12 calculates the accuracy evaluation value such that the error is larger for bounding boxes where the upper body is cut off than for bounding boxes where the lower body is cut off.

[0183] The evaluation unit 12 divides the bounding box assigned to the person into an upper body region corresponding to the person's upper body and a lower body region corresponding to the person's lower body. The evaluation unit 12 calculates the Interaction Over Union (IOU) of the upper body region and the lower body region. The evaluation unit 12 calculates an accuracy evaluation value by weighting the IOU of the upper body region and the IOU of the lower body region so that the weight value of the IOU of the upper body region is higher than the weight value of the IOU of the lower body region. The accuracy evaluation value is defined as follows.

[0184] Accuracy rating = 1 - (0.8 × upper IOU + 0.2 × lower IOU) (2) 0.8 is the weight value for the upper IOU. 0.2 is the weight value for the lower IOU. The upper IOU is the IOU for the upper body region. The lower IOU is the IOU for the lower body region. Note that the weight values ​​in the above formula are just examples, and any values ​​can be used as long as the weight value for the upper IOU is greater than the weight value for the lower IOU, under the constraint that the sum of both weight values ​​is 1.

[0185] Figure 16 is an explanatory diagram of the fourth example of the accuracy evaluation value. Image 1210 is an image in which the correct bounding box 1211 has been assigned to person 1240. Image 1220 is an image in which the bounding box 1221 has been assigned to person 1240 with the lower half of the body cut off. Image 1230 is an image in which the bounding box 1231 has been assigned to person 1240 with the upper half of the body cut off. Hereinafter, when referring to bounding boxes 1221 and 1231 collectively, they will be called the target bounding box. The evaluation unit 12 divides the target bounding box into upper and lower halves by the boundary line 1600. The boundary line 1600 is a straight line that divides the correct bounding box 1211 into two equal parts vertically.

[0186] As a result, bounding boxes 1221 and 1231 are divided into upper body and lower body regions, respectively. The evaluation unit 12 calculates the upper IOU and lower IOU of the target bounding box.

[0187] The upper IOU is expressed as the ratio of the area represented by the logical AND of the upper body area of ​​the target bounding box to the area represented by the logical OR of the upper body area of ​​the target bounding box.

[0188] The lower IOU is expressed as the ratio of the area represented by the logical AND of the lower body region of the target bounding box to the area represented by the logical OR of the lower body region of the target bounding box.

[0189] The evaluation unit 12 then substitutes the upper IOU and lower IOU into the above-mentioned formula (2) to calculate the accuracy evaluation value.

[0190] In the example in Figure 16, bounding box 1221 includes the entire upper body but cuts off the lower body. On the other hand, bounding box 1231 includes the entire lower body but cuts off the upper body. Therefore, bounding box 1221 has a lower accuracy rating than bounding box 1231. As a result, bounding box 1221 has a higher accuracy rating than bounding box 1231.

[0191] The fifth example of an accuracy evaluation value is when the annotation item is the assignment of a bounding box, and the original training data contains objects of anomalies. Anomalies are areas that indicate defects or dangerous parts. When detecting anomalies, it is important to prevent overlooking problem areas with anomalies. Therefore, the evaluation unit 12 calculates an accuracy evaluation value such that the error is large for bounding boxes that do not contain all of the anomalies.

[0192] Specifically, the evaluation unit 12 calculates the correct bounding box inclusion ratio, which is the ratio of the target bounding box to the correct bounding box. Based on the IOU (Interaction Over Union) and the correct bounding box inclusion ratio of the target bounding box, the evaluation unit 12 calculates the accuracy evaluation value. The accuracy evaluation value is defined as follows.

[0193] Accuracy rating = 1 - (0.5 × IOU + 0.5 × percentage of correct answers) (3) 0.5 is the weight value between IOU and the percentage of correct answers. Note that the weight value is just an example, and any appropriate value can be used.

[0194] The percentage of the correct answer included is expressed as the ratio of the area represented by the logical AND of the correct bounding box and the target bounding box to the area of ​​the correct bounding box.

[0195] Figure 17 is an explanatory diagram of the fifth example of accuracy evaluation values. Image 1310 is an image 1312 with a correct bounding box 1311 assigned to it, Image 1320 is an image with a bounding box 1321 assigned when the correct inclusion rate is 1.0, and Image 1330 is an image with a bounding box 1331 assigned when the correct inclusion rate is 0.7.

[0196] Bounding box 1321 contains all of the correct bounding box 1311. On the other hand, bounding box 1331 does not contain all of the correct bounding box 1311. Therefore, bounding box 1321 has a lower accuracy score and a higher score than bounding box 1331.

[0197] The sixth example of the accuracy score applies when the annotation item is a time segment label. In this case, the accuracy score is defined by the following formula:

[0198] Accuracy score = 1 - IOU (4) In this case, IOU is expressed as the ratio of the time represented by the logical AND of the correct interval and the target interval to the time represented by the logical OR of the correct interval and the target interval to which time segment labels have been assigned.

[0199] The seventh example of accuracy evaluation criteria applies when the annotation item is the assignment of time segment labels and the original training data contains objects in anomalous intervals. In other words, the seventh example of accuracy evaluation criteria applies when evaluating annotation that labels anomalous intervals within the entire range of a video.

[0200] Figure 18 is an explanatory diagram of the seventh example of accuracy evaluation values. Graph 1400 shows the correct interval 1401, to which correct classification labels are assigned, from the playback interval of the original training data consisting of video images. Graph 1410 shows the target interval 1411, to which the correct answer inclusion rate is 1.0, from the playback interval of the original training data consisting of video images. Graph 1420 shows the target interval 1421, to which the correct answer inclusion rate is 0.8, from the entire interval of the original image data consisting of video images.

[0201] In the seventh example, the accuracy evaluation value is calculated using the formula (3) described above.

[0202] However, in this case, the percentage of correct answers included is expressed as the ratio of the time represented by the logical AND of the correct answer interval 1401 and the target interval (1411, 1421) to the correct answer interval 1401. In the case of graph 1410, the target interval 1411 includes the entire correct answer interval 1401, so the percentage of correct answers included is 1. In the case of graph 1420, the target interval 1421 does not include the entire correct answer interval 1401, so the percentage of correct answers included is a value less than 1.

[0203] Therefore, the accuracy evaluation value calculated for case 1410 is smaller and the evaluation is higher than that for case 1420.

[0204] When detecting abnormal sections where a dangerous condition has occurred, it is important to detect all abnormal sections. Therefore, the evaluation unit 12 calculates an accuracy evaluation value such that the error is larger when not all abnormal sections are included compared to when all abnormal sections are included. This ensures that annotation is performed more carefully to prevent omissions of abnormal sections.

[0205] As explained above, according to this embodiment, a display screen for a training menu corresponding to the evaluation value of the annotation is shown on the annotator's display. Therefore, the annotator can overcome the annotations they are not good at through the training menu. This makes it possible to efficiently improve the annotator's skills. [Industrial applicability]

[0206] This disclosure will be useful in the field of technology for generating training data. [Explanation of Symbols]

[0207] 1: Information Processing Device 2: Terminal device 10: Processor 11: Acquisition part 12: Evaluation Department 13: Generation part 14: Output section 15: Settings Section 20: Memory 30: Communications Department

Claims

1. A method of information processing in a computer, The annotation work history by the annotator is obtained, and the annotations are applied to the original training data in order to generate training data for the machine learning model. Based on the work history, an evaluation value is calculated to evaluate the annotation. A training menu corresponding to the aforementioned evaluation value, which generates the training menu for improving the annotation skills, The aforementioned training menu is output to the annotator's display. Information processing methods.

2. The aforementioned evaluation value is calculated for each of the multiple evaluation criteria, The aforementioned training menu includes training menu content corresponding to evaluation criteria in which the calculated evaluation value does not meet the standard conditions, The information processing method according to claim 1.

3. The aforementioned multiple evaluation criteria include an accuracy evaluation criterion for evaluating the accuracy of the annotation and a time efficiency evaluation criterion for evaluating the time spent on the annotation work. The information processing method according to claim 2.

4. The aforementioned annotations are classified into multiple annotation items, The aforementioned evaluation value is calculated for each of the multiple evaluation criteria for each annotation item. The aforementioned training menu includes training menu content corresponding to the evaluation criteria and annotation items for which the calculated evaluation value does not meet the aforementioned standard conditions. The information processing method according to claim 2 or 3.

5. The aforementioned plurality of annotation items are at least one of the following types: one that assigns a class label to an object contained in an image, one that assigns a bounding box to the object, and one that assigns a time segment label to a time-series image. The information processing method according to claim 4.

6. The aforementioned training menu consists of a display screen that includes at least one of the following: a good annotation sample showing the annotation performed by a good annotator or a correct annotation sample to which the correct annotation has been assigned, and a bad annotation sample showing the annotation performed by a bad annotator. The information processing method according to claim 1 or 2.

7. The above-mentioned excellent annotation sample and the above-mentioned poor annotation sample are images showing the results of the annotation performed by the excellent annotator and the above-mentioned poor annotator. The information processing method according to claim 6.

8. The above-mentioned excellent annotation sample and the above-mentioned poor annotation sample are video images showing the process by which the annotation is applied by the excellent annotator and the poor annotator. The information processing method according to claim 6.

9. The aforementioned training menu includes an annotation test having test content corresponding to the evaluation value, wherein the annotation test is assigned to the annotator to perform the task of adding annotations to the test image. The information processing method according to claim 1 or 2.

10. The aforementioned annotation test is, The test score of the annotation test is calculated from the results of the annotation test, The pass / fail status of the annotation test is determined from the test score, If the annotator is determined to have passed the annotation test, the annotation process is to be resumed, including: The information processing method according to claim 9.

11. The aforementioned training menu includes outputting an annotation sample to the display if the annotator is determined to have failed the annotation test. The information processing method according to claim 10.

12. Furthermore, the display includes a threshold setting screen for evaluating the evaluation value, The information processing method according to claim 1 or 2.

13. The threshold includes a first threshold for evaluating the accuracy of the annotation, The aforementioned settings screen includes a first settings screen for setting the first threshold, The aforementioned first settings screen is, To adjust the first threshold, an adjustment unit is provided, Includes a display field for displaying annotation samples corresponding to the first threshold, The information processing method according to claim 12.

14. The threshold includes a second threshold for evaluating the accuracy of the annotation. The aforementioned settings screen is, Includes a second settings screen for setting the second threshold, The second settings screen mentioned above is, An adjustment unit for adjusting the second threshold, Includes a display field that displays the annotation work time per sample according to the second threshold, The information processing method according to claim 12.

15. The annotation includes multiple annotation items, The display of the aforementioned settings screen is as follows: The system accepts a selection instruction for one of the aforementioned multiple annotation items, This includes displaying a settings screen for setting the threshold for the annotation item in item 1, The information processing method according to claim 12.

16. The calculation of the aforementioned evaluation value is as follows: The annotations that the annotator has applied to a person included in the image are divided into an upper body region corresponding to the upper body of the person and a lower body region corresponding to the lower body of the person. To calculate the Interaction Over Union (IOU) of the upper body region and the IOU of the lower body region, This includes calculating the evaluation value based on a weighted average of the IOU of the upper body and the IOU of the lower body such that the weight value of the IOU of the upper body region is higher than the weight value of the IOU of the lower body region. The information processing method according to claim 1 or 2.

17. The calculation of the aforementioned evaluation value is as follows: The correct inclusion rate is calculated as the ratio of annotations added by the annotator to the image relative to the correct annotations, This includes calculating the evaluation value based on the Interaction Over Union (IOU) of the annotation and the correct answer inclusion rate, The information processing method according to claim 1 or 2.

18. The evaluation value is calculated based on the annotation results that the annotator assigns to the evaluation data included in the dataset on which the annotator performs the annotation. The evaluation data is the original training data with hidden ground truth annotations. The information processing method according to claim 1 or 2.

19. An information processing device including a processor, The aforementioned processor, The annotation work history by the annotator is obtained, and the annotations are applied to the original training data in order to generate training data for the machine learning model. Based on the work history, an evaluation value is calculated to evaluate the annotation. A training menu corresponding to the aforementioned evaluation value, which generates the training menu for improving the annotation skills, The process of outputting the aforementioned training menu to the annotator's display is executed. Information processing device.

20. On the computer, The annotation work history by the annotator is obtained, and the annotations are applied to the original training data in order to generate training data for the machine learning model. Based on the work history, an evaluation value is calculated to evaluate the annotation. A training menu corresponding to the aforementioned evaluation value, which generates the training menu for improving the annotation skills, The process of outputting the aforementioned training menu to the annotator's display is executed. Information processing program.