Annotation work support system, annotation work support method, and annotation work support program

The annotation support system simplifies the retraining of AI models by storing and displaying historical annotation information, making the reannotation process more efficient and effective.

JP2026004636APending Publication Date: 2026-01-15EVIDENT CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024102454
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-26
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

The process of updating an AI model for classifying tasks in video recordings of assembly processes is cumbersome due to the need for re-annotation of training videos, and the lack of information about previous annotation work complicates this process.

Method used

An annotation support system that includes a memory unit to store trained models, associated microscope videos, annotation information, and classification standards, allowing for the display of this information to assist model developers in relearning tasks.

Benefits of technology

Facilitates easier and more informed annotation work for updating AI models by providing historical annotation data, thereby reducing the burden and improving the efficiency of retraining processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026004636000001_ABST
    Figure 2026004636000001_ABST
Patent Text Reader

Abstract

To support and facilitate annotation work for relearning a learned model.SOLUTION: The storage unit stores at least one learned model learned using at least one microscope moving image, and a microscope moving image, annotation information (including classification for each moving image segment included in the microscope moving image) indicating an annotation assigned to the microscope moving image, and classification reference information indicating an assignment reference of classification, which are associated with each learned model. The control unit receives selection of at least one learned model, acquires the microscope moving image, the annotation information, and the classification criterion information associated with the selected learned model from the storage unit, displays the acquired microscope moving image and annotation information on the display in association with each other, and displays the acquired classification criterion information on the display.SELECTED DRAWING: Figure 14
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The disclosure of this specification relates to an annotation work support system, an annotation work support method, and an annotation work support program. [Background technology]

[0002] There is known a technology for supporting the creation of teaching data for machine learning of learning data, which is used to classify objects from the morphology of the objects obtained by imaging a carrier that supports cells (see, for example, Patent Document 1). This technology enables a user to classify objects by displaying a teaching image containing the object for creating teaching data on a display unit.

[0003] Furthermore, a technology called TMRNet is known as a technology for recognizing actions related to a task from a video of the task (see, for example, Non-Patent Document 1). TMRNet is an abbreviation for Temporal Memory Relation Network, and is a technology that identifies what task the action shown in the current frame is, based on the relationship between each of multiple frames. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-009314 [Non-patent literature]

[0005] [Non-Patent Document 1] Yueming Jin, 5 others, “Temporal Memory Relation Network for Workflow Recognition from Surgical Video”, IEEE Transactions on Medical Imaging, Volume 40, Issue 7, July 2021 Summary of the Invention [Problem to be solved by the invention]

[0006] An inference model (trained model) generated by machine learning can be used to classify the tasks involved in videos of work processes on an object captured using a microscope. When this trained model is actually used, re-training may be repeatedly performed to obtain an updated version of the trained model, for example, whenever a change occurs in the work process or to meet a request for improved classification accuracy.

[0007] When relearning is performed, annotation work may be performed to re-annotate the training video used for relearning in order to change the annotations that were added to the training video. Note that annotation work refers to the work of marking the boundary between two consecutive video segments when dividing the training video into multiple video segments for each task, and assigning a classification for each task to each video segment.

[0008] When annotating a training video to be used for relearning, information about the annotation work previously performed on that training video (such as the division positions of the training video, the classifications assigned to each video segment, and the criteria for assigning classifications) is important in supporting the various decisions required of the annotator (such as where to divide the training video and which task to classify it into).

[0009] In view of the above, an object of one aspect of the present invention is to support annotation work for relearning a trained model and to make the annotation work easier. [Means for solving the problem]

[0010] An annotation support system according to one aspect of the present invention includes a memory unit and a control unit. The memory unit stores at least one trained model trained using at least one microscope video, a microscope video associated with each of the trained models, annotation information associated with each of the trained models indicating annotations assigned to the microscope video, the annotation information including a classification for a video segment included in the microscope video, and classification standard information associated with each of the trained models indicating a standard for assigning the classification. The control unit accepts selection of at least one trained model, acquires from the memory the microscope video, the annotation information, and the classification standard information associated with the selected trained model, associates the acquired microscope video and annotation information with each other, and displays the acquired microscope video and annotation information on a display. [Effects of the Invention]

[0011] According to the above aspect, it is possible to support annotation work for relearning a trained model, and to make the annotation work easier. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a flowchart showing an example of a procedure for creating an AI model. [Figure 2] 10 is a flowchart showing an example of a procedure for updating an AI model. [Figure 3] FIG. 1 illustrates an example of an annotation work support system. [Figure 4] FIG. 1 illustrates an example of a hardware configuration of a computer. [Figure 5] FIG. 2 is a diagram illustrating an example of a data structure of a model design information DB. [Figure 6] FIG. 10 is a diagram illustrating the number of classes classified into moving images. [Figure 7] FIG. 10 is a diagram illustrating an example of the data structure of a model-related file information DB. [Figure 8] 10 is a flowchart illustrating an example of an annotation work support process for relearning. [Figure 9] FIG. 10 is a diagram showing an example of a model selection screen. [Figure 10] 10 is a flowchart showing an example of an annotation work screen process; [Figure 11] FIG. 10 is a diagram showing an example of an annotation work screen. [Figure 12] 10 is a flowchart showing an example of a moving image segment screen process. [Figure 13] 10 is a flowchart showing processing details of an example of classification criteria information screen processing. [Figure 14] FIG. 10 is a diagram showing an example of an annotation work screen. [Figure 15] 10 is a flowchart showing the processing content of an example of a first annotation information editing process. [Figure 16] FIG. 10 is a diagram showing an example of an annotation work screen. [Figure 17] 10 is a flowchart showing the processing content of an example of a second annotation information editing process. [Figure 18] FIG. 10 is a diagram showing a display example when a video segment is further divided. [Figure 19] 10 is a flowchart showing an example of a moving image tag information screen process; [Figure 20] FIG. 10 is a diagram showing an example of a moving image tag information screen. [Figure 21] 10 is a flowchart showing the processing contents of an example of a re-learning process. [Figure 22] FIG. 10 is a diagram showing another example of the classification criteria information screen. [Figure 23] FIG. 10 is a diagram showing an example of editing classification criterion information. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, the embodiments will be described in detail with reference to the drawings.

[0014] Even today, with the increasing automation of work using robots and other tools, there are still many products that require manual assembly, and medical equipment is one example. Because the assembly of precision equipment like medical equipment involves many delicate tasks, it is often done under a microscope, and stereomicroscopes that allow the object to be viewed in three dimensions with both eyes are often used for such work. Working under such a microscope is difficult, and variations in the work are likely to occur.

[0015] To reduce variations in work, the work may be limited to only trained workers. However, since variations in work due to individual proficiency are unavoidable, the work may be recorded under a microscope to enable confirmation of the appropriateness of the work and results. This recording method may involve capturing video of the work or results using a microscope camera.

[0016] The amount of video obtained in this way is large in daily product production. For this reason, it is not practical for a reviewer to check each video segment one by one to determine which task in a series of assembly processes each video segment represents. Therefore, a method has recently been proposed in which an AI model divides a video recording of a series of assembly processes and classifies them by task. "AI" stands for artificial intelligence. For example, the aforementioned TMRNet can be used as an AI model for this purpose.

[0017] Here, we will explain the process of creating an AI model that classifies each task in the product assembly process. Figure 1 is a flowchart showing an example of the procedure for creating an AI model.

[0018] In the AI ​​model creation process, first, the overall design of the AI ​​model to be created (e.g., how many video segments to divide the video into, which tasks each video segment should be classified into, etc.) is considered (S11). Next, the learning video showing the assembly process tasks is acquired (S12).

[0019] Next, annotation is performed on each video acquired in S12 according to the results of the review in S11 (S13).

[0020] Next, various conditions (learning conditions) for machine learning to create an AI model are set (S14), such as the number of learning iterations and the threshold for determining convergence of the learning.

[0021] Next, under the learning conditions set by the work in S14, machine learning and validation of the learning results are carried out using the annotated learning videos obtained by the work up to S13 as training data (S15).

[0022] Next, to test the AI ​​model obtained as a result of the learning process in S15, the AI ​​model is made to classify tasks shown in videos other than the training data (S16).Then, it is determined whether the results of this test are valid (S17).

[0023] If the test results are determined to be valid, the AI ​​model creation process ends. On the other hand, if the test results are determined to be invalid, the AI ​​model is recreated. Specifically, the trial and error process of re-acquiring the learning video (S18), re-annotating (S19), and re-setting the learning conditions (S20) is repeated until the test results of S16 are determined to be valid.

[0024] An AI model is completed through this kind of creation process, for example.

[0025] After creating an AI model, it may become necessary to update the AI ​​model due to circumstances such as changes to the assembly process or a desire to improve the accuracy of task classification. Next, we will explain the process of updating an AI model. Figure 2 is a flowchart showing an example of the procedure for updating an AI model.

[0026] When an AI model needs to be updated, first, the design information at the time of creation of the current version of the AI ​​model to be updated and each version prior to the current version is referenced (S21). Next, based on the design information referenced in S21, the overall design of the updated version of the AI ​​model is reviewed (S22). Note that the design information includes, for example, the time of model creation, information about the video that serves as training data, the shooting conditions of the video, the model's performance at the time of creation, and information about the worker who was the subject of the model.

[0027] By referring to the design information through the S21 process, model developers who are updating an AI model can understand the design intent and process at the time the previous version of the AI ​​model was created, and can obtain an updated version of the AI ​​model that allows them to understand the changes from the previous version, regardless of whether there were any problems with the previous version that can be derived from the design intent and process.

[0028] Next, based on the results of the review work in S22, a decision is made as to whether or not it is necessary to acquire additional new learning videos, which are different from those used in the machine learning to create the previous version of the AI ​​model (S23).

[0029] If it is determined that additional acquisition is necessary, the following steps are performed in sequence: additional acquisition of training videos (S24), annotation of the additionally acquired videos (S26), and resetting of the training conditions associated with the additional acquisition of videos (S28). Then, as a re-learning task, the machine learning and validation tasks of S15 in the AI ​​model creation task shown in Figure 1, and the subsequent tasks from S16 onwards, are performed in sequence.

[0030] On the other hand, if it is determined that there is no need to acquire additional training videos, the next step is to determine whether or not changes are needed to the annotations attached to the already acquired training videos that were used to create the old version of the AI ​​model, based on the results of the review work in S22 (S25).

[0031] If it is determined that annotation changes are necessary, the acquired learning video is annotated (S26) and the learning conditions are reset (S28) to accommodate the annotation changes. After that, the AI ​​model creation process shown in Figure 1, including the machine learning and validation steps in S15 and the steps from S16 onward, is carried out as a re-learning process.

[0032] On the other hand, if it is determined that neither additional acquisition of learning videos nor changes to annotations is necessary, the next step is to determine whether or not the learning conditions need to be reset based on the results of the review work in S22 (S27).

[0033] If it is determined that the learning conditions need to be reset, the learning conditions are reset (S28). After that, as a re-learning task, the machine learning and validation tasks of S15 in the AI ​​model creation task shown in Figure 1, and the subsequent tasks from S16 onwards are performed in sequence.

[0034] On the other hand, if it is determined that there is no need to acquire additional learning videos, change annotations, or reset learning conditions, the review process in S22 is performed again, and then the process from S23 onwards is performed again.

[0035] For example, the above steps are performed when updating an AI model. In this procedure, if it is determined in S25 that an annotation change is necessary, annotation work to make the change is performed on the acquired training video in S26. At this time, if information about the annotation work previously performed on the training video can be obtained (information such as the division positions of the training video, the classifications assigned to each video segment, and the criteria for assigning the classification), the purpose and intention of the annotation work at that time can be understood, thereby reducing the burden of annotation work in updating an AI model, i.e., annotation work for re-training a trained AI model.

[0036] Therefore, in the following description, as an embodiment of the present invention, we will explain a system that supports annotation work by model developers by displaying information about previously performed annotation work when annotation work is being performed to re-train a trained AI model.

[0037] First, a description will be given of Fig. 3. Fig. 3 shows the overall configuration of an example of an annotation work support system 1.

[0038] The annotation work support system 1 includes a microscope 100, a control device 200, a monitor 300, and a plurality of input devices 400 (a mouse 401, a keyboard 402, a foot switch 403, and a barcode reader 404).

[0039] Microscope 100 is a stereo microscope that allows for stereoscopic viewing of a sample. A user can observe an optical image formed by the microscope optical system on the object side of eyepiece 106 with both eyes via eyepiece 106, allowing for stereoscopic viewing of the object. Microscope 100 is suitable for applications such as precision equipment assembly work.

[0040] The microscope 100 is equipped with a zoom lens that can be operated with a zoom handle 130. By operating the zoom handle 130, the user can change the observation magnification while continuing to look through the eyepiece 106 and observe the object.

[0041] The microscope 100 includes a focusing handle 140. By operating the focusing handle 140, the user can change the distance between the object and the objective lens 101 and bring the object into focus.

[0042] The microscope 100 is equipped with an imaging device 112 that captures an image of an object and acquires a video of the object (microscope video). The eyepiece tube 120 to which the eyepiece 106 is attached is a triplet tube, and the imaging device 112 is attached to the eyepiece tube 120. The imaging device 112 is provided with a two-dimensional image sensor. The image sensor is not particularly limited, but may be, for example, a CCD image sensor or a CMOS image sensor. The video acquired by the imaging device 112 is output to the control device 200. The video may also be output directly to the monitor 300.

[0043] Light split from the optical path of an optical system (not shown) included in the microscope 100 by a beam splitter such as a half mirror enters the imaging device 112 via an imaging lens (not shown).

[0044] The microscope 100 includes a projector 113 that projects an auxiliary image onto an image plane where an imaging lens forms an optical image. The projector 113 is a device that projects and superimposes an auxiliary image onto the image plane in accordance with a command from the control device 200. More specifically, the projector 113 superimposes the auxiliary image onto the image plane based on auxiliary image data, which will be described later. The type of the projector 113 is not particularly limited. The projector 113 may be configured using, for example, a liquid crystal device or a digital mirror device.

[0045] The projector 113 is provided inside the eyepiece tube 120. Light from the projector 113 is guided to the optical path of the optical system of the microscope 100.

[0046] The eyepiece tube 120 is provided with an operation unit 121. By operating the operation unit 121, the user can switch the projector 113 on and off and instruct it to start or stop superimposing the auxiliary image on the image plane.

[0047] The control device 200 controls the microscope 100. The control device 200 generates the auxiliary image data described above and outputs it to the microscope 100 (projector 113).

[0048] The monitor 300 and the input device 400 are connected to the control device 200. The monitor 300 is, for example, a liquid crystal display or an organic EL display, and functions as a display in the annotation work support system 1. Note that "EL" is an abbreviation for electro-luminescence.

[0049] 4 shows an example of the hardware configuration of a computer 200a for realizing the control device 200 in the annotation work support system 1. The computer 200a includes, as hardware, a processor 201, a memory 202, a storage device 203, a reading device 204, a communication interface 206, and an input / output interface 207. The processor 201, the memory 202, the storage device 203, the reading device 204, the communication interface 206, and the input / output interface 207 are connected to one another via, for example, a bus 208.

[0050] The processor 201 may be, for example, a single processor, a multiprocessor, or a multi-core processor. The processor 201 reads and executes programs stored in the storage device 203 to perform various control processes including annotation work support processing for relearning, which will be described later, and provides a function as a control unit in the annotation work support system 1.

[0051] The memory 202 may be, for example, a semiconductor memory, and may include a RAM area and a ROM area. Note that "RAM" is an abbreviation for Random Access Memory, and "ROM" is an abbreviation for Read Only Memory.

[0052] The storage device 203 is, for example, a semiconductor memory such as a hard disk or a flash memory, or an external storage device, and provides a function as a storage unit in the annotation work support system 1. More specifically, the storage device 203 stores, for example, configuration data of at least one trained model trained using at least one video captured by the imaging device 112 of the microscope 100. The storage device 203 also stores a model design information DB 500 and a model-related file information DB 600, which will be described later. Note that "DB" is an abbreviation for database.

[0053] The reader 204 accesses the removable storage medium 205 in accordance with, for example, an instruction from the processor 201. The removable storage medium 205 is realized by, for example, a semiconductor device, a medium for inputting and outputting information by magnetic action, or a medium for inputting and outputting information by optical action. An example of a semiconductor device is a USB (Universal Serial Bus) memory. An example of a medium for inputting and outputting information by magnetic action is a magnetic disk. An example of a medium for inputting and outputting information by optical action is a CD (Compact Disc)-ROM, a DVD (Digital Versatile Disc), or a Blu-ray (registered trademark) disk.

[0054] The communication interface 206 communicates with other devices (such as the microscope 100) according to instructions from the processor 201, for example. The input / output interface 207 is, for example, an interface between an input device 400 and an output device. The input device 400 is, for example, a device such as a mouse 401, a keyboard 402, or a foot switch 403 that receives instructions from a user. The output device is, for example, a monitor 300 and an audio device such as a speaker. Note that operations such as a "click operation" described below are described as operations performed by the mouse 401, for example, but are not limited to clicking as long as they are designation operations using the input device 400.

[0055] The program executed by the processor 201 is provided to the computer in the following form, for example. (1) It is pre-installed in the storage device 203. (2) Provided by removable storage medium 205. (3) Provided from a server such as a program server.

[0056] Note that the hardware configuration of the computer 200a for realizing the control device 200 described with reference to FIG. 4 is an example, and the embodiment is not limited to this. For example, part of the above-described configuration may be deleted, or new configuration may be added. Furthermore, in another embodiment, for example, part or all of the functions of the control device 200 may be implemented as hardware. FPGA (Field Programmable Gate Array), SoC (System-on-a-Chip), ASIC (Application Specific Integrated Circuit), and PLD (Programmable Logic Device) are examples of hardware that can implement the control device 200.

[0057] Next, a description will be given of the model design information DB 500 stored in the storage device 203. Fig. 5 shows an example of the data structure of the model design information DB 500.

[0058] Each trained model stored in the storage device 203 is associated with version information indicating the version of the trained model. In the model design information DB 500 in Figure 5, for each trained model, one or more pieces of version information indicating the version of the trained model are associated with one or more pieces of design information about the trained model. Note that in Figure 5, "model name" is the name given to the AI ​​model (trained model), and "revision" is an example of version information.

[0059] The design information is information that includes at least one of time information, personal information, learning information, and character information. In FIG. 5, "Number of Videos," "Number of Classifications," "Score," and "Number of Inferences" are examples of learning information, and "Video Tag Information" and "Updater" are examples that include personal information. Also, "Use (any word)" is an example of character information, and "Creation Date and Time" and "Last Update Date and Time" are examples of time information. In other words, in the model design information DB 500 in FIG. 5, these pieces of design information about the trained model are associated with "revisions" that identify the version of the trained model.

[0060] The design information in FIG. 5 will be further explained.

[0061] "Number of videos" is information about the number of training videos used in the machine learning performed when creating the trained model.

[0062] The "number of classifications" is information about the number of classifications that the trained model makes when classifying a single video obtained by photographing an assembly process with the microscope 100 into video segments of several tasks that make up the assembly process.

[0063] For example, in the video of the assembly process of a certain part shown in Figure 6, one video is classified into video segments from "Class 00" to "Class 06" for each task that makes up the assembly process. Therefore, in this example, the "number of classes" is "7".

[0064] Returning to the explanation of FIG. 5, the "score" is information about a value that quantitatively evaluates the trained model and represents the reliability of the classification of the assembly process video into video segments of each task performed by the trained model. In this embodiment, the "score" is calculated based on the convergence value of a loss function calculated during machine learning when creating the trained model. Note that values ​​calculated by other methods may also be used as the "score."

[0065] "Number of inferences" is information about the number of inferences performed using the trained model, i.e., the number of times the trained model was actually used to classify the assembly process video into video segments for each task.

[0066] "Video tag information" is tag information that was attached to the data of the training videos used for machine learning when creating the trained model. The training videos of the assembly process stored in the storage device 203 are attached with, for example, the name of the worker who performed the assembly process and information about the worker's dominant hand, as well as observation information such as the configuration and observation magnification of the microscope 100 used to capture the video. "Video tag information" shows all of this tag information that was attached to each of the training videos used for machine learning.

[0067] "Purpose ((any word))" is text information that expresses information about the creation of the trained model, such as the design intent and history of the trained model, and is information entered by the model developer who created or updated the trained model.

[0068] "Creation date and time" is information about the date and time when the trained model was created or updated.

[0069] The "last update date and time" is information on the date and time when the design information for the trained model was updated.

[0070] "Updater" is information about the name of the model developer who created or updated the trained model.

[0071] Next, a description will be given of the model-related file information DB 600 stored in the storage device 203. Fig. 7 shows an example of the data structure of the model-related file information DB 600.

[0072] As described above, configuration data of trained models is stored in the storage device 203. Various data files related to the trained models are also stored in this storage device 203. The model-related file information DB 600 is used to manage the association between these data files and trained models.

[0073] In the model-related file information DB600 illustrated in Figure 7, for each trained model, a "revision" as version information indicating the version of the trained model is associated with the name of the data file stored in the storage device 203.

[0074] "Video file name" is information about the file name of the video data file of the training video used in the machine learning performed when creating the trained model.

[0075] The "annotation file name" is the file name information of the annotation file for the training video indicated by the "video file name" used in the machine learning performed when creating the trained model. The annotation file is a data file that stores information (annotation information) indicating the annotations added to the training video during the annotation process. The annotation information includes information indicating the division positions when the training video is divided into multiple video segments by task, and information indicating the classification by task assigned to each video segment. The annotation file stores annotation information for each revision that indicates the version of the annotation information, and the annotation information for each revision also includes information such as the revision of the annotation information and the annotation work date.

[0076] The "classification criteria file name" is the file name information of a data file that stores classification criteria information indicating the criteria for assigning classifications by task to each video segment when annotations are assigned to each of the training videos indicated by the "video file name" used in the machine learning performed when creating the trained model.

[0077] In the model-related file information DB 600 illustrated in FIG. 7, annotation files are managed for each revision of the trained model. Alternatively, annotation files may be managed for each video file. That is, video files and annotation files may be associated one-to-one, and annotation information for the video file may be managed for each revision of the trained model within the corresponding annotation file. Furthermore, annotation information for each revision of the trained model for a video file may be embedded in the video file, and annotations for each revision of the trained model may be managed within the video file.

[0078] Next, various processes performed by the processor 201 will be described.

[0079] [1: Annotation work support processing for re-learning]

[0080] First, the annotation work support process for relearning will be described. Fig. 8 is a flowchart showing the processing contents of an example of the annotation work support process for relearning.

[0081] This annotation work support process begins when the processor 201 receives a process start instruction from the model developer who will be performing the annotation work, via the input device 400. When the process begins, first, in S101, a process is performed in which the model selection screen 700 shown in Fig. 9 is displayed on the monitor 300 connected to the input / output interface 207.

[0082] Here, a description will be given of the model selection screen 700 in Fig. 9. On the model selection screen 700, the information of "model name", "date", "dataset", and "AI model" is associated with each other.

[0083] "Model name" is the name of the AI ​​model whose configuration data is stored in the storage device 203, and "date" is the creation date of the AI ​​model. This information is obtained from the model design information DB 500 described above and displayed on the model selection screen 700.

[0084] "Dataset" represents the number of training videos planned to be used for machine learning when creating the AI ​​model, and the number of training videos actually used for that machine learning. If the numbers on both sides of the diagonal line in "Dataset" are the same, it means that all of the planned training videos were used for machine learning.

[0085] Furthermore, "AI model" indicates the creation status of the trained model, and "created" indicates that creation of the AI ​​model has already been completed. In the example of Figure 9, all "AI models" are "created," so creation work has been completed for all AI models whose model names are displayed on the model selection screen 700 in Figure 9.

[0086] The processor 201 generates the information displayed on the model selection screen 700 using the information shown in the model design information DB 500 and the model-related file information DB 600.

[0087] Returning to the explanation of Fig. 8, once the model selection screen 700 is displayed on the monitor 300, a process of acquiring an instruction operation for the input device 400 is then performed in S102. Then, a process of determining whether the acquired instruction operation is an operation for selecting the model name of an AI model is performed in S103.

[0088] 9, model selection screen 700 is arranged with model selection buttons 710, each showing the name of an AI model as a "model name." The model selection buttons 710 are icon buttons, and a click on the model selection button 710 is detected as an operation to select the model name of an AI model.

[0089] If it is determined in the determination process of S103 that the acquired operation is for selecting a model name, annotation work screen processing is performed in S104. The annotation work screen processing is processing for switching the display screen on the monitor 300 from the model selection screen 700 to the annotation work screen 800, which will be described later. Details of this processing will be described later.

[0090] Thereafter, when the annotation work screen processing is completed, the process returns to S101, and the process of displaying the model selection screen 700 is performed again.

[0091] On the other hand, if it is determined in the determination process of S103 that the acquired instruction operation was not a model name selection operation, a process is performed in S105 to determine whether the instruction operation acquired in S102 was a dataset selection operation.

[0092] In the model selection screen 700 illustrated in FIG. 9, a mouse pointer 720 points to the display position of the dataset for "Trained Model 1," and the movement of the mouse pointer 720 to this position is detected as an operation to select a dataset.

[0093] If it is determined in the determination process of S105 that the acquired instruction operation is an operation for selecting a data set, then in S106, a process is performed in which a moving image list screen 730 is popped up on the monitor 300 that is displaying the model selection screen 700. On the other hand, if it is determined in the determination process of S105 that the acquired instruction operation is not an operation for selecting a data set, then the process returns to S102, and the process of acquiring an instruction operation is performed again.

[0094] The video list screen 730 is a screen that displays a list of information about the training videos used in machine learning when generating the AI ​​model identified by the model name corresponding to the selected dataset. In the example of Figure 9, this information includes the creator's name ("ID"), creation date ("Date"), annotation information revision ("Rev"), and number of class classifications ("Number of Classes") for the training video. This information is included in the tag information attached to the training video or the annotation file for the training video.

[0095] Following the process of S106, in S107, a process is performed to determine whether the selection operation of the data set acquired in the process of S102 has been completed. This determination process is repeated until it is determined that the selection operation has been completed, and when it is determined that the selection operation has been completed, the process returns to S101, where the pop-up display of the video list screen 730 is terminated and the process of displaying the model selection screen 700 is performed again.

[0096] [2: Annotation work screen processing]

[0097] Next, annotation work screen processing will be described. The annotation work screen processing is processing that is performed as processing of S104 when it is determined that the processor 201 has accepted an operation to select an AI model (trained model) in the determination processing of S103 of the annotation work support processing for re-learning in Fig. 8. Fig. 10 is a flowchart showing the processing contents of an example of annotation work screen processing.

[0098] When the processing of Figure 10 starts, first, in S111, a process is performed to obtain various information about the trained model selected by operation in the determination process of S103 of the annotation work support process for re-learning from the storage device 203.

[0099] By the processing of S111, one or more pieces of version information about the selected trained model and design information corresponding to each of the version information are acquired from the model design information DB 500. In addition, model-related information about the selected trained model is acquired from the model-related file information DB 600. Furthermore, video (learning video) data, annotation information, and classification standard information identified by the file name indicated in the acquired model-related information are acquired from the storage device 203. In other words, the video, annotation information, and classification standard information associated with the selected trained model are acquired from the storage device 203.

[0100] Next, in S112, the various pieces of information acquired in the process of S111 are used to create the annotation work screen 800 shown in FIG.

[0101] Here, an example of the annotation work screen 800 will be described.

[0102] The annotation work screen 800 has a plurality of display areas 810, 820, 830, 840, and 850.

[0103] Processor 201 causes the name of the selected trained model to be displayed in display area 810 by the processing of S112. Furthermore, processor 201 causes display area 820 to display a list of each training video used during training of the selected trained model (or the latest revision of the trained model if there are multiple revisions of the trained model). In this list display, one of the displayed training videos is displayed in a selected manner (in the example of FIG. 11, the video file name is displayed in black text on a white background). Processor 201 causes display area 830 to display the training video displayed in the selected manner. Furthermore, processor 201 causes display areas 840 and 850 to display annotation information about the training video displayed in the selected manner (or the latest revision of the annotation information if there are multiple revisions of the annotation information). More specifically, when the learning video displayed in the selected mode is divided into multiple video segments by task, marks 841 indicating each division position are superimposed on a timeline 842 of the learning video and displayed in display area 840. Furthermore, task classifications assigned to each video segment are displayed in a list in display area 850. In the example of FIG. 11, "Task 1: No specimen" is displayed as the task classification assigned to the first video segment of the learning video displayed in the selected mode, and "Task 2: Part placement" is displayed as the task classification assigned to the second video segment. In this way, processor 201 associates the learning video displayed in the selected mode with annotation information about the learning video and displays them on monitor 300.

[0104] In addition, when a click operation is performed on an unselected learning video in display area 820, the selection of learning video is changed, and processor 201 changes the display in display areas 830, 840, and 850 according to the newly selected learning video.

[0105] Returning to the explanation of Fig. 10, once the annotation work screen 800 is displayed on the monitor 300 by the processing of S112, next, processing is performed in S113 to acquire an instruction operation for the input device 400. Then, processing is performed in S114 to determine whether the instruction operation acquired by the processing of S113 was a click operation on the back button 863 of the annotation work screen 800. If it is determined in this determination processing that the instruction operation was a click operation on the back button 860, this annotation work screen processing is terminated, and processing is returned to the annotation work support processing for relearning in Fig. 8.

[0106] On the other hand, if it is determined in the determination process of S114 that the acquired instruction operation is not a click operation on the back button 863, then in S115, various processes corresponding to the acquired instruction operation are executed. Details of these processes will be described later. Then, when the process of S115 ends, the process returns to S113, and the process of acquiring the instruction operation is executed again.

[0107] The above processing is the annotation work screen processing.

[0108] In the following explanation, the main processing performed as the processing of S115 in the annotation work screen processing will be explained.

[0109] [3: Video segment screen processing]

[0110] First, the video segment screen processing will be described with reference to a flowchart of FIG.

[0111] The video segment screen processing is a process executed by the processor 201 as the process of S115 when the instruction operation acquired by the process of S113 in the annotation work screen processing of Figure 10 is an operation to select one of the categories listed in the display area 850.

[0112] In the display area 850 of the annotation work screen 800 illustrated in FIG. 11, the mouse pointer 865 points to the display area for the category "Task 1: No specimen," and the movement of the mouse pointer 865 to this position is detected as an operation to select the category "Task 1: No specimen."

[0113] When the processing of FIG. 12 starts, first, in S121, a process is performed in which a video segment screen 870 is displayed as a pop-up on the monitor 300 on which the annotation work screen 800 is being displayed.

[0114] The video segment screen 870 is a screen that displays a frame image 871 (for example, the last frame image of the video segment) that marks a turning point in the work in the video segment to which the classification selected in the display area 850 has been assigned, the name of the model developer ("ID"), the annotation work date ("Date"), and the revision of the annotation information ("Rev") Here, the name of the model developer is included in the design information, and the annotation work date and the revision of the annotation information are included in the annotation information.

[0115] In addition, popping up the video segment screen 870 on the monitor 300 while the annotation work screen 800 is displayed also means displaying a portion of the design information in association with the learning video being displayed in the selected manner.

[0116] When the video segment screen 870 is popped up on the monitor 300 by the process of S121, a process of acquiring an instruction operation on the input device 400 is then performed in S122. Then, a process of determining whether the acquired instruction operation is a click operation on the revision change button 872a or 872b of the video segment screen 870 is performed in S123.

[0117] If it is determined in the determination process of S123 that the acquired instruction operation is a click operation on the revision change button 872a or 872b, then in S124, a process is performed to change the display contents (frame image 871, annotation work date, and revision of annotation information) of the video segment screen 870 to that of the annotation information of the revision corresponding to the revision change button 872a or 872b that was clicked. More specifically, if the acquired instruction operation is a click operation on the revision change button 872a, a process is performed to change the display contents of the video segment screen 870 to that of the annotation information of the previous (or next) revision. On the other hand, if the acquired instruction operation is a click operation on the revision change button 872b, a process is performed to change the display contents of the video segment screen 870 to that of the annotation information of the next (or previous) revision. Then, when the process of S124 ends, the process returns to S122, and a process is performed to acquire an instruction operation on the input device 400 again.

[0118] If it is determined in the determination process of S123 that the acquired instruction operation was not a click operation on the revision change button 872a or 872b, then in S125, a process is performed to determine whether the instruction operation acquired by the process of S122 was a click operation on the close button 873 of the video segment screen 870.

[0119] If it is determined in the determination process of S125 that the acquired instruction operation is not a click operation on the close button 873, the process returns to S122, and a process of acquiring an instruction operation on the input device 400 again is performed.

[0120] On the other hand, if it is determined in the determination process of S125 that the acquired instruction operation is a click operation on the close button 873, then in S126, a process is performed to hide the pop-up displayed video segment screen 870. Then, when the process of S126 ends, this video segment screen process ends, and the process returns to the annotation work screen process of FIG.

[0121] With this type of video segment screen processing, the model developer performing the annotation work can check information such as at what point the video segment was divided and who performed the annotation work and when for each revision of the annotation information.

[0122] [4: Video segment playback processing]

[0123] Next, the video segment playback process will be described. The video segment playback process is executed by processor 201 as the process of S115 when the instruction operation acquired by the process of S113 in the annotation work screen process of Fig. 10 is a click operation on any of the display areas of the categories displayed in the list in display area 850.

[0124] In the video segment playback process, a video segment that is assigned the classification of the display area where the click operation was performed, or a part of that video segment, is played and displayed in display area 830. Then, when this process ends, the process returns to S113, and a process of acquiring a new instruction operation on input device 400 is performed.

[0125] Such video segment playback processing allows the model developer performing the annotation work to check the content of the video segment or part of the content thereof.

[0126] [5: Classification criteria information screen processing]

[0127] Next, the classification criteria information screen processing will be described. The classification criteria information screen processing is processing executed by the processor 201 as processing of S115 when the instruction operation acquired by the processing of S113 in the annotation work screen processing of Fig. 10 is a click operation on the process list icon 851 included in the annotation work screen 800. Fig. 13 is a flowchart showing the processing contents of an example of the classification criteria information screen processing.

[0128] When the processing of FIG. 13 starts, first, in S131, a processing is performed in which a classification criteria information screen 880 is displayed as a pop-up on the monitor 300 displaying the annotation work screen 800 exemplified in FIG.

[0129] The classification criteria information screen 880 displays the classification criteria information acquired by the processing of S111 in the annotation work screen processing of FIG. 10. In the example of FIG. 14, the classification criteria information is shown as a flowchart illustrating the classification procedure. By referring to this classification criteria information screen 880, a model developer performing annotation work can understand the criteria by which the annotation information displayed on the annotation work screen 800 (display areas 840, 850) divided and classified the selected training video into multiple video segments. For example, the model developer can understand that, in the selected training video, a video segment that shows a specimen and the task of component placement is classified as the task of "component placement." Therefore, the model developer can easily perform annotation work for re-learning by referring to the classification criteria information displayed on the classification criteria information screen 880.

[0130] When the classification criteria information screen 880 is popped up and displayed on the monitor 300 by the process of S131, next, in S132, a process of acquiring an instruction operation to the input device 400 is performed. Then, in S133, a process of determining whether the acquired instruction operation is a click operation on the edit button 881 of the classification criteria information screen 880 is performed.

[0131] If it is determined in the determination process of S133 that the acquired instruction operation is a click operation on the edit button 881, then in S134, editing process of the classification standard information is performed. This process is a process of editing the classification standard information displayed on the classification standard information screen 880, such as adding, deleting, or correcting a classification procedure, in accordance with an operation on the input device 400. This process allows the model developer performing the annotation work to edit the classification standard information. When the process of S134 ends, the process returns to S132, and a process of acquiring an instruction operation on the input device 400 again is performed.

[0132] On the other hand, if it is determined in the determination process of S133 that the acquired instruction operation was not a click operation on the edit button 881, then in S135, a process is performed to determine whether the instruction operation acquired by the process of S132 was a click operation on the display area of ​​any of the classification determination steps shown on the classification criteria information screen 880 (for example, the classification determination step ``Part Placed?'').

[0133] If it is determined in the determination process of S135 that the acquired instruction operation is a click operation on one of the display areas of the classification determination steps, then in S136 a comment addition process is performed. This process is a process of displaying a comment display field near the classification determination step in the display area where the click operation was performed, and displaying a comment entered by an operation on the input device 400 in the comment display field. This process allows the model developer performing the annotation work to add a comment to the classification determination step shown on the classification standard information screen 880. When the process of S136 ends, the process returns to S132, and a process of acquiring an instruction operation on the input device 400 is performed again.

[0134] On the other hand, if it is determined in the judgment process of S135 that the acquired instruction operation was not a click operation on any of the display areas of the classification judgment step, then in S137, a process is performed to determine whether the instruction operation acquired by the process of S132 was a click operation on the save button 882 of the classification criteria information screen 880.

[0135] If it is determined in the determination process of S137 that the acquired instruction operation is a click operation on the save button 882, then a save process is performed in S138. This process is a process of saving a classification standard file that stores the classification standard information displayed on the classification standard information screen 880 (including the comments if a comment display field is displayed) in the storage device 203. This process allows the model developer performing the annotation work to save the classification standard information to which edits and comments have been added. Note that when a trained model is subsequently retrained and a new trained model is created, the classification standard file saved in the storage device 203 at this time is associated with the newly created trained model by the model-related file information DB 600. When the process of S138 ends, the process returns to S132, and a process of newly acquiring an instruction operation for the input device 400 is performed.

[0136] On the other hand, if it is determined in the judgment process of S137 that the acquired instruction operation was not a click operation on the save button 882, then in S139, a process is performed to determine whether the instruction operation acquired by the process of S132 was a click operation on the close button 883 of the classification criteria information screen 880.

[0137] If it is determined in the determination process of S139 that the acquired instruction operation is not a click operation on the close button 883, the process returns to S132, and a process of acquiring an instruction operation on the input device 400 again is performed.

[0138] On the other hand, if it is determined in the determination process of S139 that the acquired instruction operation is a click operation on the close button 883, then in S140, a process is performed to hide the classification standard information screen 880 that is displayed as a pop-up. Then, when the process of S140 ends, this classification standard information screen process ends, and the process returns to the annotation work screen process of FIG.

[0139] [6: First annotation information editing process]

[0140] Next, the first annotation information editing process will be described with reference to Fig. 15, which is a flowchart showing an example of the first annotation information editing process.

[0141] The first annotation information editing process is a process executed by the processor 201 as the process of S115 when the instruction operation acquired by the process of S113 in the annotation work screen process of Figure 10 is an operation to modify the position of the boundary between video segments.

[0142] In the annotation work screen 800 shown in FIG. 11, an operation of dragging and dropping a mark 841 on a timeline 842 in the left-right direction is detected as an operation of correcting the position of the boundary between video segments.

[0143] When the processing of FIG. 15 starts, first, in S141, processing is performed to correct the position of the boundary between two adjacent video segments that sandwich the mark 841, depending on the position of the mark 841 where the drag-and-drop operation was performed.

[0144] Next, in S142, a process is performed to calculate the similarity between the frame image immediately before (or immediately after) the position of the mark 841 where the drag-and-drop operation was performed and the frame image immediately before (or immediately after) the position of the mark 841 before the drag-and-drop operation was performed. Then, in S143, a process is performed to determine whether the calculated similarity is equal to or greater than a threshold value.

[0145] If it is determined in the determination process of S143 that the similarity is not equal to or greater than the threshold, then in S144 an alert is issued to prompt the user to confirm whether the correction of the boundary position between the video segments is appropriate. This process issues an alert if the content of the frame images before and after the correction of the boundary position between the video segments is significantly different.

[0146] 16, when the position of the boundary between the video segment classified as "Task 1: no specimen" and the video segment classified as "Task 2: component placement" is corrected, an alert message 843 is displayed as an alert. The alert may be audio or may be both audio and a message.

[0147] On the other hand, if it is determined in the determination process of S143 that the similarity is equal to or greater than the threshold value, the first annotation information editing process ends, and the process returns to the annotation work screen process of FIG.

[0148] According to this first annotation information editing process, the model developer who performs the annotation work can correct the division positions of the learning video. That is, the model developer can correct the annotation. Furthermore, if the division positions may be inappropriately corrected, the model developer can be alerted.

[0149] [7: Second annotation information editing process]

[0150] Next, the second annotation information editing process will be described with reference to Fig. 17, which is a flowchart showing an example of the second annotation information editing process.

[0151] The second annotation information editing process is a process executed by the processor 201 as the process of S115 when the instruction operation acquired by the process of S113 in the annotation work screen processing of Figure 10 is either an operation to further divide a video segment or an operation to delete one of the further divided video segments.

[0152] When the processing in FIG. 17 starts, first, in S151, a process is performed in which it is determined whether the instruction operation acquired in the processing in S113 is an operation for further dividing a video segment.

[0153] In the annotation work screen 800 illustrated in FIG. 11, a click operation on any position on the timeline 842 is detected as an operation to further divide a video segment.

[0154] If it is determined in the determination process of S151 that the acquired instruction operation is an operation to further divide the video segment, then in S152, a process to further divide the video segment is performed. In this process, the video segment at the position on timeline 842 where the click operation was performed is further divided at that position. In addition, a process to change the display content of display area 850 is performed accordingly.

[0155] 11, for example, when a click operation is performed on any position on timeline 842 for a video segment classified as "Task2: component placement," the video segment of "Task2: component placement" is further divided at that position. In response to this, in display area 850, as shown in FIG. 18, the classification of "Task2: component placement" before the division is displayed as being divided into classifications of "Task2: component placement_01" and "Task2: component placement_02" after the division.

[0156] When the process of S152 ends, the second annotation information editing process ends, and the process returns to the annotation work screen process of FIG.

[0157] On the other hand, if it is determined in the judgment process of S151 that the instruction operation acquired by the process of S113 was not an operation to further divide the video segment, then in S153, a process is performed to determine whether the instruction operation acquired by the process of S113 was an operation to delete any of the further divided video segments.

[0158] 18, an operation such as a double-click or right-click on the display area of ​​either "Task2: component placement_01" or "Task2: component placement_02," which are classifications of the further divided video segments, is detected as an operation to delete one of the further divided video segments. For example, when a right-click operation is performed, an operation to select (click) a deletion instruction item from the menu screen that is popped up by that operation is detected as an operation to delete one of the further divided video segments.

[0159] If it is determined in the determination process of S153 that the acquired instruction operation is an operation to delete one of the further divided video segments, then in S154, a process is performed to restore the further divided video segments to their original state. In this process, the further divided video segments are integrated to restore the original (pre-division) video segments. In response to this, a process is also performed to restore the display content of display area 850 to the display content before division.

[0160] In the divided display area 850 shown in Fig. 18, when a double-click operation is performed on the display area of ​​"Task2: component placement_01" or "Task2: component placement_02," the video segment of "Task2: component placement_01" and the video segment of "Task2: component placement_02" are integrated and restored to the original video segment of "Task2: component placement." In response to this, the display in the divided display area 850 returns to the display in the pre-division display area 850 shown in Fig. 18.

[0161] When the process of S154 ends, the second annotation information editing process ends, and the process returns to the annotation work screen process of FIG.

[0162] On the other hand, if it is determined in the judgment process of S153 that the acquired instruction operation is not an operation to delete one of the further divided video segments, this second annotation information editing process is terminated and the process returns to the annotation work screen process of Figure 10.

[0163] According to this second annotation information editing process, the model developer performing the annotation work can further divide a video segment or return a further divided video segment to the original video segment. In other words, the model developer can add or delete annotations.

[0164] [8: Annotation information saving process]

[0165] Next, the annotation information saving process will be described. The annotation information saving process is executed by the processor 201 as the process of S115 when the instruction operation acquired by the process of S113 in the annotation work screen process of Fig. 10 is a click operation on the save button 862 of the annotation work screen 800.

[0166] In the annotation information saving process, the annotation information displayed in the display areas 840 and 850 of the annotation work screen 800 is saved in the storage device 203. More specifically, the annotation information is stored as a new revision of the annotation information in the annotation file associated with the learning video displayed in the selected format. Then, when the annotation information saving process is completed, the process returns to the annotation work screen process of FIG. 10.

[0167] [9: Video tag information screen processing]

[0168] Next, the video tag information screen process will be described with reference to a flowchart of FIG.

[0169] The video tag information screen processing is executed by the processor 201 as processing of S115 when the instruction operation acquired by processing of S113 in the annotation work screen processing of Figure 10 is a click operation on the tag information button 861 of the annotation work screen 800.

[0170] When the processing of FIG. 19 is started, first, in S161, processing is performed to display the moving image tag information screen 900 exemplified in FIG. 20 on the monitor 300.

[0171] Here, the moving image tag information screen 900 shown in FIG. 20 will be described.

[0172] The video tag information screen 900 is a screen that displays, for each training video used in machine learning when creating a trained model, the tag information attached to the training video data in association with the training video. By referring to this video tag information screen 900, the model developer can easily understand the situation when the training video was acquired.

[0173] The moving image tag information screen 900 has a design information display area 910 , a tag information list display area 920 , and a tag information selection area 930 .

[0174] The design information display area 910 is an area where design information about the selected trained model (if there are multiple revisions of the trained model, the latest revision of the trained model) is displayed.

[0175] The tag information list display area 920 is an area that displays a list of the learning videos used for learning when creating the trained model of the revision whose design information is displayed in the design information display area 910, in association with the tag information attached to each piece of data in the learning videos.

[0176] The tag information selection area 930 is an area for individually selecting tag information from the design information of the trained model of the revision whose design information is displayed in the design information display area 910.

[0177] In the example of FIG. 20, the design information display area 910 displays five items of tag information: "Worker XX," "Right-handed," "Zoom 2X," "Worker YY," and "Left-handed." These are tag information items that were attached to any of the training videos used for training when creating the trained model of the revision whose design information is displayed in the design information display area 910. These five items are displayed in the tag information selection area 930.

[0178] When any of these items displayed in the tag information selection area 930 is clicked, the item clicked on is displayed in inverted display mode (white characters representing the item are displayed on a black background). At this time, the tag information of the same item displayed in association with the learning video in the tag information list display area 920 also changes to inverted display mode.

[0179] In the example of Fig. 20, three of the five items displayed in the tag information selection area 930, "Worker XX," "Right-handed," and "Zoom 2X," are displayed in inverted mode, indicating that these three items have been selected. Fig. 20 also shows that, as a result of this selection, the tag information of "Worker XX," "Right-handed," and "Zoom 2X," which are displayed in association with the learning video in the tag information list display area 920, are now displayed in inverted mode.

[0180] As described above, the video tag information screen 900 displays the selected tag information in association with the learning videos related to that tag information in response to the selection of tag information received in the tag information selection area 930. Displaying such a video tag information screen 900 on the monitor 300 can provide the model developer with information to help them select learning videos to use in re-learning to update a trained model.

[0181] Returning to the explanation of Figure 19, in the processing of S161, first, design information for the selected trained model (if there are multiple revisions of the trained model, the latest revision of the trained model) is acquired. Then, the display of the design information display area 910 is created using the acquired design information. At this time, the video file of each training video used for training when creating the selected trained model is identified by referring to the model-related file information DB600. Then, tag information is acquired from the video file, and the display of the tag information list display area 920 is created by associating the acquired tag information with the training video for each training video. Furthermore, the display of the tag information selection area 930 is created using the tag information included in the design information displayed in the design information display area 910. The processor 201 displays the video tag information screen 900, in which the display of each area has been created in this manner, on the monitor 300.

[0182] Next, in S162, a process is performed to acquire an instruction operation on the input device 400. Then, in S163, a process is performed to determine whether the acquired instruction operation is a click operation on any of the tag information displayed in the tag information selection area 930.

[0183] If it is determined in the process of S163 that the instruction operation is a click operation on tag information, then in S164, the same tag information as that clicked is displayed in reverse video in the tag information list display area 920 and the tag information selection area 930. Thereafter, the process returns to S162, and the process of acquiring the instruction operation on the input device 400 continues.

[0184] On the other hand, if it is determined in the process of S163 that the instruction operation was not a click operation on tag information, then in S165, a process is performed to determine whether or not the instruction operation was a click operation on the back button 940 of the video tag information screen 900. If it is determined in this determination process that the instruction operation was a click operation on the back button 940, then the video tag information screen process is terminated, and the process returns to the original process, the annotation work screen process of Fig. 10.

[0185] On the other hand, if it is determined in the determination process of S165 that the instruction operation is not a click operation on the back button 940, the process returns to S162, and the instruction operation on the input device 400 is acquired again.

[0186] [10: Re-learning process]

[0187] Next, the re-learning process will be described with reference to the flowchart of FIG.

[0188] The re-learning process is executed by the processor 201 as processing of S115 when the instruction operation acquired by processing of S113 in the annotation work screen processing of Figure 10 is a click operation on the AI ​​model creation button 864 on the annotation work screen 800.

[0189] The model developer performs the work in S21 of Figure 2 to understand the design intent and process behind the creation of the previous version of the AI ​​model. The model developer then performs the review work in S22 and performs the work from S23 onwards based on the results of that review. The re-learning process is the process for the work from S23 onwards.

[0190] 21 starts, first, in S171, a process of acquiring an instruction operation on the input device 400 is performed. Then, in the subsequent determination processes of S172, S174, S176, and S178, a process of determining the instruction content indicated by the instruction operation is performed.

[0191] When it is determined in the determination process of S172 that the instruction operation indicates an instruction to acquire a video, a learning video acquisition process is performed in S173. This process is a process for acquiring a learning video, and is a process for the process of acquiring additional learning videos, which is the process of S24 in the AI ​​model update process shown in Figure 2.

[0192] If it is determined in the determination process of S174 that the instruction operation indicates an instruction to perform annotation, annotation process of S175 is performed. This process is a process of adding annotations to the learning video, and is a process for adding annotations to the learning video or changing annotations that have already been added, which is the process of S26 in the AI ​​model update process shown in Figure 2. Note that this process includes the annotation work support process for re-learning shown in Figure 8.

[0193] If it is determined in the determination process of S176 that the instruction operation indicates an instruction to set learning conditions, the process proceeds to a learning condition setting process of S177. This process sets learning conditions in machine learning for creating an AI model, and is a process for resetting the learning conditions, which is the process of S28 in the AI ​​model update process shown in Figure 2.

[0194] When the process of S173, S175, or S177 described above is completed, the process returns to S171, and the process of acquiring a new instruction operation is performed again.

[0195] On the other hand, if it is determined in the determination process of S178 that the instruction operation indicates an instruction to execute machine learning, the machine learning process of S179 is performed. This process is a process of performing machine learning for creating an AI model according to the set learning conditions, and is a process for the work of each procedure from S15 to S20 in Figure 1 as re-learning in the update work of the AI ​​model.

[0196] Then, when the machine learning process of S179 is completed, a process of S180 is performed in which configuration data for the re-trained trained model created by the machine learning process is saved in the storage device 203. Then, in the following S181, a process of storing design information for the re-trained trained model in the model design information DB 500 of the storage device 203 in association with a revision, which is version information indicating the version of the re-trained trained model. Also in S181, a process of registering the re-trained trained model and various data files associated therewith (learning video file, annotation file, classification standard file) in the model-related file information DB 600 of the storage device 203 in association with the revision of the re-trained trained model.

[0197] After the process of S181 is completed, this re-learning process is terminated, and the process proceeds to the annotation work support process for re-learning in FIG. 8, where the process of S101, which is the process of displaying the model selection screen 700, is performed.

[0198] If the instruction content indicated by the instruction operation cannot be determined by any of the determination processes of S172, S174, S176, and S178, the process returns to S171, and the process of acquiring a new instruction operation is performed again.

[0199] The above processing is the re-learning processing.

[0200] As described above, the annotation support system 1 is configured to associate and present annotation information associated with a trained model with a learning video, and to present classification criteria information associated with the trained model. This makes it easier to understand information related to previously performed annotation work. Because the annotation support system 1 is configured in this way, it can support annotation work for relearning a trained model, and model developers who perform annotation work can easily perform this annotation work.

[0201] The above-described embodiments are illustrative examples provided to facilitate understanding of the invention, and the present invention is not limited to these embodiments. Modifications of the above-described embodiments and alternatives to the above-described embodiments may be included. In other words, the components of the above-described embodiments may be modified without departing from the spirit and scope of the present invention. Furthermore, new embodiments can be implemented by appropriately combining multiple components disclosed in the embodiments. Furthermore, some components may be deleted from the components shown in the embodiments, or some components may be added to the components shown in the embodiments. Furthermore, the order of the processing steps shown in the embodiments may be reversed as long as there is no contradiction. In other words, the annotation support system of the present invention can be modified and changed in various ways without departing from the scope of the claims.

[0202] For example, in the above-described embodiment, the classification criteria information screen 880 displayed as a pop-up on the monitor 300 while the annotation work screen 800 shown in Fig. 14 is displayed shows the classification criteria information in the form of a flowchart. Alternatively, the classification criteria information may be displayed in another form. Fig. 22 shows another example of the classification criteria information screen 880.

[0203] The classification criteria information screen 880 shown in Fig. 22 has display areas 884 and 885. Display area 884 is an area where the classification of each video segment is displayed. Display area 885 is an area where the tasks included in the classification of each video segment are displayed. The example of Fig. 22 shows that the tasks included in the classification of "Installation on jig" are "Placing on jig" and "Lock jig."

[0204] By displaying the classification criteria information screen 880 illustrated in Figure 22, the model developer performing the annotation work can confirm the specific tasks that were included in the classification of each video segment.

[0205] The display area 884 further has a delete category button 8841, an add category button 8842, and a demote button 8843, corresponding to each displayed category. When these buttons are clicked, the processor 201 performs the following processing. When the delete category button 8841 is clicked, a process of deleting the corresponding category is performed. When the add category button 8842 is clicked, a process of adding a new category as a category that is the previous or next step of the corresponding category is performed. When the demote button 8843 is clicked, a process of demoting the corresponding category to a task is performed. More specifically, a process of making the task included in the corresponding category a task that is performed after the last task included in the category that is the previous step of the relevant category, or a process of making the task a task that is performed before the first task included in the category that is the next step of the relevant category is performed.

[0206] The display area 885 further has a delete operation button 8851, an add operation button 8852, and a promote button 8853 corresponding to each displayed operation. When these buttons are clicked, the processor 201 performs the following processing. When the delete operation button 8851 is clicked, a process of deleting the corresponding operation is performed. When the add operation button 8852 is clicked, a process of adding a new operation as an operation before or after the corresponding operation is performed. When the promote button 8853 is clicked, a process of promoting the corresponding operation to a classification is performed. More specifically, a process of classifying the corresponding operation as a subsequent process in the classification to which the operation belongs is performed. For example, as illustrated in FIG. 23 , when the promote button 8853 corresponding to the operation "foam discharge" included in the classification "parts assembly" is clicked, the operation "foam discharge" is classified as "foam discharge," which is a subsequent process in the classification "parts assembly." In this case, since the task of "discharging foam" is one task, the only task included in the promoted classification of "discharging foam" is the task of "discharging foam."

[0207] 22, a model developer performing annotation work can edit the classification criteria information by clicking each of the buttons: Delete Classification button 8841, Add Classification button 8842, Demote button 8843, Delete Task button 8851, Add Task button 8852, and Promote button 8853. Editing the classification criteria information in this way is effective, for example, when something that has been managed as a classification is desired to be managed as a task included in a classification, or when something that has been managed as a task included in a classification is desired to be managed as a classification.

[0208] Also, for example, in the above-described embodiment, the video segment screen 870 (see FIG. 11) and the classification criteria information screen 880 (see FIG. 14) that are popped up on the monitor 300 while the annotation work screen 800 is being displayed may be configured to have their display size changed in response to an instruction operation on the input device 400.

[0209] Also, for example, in the above-described embodiment, the configuration data of the trained model, the model design information DB 500, and the model-related file information DB 600 are stored separately in the storage device 203 of the computer 200a serving as the control device 200. Alternatively, the design information stored in the model design information DB 500 and the information on various file names stored in the model-related file information DB 600 may be embedded in the configuration data of the corresponding version of the trained model and stored separately in the storage device 203.

[0210] In this specification, the expression "based on A" does not mean "based only on A," but rather "based at least on A," and further means "based at least partially on A." In other words, "based on A" may be based on B in addition to A, or may be based on a part of A. [Explanation of symbols]

[0211] 1 Annotation support system 100 microscopes 101 Objective Lens 106 Eyepiece 112 Imaging device 113 Projector 120 Eyepiece tube 121 Operation section 130 Zoom Handle 140 Aiming Handle 200 control device 200a Computer 201 processor 202 memory 203 Storage device 204 Reading device 205 Removable storage media 206 Communication Interface 207 Input / Output Interface 208 Bus 300 monitors 400 Input Device 401 Mouse 402 Keyboard 403 Foot Switch 404 Barcode reader 500 Model Design Information DB 600 model related file information DB 700 Model Selection Screen 710 Model selection button 720 Mouse Pointer 730 Video list screen 800 Annotation work screen 810, 820, 830, 840, 850 display area 841 mark 842 Timeline 843 Alert Message 851 Process List Icon 861 Tag information button 862 Save button 863 Back button 864 AI model creation button 865 Mouse Pointer 870 Video Segment Screen 871 frame images 872a, 872b Revision change button 873 Close button 880 Classification criteria information screen 881 Edit button 882 Save button 883 Close button 884, 885 display area 900 Video tag information screen 910 Design information display area 920 Tag information list display area 930 Tag information selection area 940 Back button 8841 Category Delete Button 8842 Add Category button 8843 Demote button 8851 Delete Work Button 8852 Add Task Button 8853 Promotion Button

Claims

1. At least one trained model trained using at least one microscope video; The microscope video associated with each of the trained models; and Annotation information indicating annotations assigned to the microscope video, which is associated with each of the trained models, and which includes a classification for each video segment included in the microscope video; and Classification criterion information corresponding to each of the trained models and indicating a criterion for assigning the classification; a storage unit that stores the Accepting a selection of at least one of the trained models; Acquire the microscope video, the annotation information, and the classification standard information associated with the selected trained model from the storage unit; displaying the acquired microscope video and the annotation information on a display in association with each other; The acquired classification standard information is displayed on the display. A control unit; An annotation work support system comprising:

2. The storage unit further stores design information associated with each of the trained models, The control unit further causes a part of the design information associated with the selected trained model to be associated with the microscope video and displayed on the display.

2. The annotation work support system according to claim 1, wherein:

3. the design information includes tag information, The storage unit further stores the tag information in association with the microscope video; The control unit further Accepting a selection of the tag information; The selected tag information and the microscope video associated with the selected tag information are displayed on the display in association with each other.

3. The annotation work support system according to claim 2, wherein:

4. The annotation work support system according to claim 1, wherein the control unit accepts the selection of at least one of the trained models using a selection screen displayed on the display.

5. The annotation work support system according to claim 1 , wherein the control unit causes the display to display the classifications included in the annotation information as a list.

6. The control unit further Accepting a designation of the classification in the list; Displaying the video segments to which the designated classification is assigned on the display.

6. The annotation work support system according to claim 5, wherein:

7. the control unit further receives addition, deletion, or modification of the annotation in the displayed annotation information; The storage unit further stores the updated annotation information.

2. The annotation work support system according to claim 1, wherein:

8. The annotation work support system according to claim 1 , wherein the classification criteria information includes a procedure for assigning the classification when the annotation is assigned to the microscope video.

9. The control unit further receives addition, deletion, or modification of the classification assignment procedure in the displayed classification standard information, The storage unit further stores the updated classification standard information.

9. The annotation work support system according to claim 8, wherein:

10. Accepting a selection of at least one trained model trained using at least one microscopy video; Acquire the microscope video, the annotation information, and the classification standard information associated with the selected trained model from a storage unit that stores at least one of the trained models, the microscope video associated with each of the trained models, annotation information associated with each of the trained models and indicating annotations assigned to the microscope video, the annotation information including a classification for each video segment included in the microscope video, and classification standard information associated with each of the trained models and indicating a standard for assigning the classification; displaying the acquired microscope video and the annotation information on a display in association with each other; The acquired classification standard information is displayed on the display. The annotation work support method is characterized in that the above steps are performed by a computer.

11. Accepting a selection of at least one trained model trained using at least one microscopy video; Acquire the microscope video, the annotation information, and the classification standard information associated with the selected trained model from a storage unit that stores at least one of the trained models, the microscope video associated with each of the trained models, annotation information associated with each of the trained models and indicating annotations assigned to the microscope video, the annotation information including a classification for each video segment included in the microscope video, and classification standard information associated with each of the trained models and indicating a standard for assigning the classification; displaying the acquired microscope video and the annotation information on a display in association with each other; The acquired classification standard information is displayed on the display. An annotation work support program characterized by causing a computer to perform processing.

Citation Information

Patent Citations

  • Method and device for supporting creation of teaching data, program, and program recording medium

    JP2017009314A