Model generation program, model generation method, and information processing device

By generating two classification models with different similarity thresholds, the information processing device improves task detection accuracy in assembly work by distinguishing between target tasks and similar tasks, reducing false positives.

WO2025203479A1PCT designated stage Publication Date: 2025-10-02FUJITSU LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/012787
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-28
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing classification models struggle to accurately distinguish between similar tasks in assembly work, leading to false detections due to similarities in tool usage and work state transitions.

Method used

The information processing device generates two classification models through machine learning, using first and second image data with varying similarity thresholds to improve task detection accuracy by distinguishing between target tasks and similar tasks.

Benefits of technology

This approach enhances task detection accuracy by differentiating between target tasks and similar tasks, reducing false positives and improving overall detection precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024012787_02102025_PF_FP_ABST
    Figure JP2024012787_02102025_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device acquires first image data obtained by capturing an image of a state in which work to be detected is being performed, second image data in which the degree of similarity to the first image data is equal to or greater than a threshold value, and third image data in which the degree of similarity to the first image data is less than the threshold value. The information processing device executes the following: machine learning of a first machine learning model using the first image data and the second image data as positive example data and the third image data as negative example data; and machine learning of a second machine learning model using the first image data as positive example data and the second image data as negative example data.
Need to check novelty before this filing date? Find Prior Art

Description

Model generation program, model generation method and information processing device

[0001] The present invention relates to a model generation program, a model generation method, and an information processing device.

[0002] Although assembly work is becoming increasingly automated, much manual work remains, and in order to improve the quality and productivity of manual work, it is important to ensure that the work is being carried out in the correct process and to transfer the know-how of experts to less skilled workers.Video analysis is being used as a technology to achieve this.

[0003] Video analysis uses models such as neural networks (NNs) to detect specific tasks from image data or tasks by detecting the body and hand postures associated with the actions, thereby checking for omissions in tasks. These models include models trained using similar training data collected using NNs capable of detecting objects included in a seed dataset, and models trained using manually labeled training data and capable of identifying tasks with assistance from the worker's speech or pointing.

[0004] JP 2020-204800 A JP 2021-114700 A U.S. Patent Application Publication No. 2011 / 0194780 U.S. Patent Application Publication No. 2022 / 0076683

[0005] However, task detection using the above model can sometimes result in a decrease in detection accuracy. For example, when detecting tasks involving tools, tasks involving tools with similar shapes or tasks involving the same hand movements may be falsely detected.

[0006] In one aspect, an object of the present invention is to provide a model generation program, a model generation method, and an information processing device that can improve the accuracy of task detection.

[0007] In the first proposal, the model generation program causes a computer to execute the following processes: acquire first image data capturing an image of a state in which a detection target task is being performed, second image data whose similarity to the first image data is equal to or greater than a threshold, and third image data whose similarity to the first image data is less than a threshold; perform machine learning of a first machine learning model using the first image data and the second image data as positive example data and the third image data as negative example data; and perform machine learning of a second machine learning model using the first image data as positive example data and the second image data as negative example data.

[0008] According to one embodiment, the accuracy of work detection can be improved.

[0009] FIG. 1 is a diagram illustrating an activity recognition system according to a first embodiment. FIG. 2 is a diagram illustrating an activity process to be recognized. FIG. 3 is a diagram illustrating processing by an information processing device according to the first embodiment. FIG. 4 is a functional block diagram illustrating a functional configuration of an information processing device according to the first embodiment. FIG. 5 is a diagram illustrating generation of training data. FIG. 6 is a diagram illustrating an example of generation of training data. FIG. 7 is a diagram illustrating training of a classification model. FIG. 8 is a diagram illustrating classification processing using a classification model after training. FIG. 9 is a flowchart illustrating the flow of processing to generate training data. FIG. 10 is a flowchart illustrating the flow of machine learning processing of a classification model. FIG. 11 is a diagram illustrating evaluation results. FIG. 12 is a diagram illustrating an example of a hardware configuration.

[0010] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS A model generation program, a model generation method, and an information processing device according to the present invention will be described in detail below with reference to the accompanying drawings. However, the present invention is not limited to these embodiments.

[0011] (Overall Configuration) Fig. 1 is a diagram illustrating an operation recognition system according to Example 1. As shown in Fig. 1, in this system, each device in a factory 2 is connected to an information processing device 10 via a network N. Note that the network N can be any of various communication networks such as the Internet or a dedicated line.

[0012] The factory 2 is a factory that produces various products, and cameras 2a are installed in each workplace where workers perform assembly work. Note that the type of factory and the products produced are not limited, and the system can be applied to various fields, such as factories that manufacture processed products, factories that manage the distribution of goods, and automobile factories.

[0013] The camera 2a is set to capture an arbitrary area, such as the entire worker or the area where the worker is working on an object, and records the captured data as video data. The video data is composed of multiple images (frame images) captured by the camera 2a, i.e., a series of frames of a video. The number of factories 2 and cameras 2a is not limited to that shown in FIG. 1.

[0014] The information processing device 10 is an example of a computer that is connected to a camera 2a installed in the factory 2, acquires video data captured by the camera 2a, and performs video analysis. Specifically, the information processing device 10 acquires frame images (image data) from the video data and detects the work of a worker based on the results of inputting each image data into a classification model. For example, the information processing device 10 improves the quality and productivity of manual work by recognizing whether the worker has completed the work correctly and whether the worker performed the work in the correct order.

[0015] Here, various tasks included in assembly work performed by workers will be described. Fig. 2 is a diagram illustrating the work process to be recognized. As shown in Fig. 2, workers perform tasks such as "screw work" in which a screw is inserted into a target object such as an electronic component using a screwdriver, "sealing work" in which a protective film or protective sticker is applied to the target object, and "application work" in which a chemical agent is applied to the target object using a syringe.

[0016] 2, a camera 2a captures images of the work site where a worker performs the work shown in Fig. 2, and an information processing device 10 detects the work of the worker using the image data acquired from the camera 2a and a classification model. Using the results, a manager of a factory or manufacturing control center manages the work of the worker.

[0017] However, it is difficult for a general classification model to distinguish and detect similar tasks. For example, consider the case where a general classification model is used to detect whether or not an application task has been performed correctly. As described in FIG. 2, the "application task" to be detected uses a syringe, which is a sharp tool, and the task state transitions between "before task," in which nothing is applied to the target object, "during task," in which the agent is gradually applied to the application target area of ​​the target object, and "after task," in which the agent is applied to the entire application target area of ​​the target object.

[0018] In order to accurately recognize such "painting work," a classification model is generally trained by extracting features from image data of "painting work" that include "sharp tools" and "transitions in work state from before work to during work to after work." Therefore, even if a work is different from "painting work," if the work is highly likely to belong to the distribution of positive example data in the training data, it may be recognized as a positive example "painting work" by the classification model.

[0019] For example, the "sealing work" shown in Figure 2 has the same work state transition as the "painting work" that is the detection target, but since it is not a work that uses a "sharp tool," the feature amount is different. In other words, the classification model classifies the image data of "sealing work" as a negative example (other than painting work).

[0020] On the other hand, the "screw work" shown in Figure 2 has the same work state transition as the "coating work" that is the detection target, and because it is a work that uses a "sharp tool = screwdriver," the tools themselves are different but the features are similar. In other words, the classification model is likely to classify the image data of "screw work" as a positive example (coating work).

[0021] In this way, a classification model trained using general training data may classify tasks that are different from the task to be detected as tasks to be detected, if they fall within the distribution of positive example data during training, i.e., tasks with similar features.

[0022] Therefore, the information processing device 10 according to the first embodiment aims to improve classification accuracy by generating a classification model through machine learning that clearly distinguishes between the task to be detected and similar tasks that are similar to the task to be detected.

[0023] Specifically, the information processing device 10 acquires first image data capturing an image of a state in which the detection target task is being performed, second image data having a similarity to the first image data equal to or greater than a threshold, and third image data having a similarity to the first image data less than the threshold. The information processing device 10 then performs machine learning of a first classification model using the first image data and the second image data as positive example data and the third image data as negative example data, and performs machine learning of a second classification model using the first image data as positive example data and the second image data as negative example data.

[0024] FIG. 3 is a diagram illustrating the processing of the information processing device 10 according to the first embodiment. An example in which the above-mentioned "application work" is the detection target will be described using FIG. 3 . As shown in FIG. 3 , the information processing device 10 generates a first classification model using training data in which "target data of the detection target (image data of the application work)" and "high similarity data (image data of the screwing work)" that have a high similarity to the target data are used as positive examples, and "low similarity data (image data of the sealing work)" are used as negative examples. Furthermore, the information processing device 10 generates a second classification model using training data in which "target data of the detection target (image data of the application work)" are used as positive examples, and "high similarity data (image data of the screwing work)" that have a high similarity to the target data are used as negative examples.

[0025] In other words, the information processing device 10 generates two classification models with different distributions of positive example data. As a result, the information processing device 10 can use these two classification models in stages to perform detection processing with a distribution of positive example data that is neither too broad nor too narrow, thereby improving the accuracy of task detection.

[0026] 4 is a functional block diagram showing the functional configuration of the information processing device 10 according to Example 1. As shown in FIG. 4, the information processing device 10 includes a communication unit 11, a storage unit 12, and a control unit 20.

[0027] The communication unit 11 is a processing unit that controls communication with other devices, and is realized by, for example, a communication interface, etc. For example, the communication unit 11 receives various data and instructions from an administrator terminal used by an administrator, and outputs the results of work detection, etc. to a designated terminal such as the administrator terminal.

[0028] The storage unit 12 is a processing unit that stores various data and programs executed by the control unit 20, and is realized by, for example, a memory, a processor, etc. The storage unit 12 stores a video data DB 13, an image data DB 14, a training data DB 15, a first classification model 16, and a second classification model 17.

[0029] The video data DB 13 is a database that stores video data captured by the camera 2 a. For example, the video data DB 13 stores video data acquired from the camera 2 a by the control unit 20. If there are multiple cameras, the video data DB 13 may store video data for each camera or for each worker.

[0030] The image data DB 14 is a database that stores image data showing a state in which a worker is working. For example, the image data DB 14 stores each image data (frame image) that constitutes video data obtained by the control unit 20 or manually by an administrator.

[0031] The training data DB 15 is a database that stores training data used for machine learning of the first classification model 16 and the second classification model 17. For example, the training data DB 15 stores training data in which the explanatory variable is "image data" and the objective variable is "(application work) or (other than application work)." Note that each piece of training data stored here is generated by the training data generation unit 31, which will be described later.

[0032] The first classification model 16 is an example of a machine learning model that outputs a prediction result (inference result) as to whether or not the work to be detected is a work to be detected in response to input image data. Explaining the above example, when image data captured by the camera 2a is input, the first classification model 16 outputs a prediction result as to whether or not the work to be detected is "application work." Note that the first classification model 16 is generated by machine learning by the machine learning unit 30.

[0033] Similar to the first classification model 16, the second classification model 17 is an example of a machine learning model that outputs a prediction result of whether or not the work to be detected is a work to be detected in response to input image data. The second classification model 17 differs from the first classification model 16 in that the distribution area for determining positive example data is smaller. In other words, the second classification model 17 performs more precise predictions than the first classification model 16. Explaining the above example, the second classification model 18 outputs a prediction result of whether or not the image data predicted as "application work" by the first classification model 16 is "application work," which is a work to be detected. The second classification model 17 is generated by machine learning by the machine learning unit 30.

[0034] The control unit 20 is a processing unit that controls the entire information processing device 10, and is realized by, for example, a processor. The control unit 20 has a machine learning unit 30 and a classification processing unit 40. The machine learning unit 30 and the classification processing unit 40 are realized by electronic circuits included in the processor, processes executed by the processor, etc.

[0035] The machine learning unit 30 has a training data generation unit 31 and a training execution unit 32, and is a processing unit that generates each classification model by machine learning using the training data.

[0036] The training data generation unit 31 is a processing unit that generates training data to be used for machine learning of each classification model. Specifically, the training data generation unit 31 collects training data using a trained VL (Vision-Language) model and stores the training data in the training data DB 15.

[0037] For example, the training data generation unit 31 identifies image data that shows an object similar to the target object using sentences or keywords that describe the characteristics of the target object. Then, the training data generation unit 31 labels the image data from which the similar object has been extracted with the result of automatic discrimination between positive and negative examples based on annotations of the time from the start to the end of the work. For example, the training data generation unit 31 generates training data in which the target image data is labeled as a "positive example (e.g., painting work)" if the image data was captured during the time period from the start to the end of the work, and as a "negative example (e.g., work other than painting work)" if the image data was captured during any other time period.

[0038] Here, we will explain an example of generating training data using linear-probing of the Vision-Language model, using "application work = work using a syringe containing black liquid" as an example of a positive example. Figure 5 is a diagram for explaining the generation of training data.

[0039] 5, the training data generation unit 31 acquires a text query "Grabbing a syringe with black liquid" for "application work," which is a positive example, from an input by an administrator, etc. Then, the training data generation unit 31 inputs the text query "Grabbing a syringe with black liquid" into a text feature extractor of the VL model to acquire text features.

[0040] Meanwhile, the training data generation unit 31 acquires frame images of one section (e.g., 20 minutes or 30 frames) extracted from video data showing all processes including the "application work." Next, the training data generation unit 31 generates patch images by extracting the periphery of a body part such as a hand using existing object detection or pre-specified area parameters that specify the area to be extracted for each frame image. The training data generation unit 31 then inputs each patch image to an image feature extractor of the VL model to acquire image features.

[0041] The training data generation unit 31 then calculates the similarity between the text feature and each image feature, identifies the patch image (target patch image) with the highest similarity, and determines whether the timestamp of the original image data of the target patch image is within the work time of the "application work."

[0042] Here, if the timestamp of the original image data of the target patch image is within the work time of the ``application work,'' the training data generation unit 31 determines that the target patch image is target work data and generates training data by adding the label ``application work'' to the target patch image.

[0043] On the other hand, if the timestamp of the original image data of the target patch image is outside the work time for "application work," the training data generation unit 31 determines that the target patch image is not target work data and determines whether the similarity is equal to or greater than a threshold. If the similarity is equal to or greater than the threshold, the training data generation unit 31 determines that the target patch image is high-similarity data and generates training data in which the label "other than application work" is added to the target patch image. On the other hand, if the similarity is less than the threshold, the training data generation unit 31 determines that the target patch image is low-similarity data and generates training data in which the label "other than application work" is added to the target patch image.

[0044] An example of training data generated by the above-described process will be described. FIG. 6 is a diagram illustrating an example of training data generation. As shown in FIG. 6, the training data generation unit 31 generates positive example training data by adding the label "application work" to image data of an "application work" in which a medicine is applied. The training data generation unit 31 also generates negative example training data (high similarity data) by adding the label "other than application work" to image data of a "screw work" in which a screw is driven in using a screwdriver, which has similar work content and tools. Similarly, the training data generation unit 31 generates negative example training data (low similarity data) by adding the label "other than application work" to image data of a state before the work or image data of a "seal work" in which a protective seal or the like is attached, which has dissimilar work content and tools.

[0045] 4 , the training execution unit 32 is a processing unit that executes machine learning for each classification model using each training data generated by the training data generation unit 31 and stored in the training data DB 15. Specifically, the training execution unit 32 generates two classification models that have different distributions for determining positive example data.

[0046] 7 is a diagram illustrating training of a classification model. As shown in FIG. 7, the training execution unit 32 treats, among the training data, low-similarity data labeled “other than application work” as negative example training data, and high-similarity data labeled “other than application work” and target work data labeled “application work” as positive example training data, thereby performing supervised learning of the first classification model 16.

[0047] In addition, the training execution unit 32 treats high-similarity data labeled "other than application work" from the training data as negative example training data, and target work data labeled "application work" as positive example training data, and performs supervised learning of the second classification model 17.

[0048] The classification processing unit 40 is a processing unit that detects the relevant task using the trained first classification model 16 and second classification model 17. For example, the classification processing unit 40 generates a patch image from image data of the prediction target by object detection or the like, and inputs the patch image to the first classification model 16. Then, when the first classification model 16 determines that the patch image is a positive example, the classification processing unit 40 inputs the patch image to the second classification model 17 and determines whether the task is a target task to be detected based on the output result of the second classification model 17. Then, the classification processing unit 40 displays the determination result on a display unit such as a monitor, or transmits it to a designated terminal.

[0049] FIG. 8 is a diagram illustrating classification processing using a classification model after training. As shown in FIG. 8, the classification processing unit 40 generates patch images from input image data using the same method as during learning. For example, the classification processing unit 40 inputs the text query "Grabbing a syringe with black liquid" for the "application work" to be detected into a text feature extractor of the VL model to acquire text features. The classification processing unit 40 also acquires frame images of a predetermined section and performs object detection or the like on each frame image to generate patch images. The classification processing unit 40 then inputs each patch image into an image feature extractor of the VL model to acquire image features. The classification processing unit 40 then calculates the similarity between the text features and each image feature and identifies the patch image to be predicted that has the highest similarity.

[0050] Then, the classification processing unit 40 inputs the identified patch image of the prediction target into the first classification model 16 to obtain the prediction result of the first classification model 16. Here, if the prediction result of the first classification model 16 is "non-applicable," that is, if it is a "negative example," the classification processing unit 40 determines that the patch image of the prediction target is "other than the target task."

[0051] On the other hand, when the prediction result of the first classification model 16 is "relevant," i.e., determined to be a "positive example," the classification processing unit 40 inputs the patch image to be predicted into the second classification model 17 to obtain the prediction result of the second classification model 17. Then, when the prediction result of the second classification model 17 is "non-relevant," i.e., a "negative example," the classification processing unit 40 determines the patch image to be predicted as "other than the target task." On the other hand, when the prediction result of the second classification model 17 is "relevant," i.e., a "positive example," the classification processing unit 40 determines the patch image to be predicted as "target task."

[0052] (Training Data Generation Process) Fig. 9 is a flowchart showing the flow of the training data generation process. As shown in Fig. 9, when an instruction to start processing is received (S101: Yes), the machine learning unit 30 acquires a text query for the target task (S102). Next, the machine learning unit 30 acquires frame images of a predetermined section (S103) and generates patch images from each frame image (S104).

[0053] The machine learning unit 30 then extracts features of the text query (text features) (S105) and extracts features of each patch image (image features) (S106). After that, the machine learning unit 30 calculates the similarity between the text features and each image feature (S107) and identifies the patch image with the highest similarity (target patch image) (S108).

[0054] Then, the machine learning unit 30 determines whether the target patch image corresponds to the target task based on annotations such as a timestamp (S109). If the target patch image corresponds to the target task (S109: Yes), the machine learning unit 30 registers the target patch image as training data for positive example data (S110).

[0055] On the other hand, if the target patch image does not correspond to the target task (S109: No), the machine learning unit 30 determines whether the similarity of the target patch image is equal to or greater than a threshold (S111). If the similarity of the target patch image is equal to or greater than the threshold (S111: Yes), the machine learning unit 30 registers the target patch image as training data for negative example data (high similarity data) (S112). On the other hand, if the similarity of the target patch image is less than the threshold (S111: No), the machine learning unit 30 registers the target patch image as training data for negative example data (low similarity data) (S113).

[0056] Thereafter, if the machine learning unit 30 continues generating training data (S114: No), it repeats S103 and subsequent steps, and if the machine learning unit 30 ends generating training data (S114: Yes), it ends the processing.

[0057] 10 is a flowchart showing the flow of machine learning processing of a classification model. As shown in FIG. 10, when an instruction to start processing is received (S201: Yes), the machine learning unit 30 acquires training data from the training data DB 15 (S202).

[0058] Next, the machine learning unit 30 treats the low-similarity data of the training data as negative examples, and the high-similarity data and the target task data as positive examples, and performs machine learning of the first classification model 16 (S203). Here, if the machine learning unit 30 decides to continue training (S204: No), it repeats S202 and subsequent steps.

[0059] On the other hand, if the training is to be ended (S204: Yes), the machine learning unit 30 acquires training data from the training data DB 15 (S205). Then, the machine learning unit 30 treats the highly similar data in the training data as negative examples and the target task data as positive examples, and performs machine learning of the second classification model 17 (S206). Here, if the training is to be continued (S207: No), the machine learning unit 30 repeats S205 and subsequent steps, and if the training is to be ended (S207: Yes), the machine learning unit 30 ends the processing.

[0060] The training of the first classification model 16 and the training of the second classification model 17 may be performed in either order, or may be performed in parallel.

[0061] As described above, the information processing device 10 according to the first embodiment can automatically collect images of a target task and images of similar tasks that are not the target task at the same time, and can introduce a detection algorithm incorporating a classification model for distinguishing between the target task and the similar tasks. As a result, the information processing device 10 can improve the accuracy of task detection.

[0062] The information processing device 10 of Example 1 can efficiently collect training data by utilizing the Vision-Language model, and therefore can easily construct a model that detects sections where specific work is being performed from video data of assembly work.

[0063] The information processing device 10 according to the first embodiment can incorporate the above detection algorithm, thereby enabling differentiation from similar tasks and realizing a model with fewer false positives. Furthermore, compared to the conventional technology, the information processing device 10 according to the first embodiment does not require advance model preparation for extracting similar tasks, does not require manual sorting of target tasks and similar tasks, and can identify similar tasks without requiring additional operator action for prediction.

[0064] Here, an evaluation of the classification model constructed by the method according to Example 1 will be described. Here, the target task is the application of sealant using a syringe, which is one of the tasks involved in assembling a product. Video data of assembly by worker A was used as training data, and video data of assembly by worker A that was different from that used for training was used as evaluation data.

[0065] As an evaluation process, training data was collected for training video data using the method according to Example 1, and the two classification models were trained using the collected training data. Then, the evaluation video data was applied to each classification model to perform predictions, and the prediction results were compared with the correct labels for evaluation.

[0066] FIG. 11 illustrates the evaluation results. The horizontal axis of FIG. 11 represents frames, and the vertical axis represents the percentage of detected positive examples. As shown in FIG. 11, the frames predicted as "positive examples" in the frame-by-frame predictions corresponding to the upper portion of the figure often correspond to frames where the correct task was performed, and the classification model also determines them as correct, demonstrating the high accuracy of the classification model. Furthermore, in the portion (A) of FIG. 11, some frames are determined to be "positive examples," but since they are not consecutively determined to be "positive examples," it can be determined that they are not the target task, and therefore the accuracy of the classification model cannot be said to be low. While there is concern about a deterioration in the accuracy of the classification model when video data of another worker B is used as the video data for evaluation, this concern is believed to be resolved by increasing the amount of training data. Furthermore, the method of Example 1 allows for accurate and easy collection of training data, eliminating the cost and problems of increasing the training data.

[0067] Although the embodiments of the present invention have been described above, the present invention may be embodied in various different forms other than the above-described embodiments.

[0068] (Numerical Values, etc.) The numerical values, graphs, examples of training data, positive examples, negative examples, etc. used in the above examples are merely examples and can be changed as desired. Furthermore, the process flow described in each flowchart can also be changed as appropriate within a consistent range. Furthermore, in the above examples, the work to be detected is described as "application work," but this is not limited to this, and any work, such as "work using a screwdriver" or "work marking with a pen," can be used as the target.

[0069] (System) The information including the processing procedures, control procedures, specific names, various data and parameters shown in the above documents and drawings may be changed arbitrarily unless otherwise specified.

[0070] Furthermore, the specific form of distribution or integration of the components of each device is not limited to that shown in the figure. For example, the machine learning unit 30 and the classification processing unit 40 may be integrated. That is, all or some of the components may be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc. Furthermore, all or any part of the processing functions of each device may be realized by a CPU and a program analyzed and executed by the CPU, or may be realized as hardware using wired logic.

[0071] Furthermore, all or any part of the processing functions performed by each device may be realized by a CPU and a program analyzed and executed by the CPU, or may be realized as hardware using wired logic.

[0072] (Hardware) Fig. 12 is a diagram illustrating an example of a hardware configuration. Here, an information processing device 10 will be described as an example. As shown in Fig. 12, the information processing device 10 includes a communication device 10a, a hard disk drive (HDD) 10b, a memory 10c, and a processor 10d. The components illustrated in Fig. 12 are connected to each other via a bus or the like.

[0073] The communication device 10a is a network interface card or the like, and communicates with other devices. The HDD 10b stores programs and DBs that operate the functions shown in FIG.

[0074] The processor 10d reads out a program that executes the same processes as the respective processing units shown in Fig. 4 from the HDD 10b or the like and loads it into the memory 10c, thereby operating a process that executes the respective functions described in Fig. 4 or the like. For example, this process executes the same functions as the respective processing units of the information processing device 10. Specifically, the processor 10d reads out a program that has the same functions as the machine learning unit 30, the classification processing unit 40, or the like from the HDD 10b or the like. Then, the processor 10d executes a process that executes the same processes as the machine learning unit 30, the classification processing unit 40, or the like.

[0075] In this way, the information processing device 10 operates as an information processing device that executes a simulation method by reading and executing a program. The information processing device 10 can also realize functions similar to those of the above-described embodiment by reading the program from a recording medium using a medium reading device and executing the read program. Note that the program in these other embodiments is not limited to being executed by the information processing device 10. For example, the above-described embodiment may also be applied in the same way when another computer or server executes the program, or when these computers cooperate to execute the program.

[0076] This program may be distributed via a network such as the Internet. Alternatively, this program may be recorded on a computer-readable recording medium such as a hard disk, a flexible disk (FD), a CD-ROM, a magneto-optical disk (MO), or a digital versatile disk (DVD), and may be read out from the recording medium and executed by a computer.

[0077] REFERENCE SIGNS LIST 10 Information processing device 11 Communication unit 12 Storage unit 13 Video data DB 14 Image data DB 15 Training data DB 16 First classification model 17 Second classification model 20 Control unit 30 Machine learning unit 31 Training data generation unit 32 Training execution unit 40 Classification processing unit

Claims

1. A model generation program that causes a computer to execute the following processes: acquire first image data capturing an image of a state in which a detection target task is being performed, second image data whose similarity to the first image data is equal to or greater than a threshold, and third image data whose similarity to the first image data is less than the threshold; perform machine learning of a first machine learning model using the first image data and the second image data as positive example data and the third image data as negative example data; and perform machine learning of a second machine learning model using the first image data as positive example data and the second image data as negative example data.

2. The model generation program according to claim 1, characterized in that the acquisition process executes the following process: acquire frame images of a predetermined section from video data of a period during which a worker performed work; generate the first image data from the frame images if the frame images belong to a time period during which the work to be detected was performed; and generate the second image data and the third image data from the frame images if the frame images do not belong to a time period during which the work to be detected was performed.

3. The model generation program according to claim 2, characterized in that the obtaining process executes the following processes: extracting text features that are features of a text query that represents the work to be detected; generating patch images by extracting parts of the worker's body from each frame image in the specified section, and extracting image features that are features of each patch image; calculating the similarity between the text features and each image feature; and generating the first image data, the second image data, and the third image data depending on whether a target patch image that is a patch image with a high similarity belongs to a time period in which the work to be detected was performed.

4. The model generation program according to claim 3, characterized in that the generating process executes the following processes: if the target patch image, which is a patch image with a high degree of similarity, belongs to the time period when the work to be detected was performed, generate the target patch image as the first image data; if the target patch image does not belong to the time period when the work to be detected was performed and the similarity is equal to or greater than a threshold, generate the target patch image as the second image data; and if the target patch image does not belong to the time period when the work to be detected was performed and the similarity is less than a threshold, generate the target patch image as the third image data.

5. The model generation program according to claim 1, characterized in that it executes the following process: inputting image data to be predicted into the first machine learning model that has been trained; if the prediction result of the first machine learning model is a negative example, determining that the image data to be predicted is not the work to be detected; if the prediction result of the first machine learning model is a positive example, inputting image data to be predicted into the second machine learning model that has been trained; if the prediction result of the second machine learning model is a negative example, determining that the image data to be predicted is not the work to be detected; and if the prediction result of the second machine learning model is a positive example, determining that the image data to be predicted is the work to be detected.

6. A model generation method characterized by the following processing performed by a computer: acquiring first image data capturing an image of a state in which a detection target task is being performed, second image data having a similarity to the first image data equal to or greater than a threshold, and third image data having a similarity to the first image data less than the threshold; performing machine learning of a first machine learning model using the first image data and the second image data as positive example data and the third image data as negative example data; and performing machine learning of a second machine learning model using the first image data as positive example data and the second image data as negative example data.

7. An information processing device having a control unit that acquires first image data capturing an image of a state in which a detection target task is being performed, second image data whose similarity to the first image data is equal to or greater than a threshold, and third image data whose similarity to the first image data is less than a threshold; performs machine learning of a first machine learning model using the first image data and the second image data as positive example data and the third image data as negative example data; and performs machine learning of a second machine learning model using the first image data as positive example data and the second image data as negative example data.

Citation Information

Patent Citations

  • Image recognition device and display device for vehicle

    JP2012164026A

  • Learning data collection device, learning device, learning data collection method, and program

    JP2022038941A

  • Learning model optimization device, learning model optimization method, and learning model optimization program

    JP2023059299A

  • Information processor and method for processing information

    JP2023060666A