Training image acquisition method, training image acquisition system, and training image acquisition program

The system automatically identifies and collects images where image recognition fails in robotic tasks, enabling efficient data collection for model improvement.

WO2025254053A1PCT designated stage Publication Date: 2025-12-11PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/019833
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-05
Filing Date
2025-06-02
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing technologies require manual analysis by operators to determine image recognition failures in robots, which is time-consuming and inefficient for collecting training data to improve recognition accuracy.

Method used

A system and method that automatically identifies failed images due to image recognition in a robot's tasks, extracts similar images, and uses them for additional learning of the image recognition model, without manual intervention.

Benefits of technology

Efficiently collects learning data for improving image recognition models by automatically identifying and utilizing images where recognition failures occur, enhancing the model's accuracy without manual analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025019833_11122025_PF_FP_ABST
    Figure JP2025019833_11122025_PF_FP_ABST
Patent Text Reader

Abstract

A training image acquisition method includes: executing image recognition on a captured captured image using a trained model for detecting an object on which prescribed work is to be executed; accumulating image log data in which the result of image recognition and the captured image are associated; searching for a failure image, in which the prescribed work has failed, from within the captured image; extracting a similar image similar to the failure image from within the image log data; and acquiring, as a training image that is to be used in additional training of the trained model, a failure image for which a factor in the failure of the prescribed work is assessed to be image recognition, on the basis of an image recognition result for the failure image and an image recognition result for the similar image.
Need to check novelty before this filing date? Find Prior Art

Description

Training image acquisition method, training image acquisition system, and training image acquisition program

[0001] The present disclosure relates to a training image acquisition method, a training image acquisition system, and a training image acquisition program.

[0002] Patent Literature 1 discloses an information processing device that causes a robot to grasp a target object. The information processing device acquires visual information by capturing an image of the target object, determines a distribution of the probability of success or failure related to grasping from the visual information, determines a grasping position for the target object based on the distribution of the probability of success or failure, moves the robot to the grasping position to control the grasping of the target object, determines whether the target object has been grasped successfully, and estimates and notifies the system state of the robot based on the determination result of whether the object has been grasped successfully at the grasping position and the probability of success or failure.

[0003] Japanese Patent Application Publication No. 2019-188516

[0004] Conventionally, there is a technology for analyzing a robot's log and estimating the cause of a failure to grasp a target object, as disclosed in Patent Document 1. The configuration of Patent Document 1 does not anticipate determining whether a target object has been grasped successfully due to a failure in image recognition. Therefore, there is a demand for improving the recognition accuracy of image recognition models to prevent image recognition failures. Here, a large amount of training data is required to improve the authentication accuracy of image recognition models.

[0005] However, analyzing whether image recognition has failed requires manual analysis by an operator, and collecting learning data requires a lot of time and manpower. Therefore, there is a demand for technology that can automatically determine whether image recognition has failed and efficiently collect captured images for which image recognition has been determined to have failed.

[0006] The present disclosure has been devised in consideration of the above-mentioned conventional circumstances, and aims to provide a learning image acquisition method, a learning image acquisition system, and a learning image acquisition program that make it more efficient to collect captured images where image recognition is a factor in picking failure.

[0007] The present disclosure provides a training image acquisition method performed by a system including a robot equipped with a camera that captures an image of an object and that performs a predetermined task on the object using a trained model to detect the object, and an image acquisition device communicatively connected to the robot, the training image acquisition method comprising: using the trained model, performing image recognition on an image captured by the camera; acquiring and accumulating image log data that associates the image recognition results from the image recognition with the captured image; searching for and acquiring failed images from the captured images in which the predetermined task has failed; extracting similar images from the image log data that are similar to the failed images; and, if it is determined, based on the image recognition results of the failed image and the image recognition results of the similar images, that the cause of failure of the predetermined task corresponding to the failed image is the image recognition, acquiring the failed image as a training image to be used for additional training of the trained model.

[0008] The present disclosure also provides a learning image acquisition system including a robot equipped with a camera that captures an image of an object and that performs a predetermined task on the object using a trained model that detects the object, and an image acquisition device communicably connected to the robot, wherein the robot uses the trained model to perform image recognition on the captured image captured by the camera and transmits the image to the image acquisition device, the image acquisition device acquires and stores image log data that associates the image recognition results from the image recognition with the captured image, searches for and acquires failed images from the captured images in which the predetermined task has failed, extracts similar images from the image log data that are similar to the failed images, and, if it determines, based on the image recognition results of the failed image and the image recognition results of the similar images, that the cause of failure of the predetermined task corresponding to the failed image is the image recognition, acquires the failed image as a learning image to be used for additional learning of the trained model.

[0009] The present disclosure also provides a learning image acquisition program executed by at least one processor that is equipped with a camera that captures an image of an object and is communicatively connected to a robot that performs a predetermined task on the object using a trained model to detect the object, the learning image acquisition program causing the processor to realize the following steps: acquiring an image captured by the camera and a result of image recognition performed on the image using the trained model; accumulating image log data that associates the result of the image recognition with the captured image; searching for and acquiring a failed image from the captured images that is an image in which the predetermined task has failed; extracting a similar image from the image log data that is similar to the failed image; and, if it is determined, based on the result of image recognition of the failed image and the result of image recognition of the similar image, that the cause of failure of the predetermined task corresponding to the failed image is the image recognition, acquiring the failed image as a learning image to be used for additional learning of the trained model.

[0010] According to the present disclosure, it is possible to more efficiently collect captured images in which the cause of picking failure is image recognition.

[0011] FIG. 1 is a diagram showing an example of a picking robot according to an embodiment; FIG. 2 is a block diagram showing an example of the internal configuration of a picking system according to an embodiment; FIG. 3 is a flowchart showing an example of the operation procedure of a database according to an embodiment; FIG. 4 is a diagram showing an example of extraction of similar scenes; FIG. 5 is a diagram explaining an example of determining the cause of a failure when the cause of a grasping failure is image recognition;

[0012] Hereinafter, with reference to the drawings as appropriate, detailed descriptions will be given of embodiments that specifically disclose the configurations and operations of the training image acquisition method, training image acquisition system, and training image acquisition program according to the present disclosure. However, more detailed descriptions than necessary may be omitted. For example, detailed descriptions of already well-known matters or redundant descriptions of substantially identical configurations may be omitted. This is to avoid unnecessary redundancy in the following description and to facilitate understanding by those skilled in the art. Note that the accompanying drawings and the following description are provided to enable those skilled in the art to fully understand the present disclosure and are not intended to limit the subject matter recited in the claims.

[0013] A picking robot RB according to an embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of a picking robot RB according to an embodiment. Note that the picking robot RB shown in Fig. 1 is an example and is not limited to this.

[0014] As an example, the picking robot RB in the present disclosure performs the task of picking up an object Tg contained in a box BX1 and transporting it to another box BX2.

[0015] The picking system 100 captures an image of an object Tg to be picked using at least one camera CM provided in the picking robot RB, and performs image recognition of the object appearing in the captured image. Using information such as the size, shape, and posture of the object Tg obtained from the image recognition results, the picking system 100 causes a hand HD provided in the picking robot RB to pick the object Tg and perform a picking operation of transporting the object Tg from box BX1 to box BX2.

[0016] The picking system 100 determines whether the object Tg was successfully grasped during the picking operation based on the control log of the picking robot RB, etc., and if it determines that the object Tg was not successfully grasped, extracts a similar scene that is an image similar to the captured image of the object Tg, and is an image of an object that was similarly unsuccessfully grasped in the past.

[0017] The picking system 100 determines whether the cause of the failure to grasp the object Tg is image recognition and estimates the likelihood that the cause of the failure to grasp the object Tg is image recognition, based on the image recognition result obtained using the captured image and the image recognition result corresponding to at least one similar scene. Furthermore, the picking system 100 determines whether to add the captured image of the object Tg to the learning data of the image recognition model, and performs additional learning of the image recognition model, based on the estimated likelihood that the cause of the failure to grasp the object Tg is image recognition.

[0018] Next, an example of the internal configuration of the picking system 100 according to the present disclosure will be described with reference to Fig. 2. Fig. 2 is a block diagram showing an example of the internal configuration of the picking system 100 according to the embodiment.

[0019] The picking system 100 includes a picking robot RB, a database DB, and a learning device P1. Note that the configuration of the picking system 100 shown in FIG. 2 is an example and is not limited to this. For example, the database DB may be a device or a server (cloud server) connected to one or more picking robots so that data can be communicated therewith. Furthermore, the learning device P1 and the database DB may be configured as an integrated unit.

[0020] The picking robot RB includes a communication unit (not shown), a processor 11, a memory 12, a camera CM, a hand HD, and a sensor SS. The communication unit (not shown) is configured using a communication interface circuit for transmitting and receiving data between the picking robot RB and the database DB and the learning device P1 via wired or wireless communication.

[0021] The processor 11 is configured using, for example, Graphics Processing Units (hereinafter referred to as "GPU"), Central Processing Units (hereinafter referred to as "CPU"), or Field Programmable Gate Arrays (hereinafter referred to as "FPGA"), and performs various processes and controls in cooperation with the memory 12. Specifically, the processor 11 references programs and data stored in the memory 12 and executes the programs to realize the functions of each unit of the processor 11, such as the image acquisition unit 111, the recognition unit 112, the control unit 113, and the control result acquisition unit 114.

[0022] The image acquisition unit 111 acquires the captured image (image data) and metadata of the captured image output from the camera CM, and outputs the acquired captured image to the recognition unit 112.

[0023] The recognition unit 112 acquires the captured image output from the image acquisition unit 111. The recognition unit 112 detects an object by performing image recognition processing on the acquired captured image using an object detection model transmitted from the learning device P1. Note that the object detection model referred to here is a learning model generated by learning for detecting an object Tg appearing in the captured image, such as a box BX1 that accommodates the object Tg or an edge (outer shape) of the object Tg.

[0024] The recognition unit 112 detects the box BX1 containing the object Tg or the object Tg through image recognition processing. Based on the detection result of the object Tg, the recognition unit 112 acquires detection information indicating the image recognition result, such as the image capture date and time, position, posture, shape, size, or number of the object Tg. The recognition unit 112 associates the acquired detection information with the captured image and outputs it to the control unit 113, or transmits it to the database DB and records it in the image log database 23 as an image log database corresponding to the object Tg.

[0025] The control unit 113 controls the hand HD and the sensor SS based on the detection information of the object Tg output from the recognition unit 112. As a result, the control unit 113 executes control to cause the picking robot RB to perform a picking task of grasping the object Tg and transporting it from the box BX1 to the box BX2. The control unit 113 acquires, as control data for the control unit 113, control details (e.g., control parameters) of the hand HD and the sensor SS, information about the object Tg (e.g., the ID of the object Tg, the size, shape, material, or color of the object Tg, production data for the object Tg, etc.), and information about the picking robot RB (e.g., the ID of the picking robot RB, version information of the software or hardware used in the picking robot RB, the number and arrangement of lights, etc.). The control unit 113 transmits the acquired control data to the database DB and records it in the control log database 24 as control log data corresponding to the object Tg.

[0026] The control result acquisition unit 114 acquires information indicating the control result from the hand HD and information indicating the detection result from the sensor SS. The control result acquisition unit 114 transmits the information indicating the control results of the hand HD and the sensor SS to the database DB and records it in the control log database 24 as control log data.

[0027] The memory 12 includes, for example, a random access memory (hereinafter referred to as "RAM") that serves as a work memory used when executing each process of the processor 11, and a read only memory (hereinafter referred to as "ROM") that stores programs and data that define the operation of the processor 11. The RAM temporarily stores data or information generated or acquired by the processor 11. The ROM stores programs that define the operation of the processor 11.

[0028] The camera CM is placed near the hand HD and moves integrally with the hand HD to capture an image of the object Tg to be picked. The camera CM associates the captured image of the object Tg with metadata corresponding to the captured image and transmits the resulting image to the image acquisition unit 111.

[0029] The hand HD is provided at the tip of the robot arm of the picking robot RB. The hand HD grasps the target object Tg in a manner that allows it to be picked, transports the target object Tg to the box BX2, and then releases the grasped state of the target object Tg to perform the picking operation. The hand HD is configured to be able to pick the target object Tg by a manner that allows it to be picked by a desired method, such as suction or clamping.

[0030] The sensor SS is installed by the operator in any number and at any position according to the detection target (detection target). At least one sensor SS is installed to detect the target Tg or the operation of the picking robot RB. The sensor SS outputs various detection results (detection results) to the processor 11.

[0031] The database DB includes a communication unit (not shown), a processor 21, a memory 22, an image log database 23, a control log database 24, and an image recognition database 25. The communication unit (not shown), which is not shown in Fig. 2, is configured using a communication interface circuit for transmitting and receiving data between the database DB and each of the picking robot RB and the learning device P1 via wired or wireless communication.

[0032] The processor 21 is configured using, for example, a GPU, a CPU, or an FPGA, and performs various processes and controls in cooperation with the memory 22. Specifically, the processor 21 references the programs and data stored in the memory 22 and executes the programs, thereby realizing the functions of each unit, such as the grasping failure data extraction unit 211, which is a function of the processor 21.

[0033] The grasping failure data extraction unit 211 searches for captured images of the target object Tg that has failed to be grasped (hereinafter referred to as "failure images") based on the control log data stored in the control log database 24. The grasping failure data extraction unit 211 starts to determine whether the cause of failure of each of the retrieved failure images is the image recognition process.

[0034] Based on the object Tg appearing in the failure image, the grasping failure data extraction unit 211 extracts at least one captured image (hereinafter referred to as a "similar scene") that depicts a scene similar to the object Tg appearing in the failure image from among the captured images stored in the image log database 23. Based on the image log data and control log data of the extracted similar scene and the image log data and control log data of the failure image, the grasping failure data extraction unit 211 determines whether the cause of the grasping failure is the image recognition process.

[0035] If the grasping failure data extraction unit 211 determines that the cause of the grasping failure is the image recognition process, it further determines whether or not to use the failure image for additional learning of the object detection model. If the grasping failure data extraction unit 211 determines that the failure image is to be used for additional learning of the object detection model, it stores (registers) the image log data and control log data corresponding to the failure image in the image recognition database 25 or transmits them to the learning unit 311 of the learning device P1.

[0036] The memory 22 includes, for example, a RAM as a work memory used when the processor 21 executes each process, and a ROM for storing programs and data that define the operation of the processor 21. The RAM temporarily stores data or information generated or acquired by the processor 21. The ROM stores programs that define the operation of the processor 21.

[0037] The image log database 23 stores image log data including captured images captured by the camera CM transmitted from the recognition unit 112 of the picking robot RB in association with each target object. The captured images stored in the image log database 23 include captured images in which the target object Tg was successfully grasped and captured images in which the target object Tg was unsuccessfully grasped (failed images). Whether the captured image stored in the image log database 23 is an image of whether the target object Tg was successfully grasped or unsuccessfully grasped is determined based on the control log data corresponding to the captured image.

[0038] The control log database 24 stores the control log data transmitted from the control unit 113 and the control result acquisition unit 114 of the picking robot RB for each target object.

[0039] The image recognition database 25 stores failed images of objects that have not been grasped in association with image log data corresponding to the objects. The various data stored in the image recognition database 25 is used for training (additional training) of the object detection model by the learning device P1.

[0040] The learning device P1 includes a communication unit (not shown), a processor 31, and a memory 32. The communication unit (not shown), which is not shown in Fig. 2, is configured using a communication interface circuit for transmitting and receiving data between the learning device P1 and each of the picking robot RB and the database DB via wired or wireless communication.

[0041] The processor 31 is configured using, for example, a GPU, a CPU, or an FPGA, and performs various processes and controls in cooperation with the memory 32. Specifically, the processor 31 references the programs and data stored in the memory 32 and executes the programs to realize the functions of each unit, such as the learning unit 311 and the object detection model generation unit 312.

[0042] The learning unit 311 additionally learns an object detection model for detecting objects (e.g., target object Tg, or boxes BX1, BX2, etc.) from captured images, using failed images in which it has been determined that image recognition is a factor in the failure of grasping and image log data of the failed images as learning data for learning erroneous object detection results. The learning unit 311 outputs the additionally learned object detection model to the object detection model generation unit 312.

[0043] The object detection model generation unit 312 tests the object detection accuracy of the trained object detection model using test data for testing the accuracy of the object detection model. If the object detection model generation unit 312 determines as a result of the test that the trained object detection model has a predetermined level of accuracy or higher, it generates this object detection model as a new object detection model and transmits it to the recognition unit 112 of the picking robot RB. The object detection model generation unit 312 updates the object detection model currently used by the picking robot RB with the newly generated object detection model.

[0044] The memory 32 includes, for example, a RAM as a work memory used when the processor 31 executes each process, and a ROM for storing programs and data that define the operation of the processor 31. The RAM temporarily stores data or information generated or acquired by the processor 31. The ROM stores programs that define the operation of the processor 31.

[0045] Next, an example of an operation procedure of the database DB will be described with reference to Fig. 3. Fig. 3 is a flowchart showing an example of the operation procedure of the database DB in the embodiment. Note that the operation procedure of the database DB shown in Fig. 3 is executed when a certain amount (e.g., 10 or 30 objects' worth of image log data and control log data of the object Tg transmitted from the picking robot RB and stored (accumulated) in the image log database 23 and the control log database 24, respectively, has been accumulated, or at a predetermined cycle (e.g., once a day or once a week).

[0046] The gripping failure data extraction unit 211 acquires each of the multiple pieces of image log data and control log data transmitted from the picking robot RB and stores them in the image log database 23 and the control log database 24, respectively (St11). The gripping failure data extraction unit 211 determines whether the picking robot RB failed to grip the object Tg shown in each captured image, based on each of the multiple pieces of control log data transmitted from the picking robot RB. The gripping failure data extraction unit 211 searches for and acquires failure images in which it is determined that the picking robot RB failed to grip the object Tg (St11).

[0047] The grasping failure data extraction unit 211 searches for similar scenes that are similar to the object Tg appearing in the failure image from among the captured images stored in the image recognition database 25 (St12), and extracts them (St13).

[0048] The grasping failure data extraction unit 211 compares the image log data and control log data corresponding to the extracted similar scene with the image log data and control log data corresponding to the failure image, and calculates the similarity between the similar scene and the failure image (St 14). For example, the grasping failure data extraction unit 211 extracts similar scenes by comparing the ID, number, shape, size, posture, or detection information by the sensor SS of each object appearing in the similar scene and the failure image, the ID of the picking robot RB that grasped each object appearing in the similar scene and the failure image, or software / hardware version information, or the image capture date and time of each of the similar scene and the failure image, or the lot of the object, etc.

[0049] The grasping failure data extraction unit 211 analyzes whether the cause of the failure to grasp the object Tg is a defect in the image recognition, or in the picking robot RB or the hand HD, based on the similarity between the similar scene and the scene of the failure image (St14). An example of determining (analyzing) the cause of the failure to grasp the object Tg will be described below.

[0050] For example, the gripping failure data extraction unit 211 determines whether there is a change in the gripping work environment of the picking robot RB, whether gripping has failed consecutively in similar scenes, etc. The gripping work environment here refers to the environment in which the picking robot RB performs the gripping work of an object, and is information included in the control log database 24, such as information about the change in object, the lot of the object, or the arrangement, number, or intensity of lighting. If the gripping failure data extraction unit 211 determines, based on the control log data for the time period before and after the failure image was captured, that there is a change in the gripping work environment of the picking robot RB and that the appearance of the object in the captured image has changed, or if it determines that gripping has failed consecutively in similar scenes, then it determines that image recognition may be a factor in the failure.

[0051] On the other hand, if it is determined that there have been no consecutive gripping failures in similar scenes, the gripping failure data extraction unit 211 determines that there is a malfunction of the picking robot RB or the hand HD. Furthermore, if the frequency of gripping failures decreases after calibration, inspection, or the like is performed, the gripping failure data extraction unit 211 determines that the cause of the gripping failure before calibration, inspection, or the like was a malfunction of the picking robot RB or the hand HD, and that the malfunction has been improved by the calibration or inspection.

[0052] The grasping failure data extraction unit 211 also determines whether the detection results (image recognition results) corresponding to similar scenes in which the same object is captured are stable, i.e., whether the similarity between the detection information of the similar scenes in which the same object is captured is high, and quantifies the likelihood that the cause of the grasping failure is image recognition based on an index obtained by combining several pieces of detection information corresponding to similar scenes. For example, the grasping failure data extraction unit 211 calculates the stability of the size of the detection frame of the object Tg captured in the similar scene, the likelihood (reliability) that the detected object Tg is the object, etc., and if it determines that the calculated stability is low, it determines that the cause of the grasping failure is image recognition. The grasping failure data extraction unit 211 also calculates each index that quantifies the similarity between the size of the detection frame and the detected shape of the detection information and the size and shape data of the object Tg, and determines whether the cause of the grasping failure is likely to be image recognition based on the calculated index.

[0053] The grasping failure data extraction unit 211 determines, as a result of the factor analysis, whether the factor of the failure to grasp the object Tg is image recognition (St15).

[0054] If the grasping failure data extraction unit 211 determines that the cause of the failure to grasp the object Tg is image recognition (St15, YES), it stores (registers) and updates the image log data and control log data of the object Tg in the image recognition database 25 (St16).

[0055] The grasping failure data extraction unit 211 associates the captured images (failure images) of multiple objects stored in the updated image recognition database 25 with a control command requesting additional learning of the object detection model using these data, and transmits them to the learning device P1 (St17).

[0056] Based on the control command transmitted from the database DB, the learning device P1 performs additional learning on the object detection model currently being used by the picking robot RB using the failed images transmitted from the database DB. The learning device P1 uses test data to verify the object detection accuracy of the object detection model obtained by the additional learning. If the learning device P1 determines that the obtained detection accuracy is equal to or greater than a predetermined value, it transmits the object detection model after the additional learning to the picking robot RB, causing the object detection model to be updated. On the other hand, if the learning device P1 determines that the obtained detection accuracy is not equal to or greater than the predetermined value, it discards the object detection model after the additional learning and omits updating the object detection model being used by the picking robot RB.

[0057] On the other hand, if the grasping failure data extraction unit 211 determines, as a result of the factor analysis, that the factor of the failure to grasp the object Tg is not image recognition (St15, NO), it ends the operation shown in FIG.

[0058] As described above, the database DB in the embodiment can automatically determine (estimate) whether the cause of the failure in grasping the target object Tg performed by the picking robot RB is image recognition. As a result, the picking system 100 can more efficiently collect (acquire) learning data by storing failed images determined to have been caused by image recognition in the image recognition database 25 as learning data for additional learning, without requiring an operator to perform image analysis.

[0059] Furthermore, by verifying the detection accuracy of the object detection model using test data, the picking system 100 can omit updating the object detection model used by the picking robot RB if it is determined that the object detection model obtained by additional learning does not satisfy the detection accuracy desired by the worker. As a result, the picking system 100 can continue to update only the object detection model that satisfies the detection accuracy desired by the worker, thereby improving the efficiency of the work performed on the object by the picking robot RB (here, grasping the object and transporting it from box BX1 to box BX2).

[0060] Next, an example of extracting similar scenes will be described with reference to Fig. 4. Fig. 4 is a diagram showing an example of extracting similar scenes. For ease of understanding, Fig. 4 describes an example of extracting similar scenes IMG11 to IMG15 (similar scenes) that are similar to a single failed image IMG10 when it is determined that grasping has failed.

[0061] The grasping failure data extraction unit 211 searches for similar scenes in which an object similar to the object Tg appearing in the failed image IMG10 is captured among the failure images stored in the image recognition database 25. The grasping failure data extraction unit 211 extracts similar scenes IMG11 to IMG15 in which a scene similar to the failed image IMG10 is captured as the similar scene corresponding to the failed image IMG10.

[0062] Next, examples of determining the cause of failure will be described with reference to Fig. 5 and Fig. 6. Fig. 5 is a diagram illustrating an example of determining the cause of failure when the cause of gripping failure is image recognition. Fig. 6 is a diagram illustrating an example of determining the cause of failure when the cause of gripping failure is something other than image recognition.

[0063] 5 and 6, for ease of understanding, an example of determining the cause of failure will be described using two similar scenes IMG11 and IMG13 out of similar scenes IMG11 to IMG15 that are similar to the failed image IMG10.

[0064] Similar scenes IMG11A and IMG13A shown in Fig. 5 are captured when the similar scenes IMG11 and IMG13 extracted in Fig. 4 are captured during a grasping failure, and the cause of the grasping failure is image recognition. Similar scene IMG11A is associated with a detection frame FRM11A in which an object Tg11 is detected as one piece of detection information. Similar scene IMG13A is associated with a detection frame FRM13A in which an object Tg13 is detected as one piece of detection information.

[0065] In such a case, the grasping failure data extraction unit 211 determines the similarity in position, size, or shape between the detection frame in which the object Tg1 was detected and the detection frames FRM11A and FRM13A in which the objects Tg11 and Tg13, which are the same objects as the object Tg1, were detected, relative to the size of the object Tg1 based on the control log data. If the grasping failure data extraction unit 211 determines that the detection frame in which the object Tg1 was detected and the detection frames when grasping of the objects Tg11 and Tg13 failed are similar (that is, the image recognition results are similar), the grasping failure data extraction unit 211 determines that the cause of the failure to grasp the object Tg1 was image recognition, which is the same as the cause of the failure of the similar scenes IMG11A and IMG13A.

[0066] Similar scenes IMG11B and IMG13B shown in Fig. 6 are captured when the similar scenes IMG11 and IMG13 extracted in Fig. 4 are captured when the grasping is successful and the image recognition of the object is successful. Similar scene IMG11B is associated with a detection frame FRM11B in which the object Tg11 is detected as one piece of detection information. Similar scene IMG13B is associated with a detection frame FRM13B in which the object Tg13 is detected as one piece of detection information.

[0067] In such a case, the grasping failure data extraction unit 211 determines the similarity in position, size, or shape between the detection frame in which the object Tg1 was detected and the detection frames FRM11B and FRM13B in which the objects Tg11 and Tg13, which are the same objects as the object Tg1, were detected, relative to the size of the object Tg1 based on the control log data. If the grasping failure data extraction unit 211 determines that the detection frame in which the object Tg1 was detected is similar to the detection frames used when the objects Tg11 and Tg13 were successfully grasped (that is, the image recognition results are similar), it determines that the cause of the failure to grasp the object Tg1 was something other than image recognition.

[0068] As described above, the database DB in the embodiment can determine whether the cause of the gripping failure corresponding to the failed image is image recognition based on information corresponding to similar scenes (e.g., control log data, success / failure information based on the control log data, image log data, etc.), thereby enabling the database DB to more efficiently acquire failed images where image recognition is the cause of the failure.

[0069] (Additional Notes) The above description of each embodiment discloses the following techniques.

[0070] (Technology 1) A training image acquisition method performed by a system (picking system 100) including: a robot (picking robot RB) equipped with a camera CM that captures an image of an object and performs a predetermined task on the object using a trained model (object detection model) that detects the object; and an image acquisition device (database DB) connected to be able to communicate with the robot (picking robot RB), the training image acquisition method comprising: performing image recognition on an image captured by the camera CM using the trained model (object detection model); acquiring and accumulating image log data that associates the image recognition results from the image recognition with the captured image; searching for and acquiring failed images from the captured images that have failed the predetermined task; extracting similar images (similar scenes) that are similar to the failed images from the image log data; and, if it is determined that the cause of failure of the predetermined task corresponding to the failed image is the image recognition based on the image recognition results of the failed image and the image recognition results of the similar images (similar scenes), acquiring the failed image as a training image to be used for additional learning of the trained model (object detection model). This allows the picking system 100 to collect (acquire) learning data more efficiently by storing failed images determined to be due to image recognition in the image recognition database 25 as learning data for additional learning, without the need for image analysis work by an operator.

[0071] (Technology 2) The learning image acquisition method described in (Technology 1) determines a similarity between the image recognition result of the failed image and the image recognition result of the similar image (similar scene), and if it is determined based on the similarity that the cause of failure of the specified task corresponding to the failed image is the image recognition, acquires the failed image as the learning image. This allows the picking system 100 to determine whether the image recognition result of the failed image is similar to (i.e., similar to) the similar scene based on the similarity of the image recognition result with the similar scene. Therefore, the picking system 100 can more efficiently determine whether the cause of failure is image recognition based on the similarity of the image recognition results and collect (acquire) learning data.

[0072] (Technology 3) The learning image acquisition method according to (Technology 2), wherein if it is determined that the similar image has failed the predetermined task and the similarity is equal to or greater than a predetermined value, the failure image is acquired as the learning image. This allows the picking system 100 to determine whether the failure factor of the failure image is the same as (i.e., similar to) the failure factor of the similar scene in the predetermined task, based on the similarity of the image recognition result with the similar scene.

[0073] (Technology 4) The learning image acquisition method according to any one of (Technology 1) to (Technology 3), comprising: acquiring control log data of the robot (picking robot RB) that performed the predetermined task on the object shown in each of the captured images, correlating the control log data with the image log data and storing the data; and searching for and acquiring the failure image from among the captured images based on the control log data. This enables the picking system 100 to determine which of the stored captured images is a failure image resulting from failure in the predetermined task.

[0074] (Technology 5) The learning image acquisition method according to (Technology 4), wherein if it is determined that the cause of failure of the predetermined task corresponding to the failed image is other than the image recognition based on the control log data of the failed image and the control log data of the similar image (similar scene), acquisition of the learning image is omitted. This allows the picking system 100 to more efficiently exclude from candidates for learning data failure images that are determined to have a cause of failure other than image recognition (for example, a malfunction of the hand HD) based on the control log data.

[0075] (Technology 6) The learning image acquisition method according to (Technology 5), wherein the control log data includes environmental information about the environment in which the predetermined work is performed. This allows the picking system 100 to more accurately determine whether the cause of the failure of the failed image is due to a factor other than image recognition (for example, a malfunction of the hand HD), based on the environmental information included in the control log data.

[0076] (Technology 7) A learning image acquisition system (picking system 100) includes: a robot (picking robot RB) equipped with a camera CM that captures an image of an object and that performs a predetermined task on the object using a trained model (object detection model) to detect the object; and an image acquisition device (database DB) communicably connected to the robot (picking robot RB), wherein the robot (picking robot RB) performs image recognition on the captured image captured by the camera CM using the trained model (object detection model) and transmits the image to the image acquisition device; the image acquisition device acquires and stores image log data that associates the image recognition results with the captured image; searches for and acquires failed images from the captured images when the predetermined task fails; and extracts similar images (similar scenes) that are similar to the failed image from the image log data. When it is determined that the cause of failure of the predetermined task corresponding to the failed image is the image recognition based on the image recognition result of the failed image and the image recognition result of the similar image (similar scene), the failed image is acquired as a learning image to be used for additional learning of the trained model (object detection model). In this way, the picking system 100 can collect (acquire) learning data more efficiently by storing the failed image, whose cause of failure is determined to be image recognition, in the image recognition database 25 as learning data for additional learning, without image analysis work by an operator.

[0077] (Technology 8) A learning image acquisition program executed by at least one processor 21 that is equipped with a camera CM that captures an image of an object and is communicably connected to a robot (picking robot RB) that performs a predetermined task on the object using a trained model (object detection model) that detects the object, the learning image acquisition program causing the processor 21 to perform the following steps: acquire an image captured by the camera CM and a result of image recognition performed on the image using the trained model (object detection model); accumulate image log data that associates the result of the image recognition with the captured image; search for and acquire a failed image from the captured images that indicates a failure in the predetermined task; extract a similar image (similar scene) that is similar to the failed image from the image log data; and, when it is determined that the cause of failure of the predetermined task corresponding to the failed image is the image recognition based on the result of the image recognition of the failed image and the result of the image recognition of the similar image (similar scene), acquire the failed image as a learning image to be used for additional learning of the trained model (object detection model). This allows the database DB to collect (acquire) learning data more efficiently by storing failed images determined to be due to image recognition in the image recognition database 25 as learning data for additional learning, without the need for an operator to perform image analysis.

[0078] Although various embodiments have been described above with reference to the accompanying drawings, the present disclosure is not limited to such examples. It is clear that those skilled in the art can conceive of various modifications, alterations, substitutions, additions, deletions, and equivalents within the scope of the claims, and it is understood that these also fall within the technical scope of the present disclosure. Furthermore, the components of the various embodiments described above may be combined in any manner without departing from the spirit of the invention.

[0079] This application is based on a Japanese patent application (Patent Application No. 2024-091465) filed on June 5, 2024, the contents of which are incorporated herein by reference.

[0080] The present disclosure is useful as a training image acquisition method, a training image acquisition system, and a training image acquisition program that make it possible to more efficiently collect captured images in which the cause of picking failure is image recognition.

[0081] 11, 21, 31 Processor 12, 22, 32 Memory 23 Image log database 24 Control log database 25 Image recognition database 100 Picking system 111 Image acquisition unit 112 Recognition unit 113 Control unit 114 Control result acquisition unit 211 Grasping failure data extraction unit 311 Learning unit 312 Object detection model generation unit CM Camera DB Database FRM11A, FRM11B, FRM13A, FRM13B Detection frame HD Hand IMG10 Failure image IMG11, IMG11A, IMG11B, IMG12, IMG13, IMG13A, IMG13B, IMG14, IMG15 Similar scene P1 Learning device RB Picking robot Tg, Tg1, Tg11, Tg13 Object

Claims

1. A training image acquisition method performed by a system comprising a robot equipped with a camera that captures an image of an object and performs a predetermined task on the object using a trained model that detects the object, and an image acquisition device connected to the robot so as to be able to communicate with the robot, the training image acquisition method comprising: performing image recognition on the image captured by the camera using the trained model; acquiring and accumulating image log data that associates the image recognition results from the image recognition with the captured image; searching for and acquiring failed images from the captured images in which the predetermined task has failed; extracting similar images from the image log data that are similar to the failed images; and, if it is determined, based on the image recognition results of the failed image and the image recognition results of the similar images, that the cause of failure of the predetermined task corresponding to the failed image is the image recognition, acquiring the failed image as a training image to be used for additional training of the trained model.

2. The learning image acquisition method according to claim 1, further comprising determining the degree of similarity between the image recognition result of the failed image and the image recognition result of the similar image, and acquiring the failed image as the learning image if it is determined based on the degree of similarity that the cause of failure of the specified task corresponding to the failed image is the image recognition.

3. The learning image acquisition method according to claim 2, wherein if it is determined that the similar image has failed the specified task and the similarity is equal to or greater than a specified value, the failed image is acquired as the learning image.

4. A learning image acquisition method as described in claim 1, further comprising: acquiring control log data of the robot that performed the specified task on the object shown in each of the captured images; storing the control log data in association with the image log data; and searching for and acquiring the failed images from among the captured images based on the control log data.

5. A training image acquisition method as described in claim 4, wherein if it is determined based on the control log data of the failed image and the control log data of the similar image that the cause of failure of the specified task corresponding to the failed image is other than the image recognition, acquisition of the training image is omitted.

6. The learning image acquisition method according to claim 5, wherein the control log data includes information about an environment in which the predetermined work is performed.

7. A learning image acquisition system comprising: a robot equipped with a camera that captures an image of an object and that performs a predetermined task on the object using a trained model to detect the object; and an image acquisition device communicably connected to the robot, wherein the robot: uses the trained model to perform image recognition on the image captured by the camera and transmits the image to the image acquisition device; the image acquisition device acquires and stores image log data that associates the image recognition results from the image recognition with the captured image; searches for and acquires failed images from the captured images where the predetermined task has failed; extracts similar images from the image log data that are similar to the failed images; and, if it is determined, based on the image recognition results of the failed image and the image recognition results of the similar images, that the cause of failure of the predetermined task corresponding to the failed image is the image recognition, acquires the failed image as a learning image to be used for additional learning of the trained model.

8. A training image acquisition program executed by at least one processor that is equipped with a camera that captures an image of an object and is communicatively connected to a robot that performs a predetermined task on the object using a trained model to detect the object, the training image acquisition program causing the processor to perform the following steps: acquire an image captured by the camera and a result of image recognition performed on the image using the trained model; accumulate image log data that associates the result of image recognition with the image; search for and acquire a failed image from the captured images that indicates a failure in the predetermined task; extract a similar image from the image log data that is similar to the failed image; and, if it is determined that the cause of failure of the predetermined task corresponding to the failed image is the image recognition based on the result of image recognition of the failed image and the result of image recognition of the similar image, acquire the failed image as a training image to be used for additional training of the trained model.

Citation Information

Patent Citations

  • Image search device and teacher data extraction method

    JP2020135494A

  • Information processing device, information processing method and information processing program

    WO2020026643A1