System, inference model generation method, and inference model generation program

The system simplifies and enhances the generation of inference models by using image generation and synthesis units to train models on virtual bulk images, resulting in faster and more accurate inference capabilities for workpieces.

JP2026052844APending Publication Date: 2026-03-25YASKAWA DENKI KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing methods for generating inference models that infer work information from captured images of scattered works are complex and inefficient.

Method used

A system that includes an image generation unit to create multiple work images from different viewpoints, an image synthesis unit to generate virtual bulk images, and a learning unit to train an inference model using these images, allowing for easier and more accurate inference model generation.

Benefits of technology

The system enables more efficient and accurate generation of inference models that can infer work information from bulk-stacked workpieces, reducing training time and improving model accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026052844000001_ABST
    Figure 2026052844000001_ABST
Patent Text Reader

Abstract

To more easily generate inference models that perform inferences on a given task. [Solution] The system comprises an image generation unit that generates multiple work images showing work from different viewpoints, an image synthesis unit that generates one or more virtual random stack images showing multiple randomly stacked work from the multiple work images, and a learning unit that trains an inference model that infers work information about one or more work shown in the random stack image from the random stack image, based on the one or more virtual random stack images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One aspect of the present disclosure relates to a system, a method for generating an inference model, and an inference model generation program.

Background Art

[0002] Techniques for generating an inference model that infers work information regarding one or more works from a captured image of the scattered works are known. For example, in Patent Document 1, there are provided an imaging unit capable of imaging a first distance image of an object from a plurality of angles, a generation unit that generates a three-dimensional model of the object based on the first distance image and generates an extraction image showing a specific part of the object corresponding to a plurality of angles based on the three-dimensional model, a first model that can estimate the position of a specific part in an arbitrary second distance image based on the extraction images related to a plurality of angles and the first distance images related to the plurality of angles, a learning unit that generates a second model capable of detecting an object in a second distance image based on a plurality of first distance images, and an image recognition unit that applies the second model to a second distance image captured at a certain angle and, when the object is detected, applies the first model to estimate the position of the specific part of the object. An information processing apparatus is described.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] It is desired to more simply generate an inference model that executes inferences on works.

Means for Solving the Problems

[0005] A system relating to one aspect of this disclosure includes an image generation unit that generates multiple work images showing work from different viewpoints, an image synthesis unit that generates one or more virtual bulk images showing multiple work in a bulk pile based on the multiple work images, and a learning unit that trains an inference model that infers work information about one or more work shown in the bulk image based on the one or more virtual bulk images. [Effects of the Invention]

[0006] According to one aspect of this disclosure, it is possible to more easily generate inference models that perform inferences on a workpiece. [Brief explanation of the drawing]

[0007] [Figure 1] This figure shows an example of the functional configuration of an inference model generation system. [Figure 2] This figure shows an example of a computer hardware configuration used for an inference model generation system. [Figure 3] This flowchart shows an example of how an inference model generation system works during the learning phase. [Figure 4] This figure shows an example of adding correct labels to a 3D model of a workpiece. [Figure 5] This figure shows an example of a series of processes that generate a virtual stacked image. [Figure 6] This figure shows an example of an inference model structure. [Figure 7] This is an example of how the inference model generation system operates during the inference phase. [Modes for carrying out the invention]

[0008] The following describes various examples in this disclosure in detail with reference to the attached drawings. In the description of the drawings, identical or equivalent elements are denoted by the same reference numeral, and redundant descriptions are omitted.

[0009] [System Overview] The system relating to this disclosure is a computer system that generates an inference model, which is a computational model that infers work information about one or more workpieces from images of a bulk-stacked workpiece (i.e., images of an actual bulk-stacked workpiece). Therefore, the system can also be called an inference model generation system, or the system can be said to include an inference model generation system. In one example, the system generates one or more virtual bulk-stacked images used to generate an inference model, and trains the inference model based on the one or more virtual bulk-stacked images. The system may input images of an actual bulk-stacked workpiece into the inference model to infer work information about one or more real workpieces, and then perform an operation on at least one real workpiece based on that work information.

[0010] A "loosely stacked workpiece" refers to a collection of multiple workpieces that are not necessarily aligned. A "loosely stacked image" refers to an image showing multiple loosely stacked workpieces. A "virtual loosely stacked image" refers to an image of loosely stacked workpieces that does not actually exist, generated by image processing by a system. A "real-world loosely stacked image" refers to an image obtained by photographing a real-world loosely stacked image.

[0011] "Training an inference model" refers to generating an inference model using machine learning, a method that autonomously discovers laws or rules by iteratively learning based on given information. The generation of an inference model corresponds to the learning phase, and the use of the generated (trained) inference model corresponds to the inference phase (operation phase).

[0012] [System Configuration] Figure 1 shows the functional configuration of an example of an inference model generation system 1. In this example, system 1 comprises an acquisition unit 11, a 3D model generation unit 12, a labeling unit 13, a deformation unit 14, an image generation unit 15, an image synthesis unit 16, a storage unit 17, a learning unit 18, an inference unit 19, and a work execution unit 20 as functional components. System 1 is connected via a communication network to a machine 2 that is placed in a real workspace and performs work on a real workpiece. Machine 2 is, for example, a robot. Machine 2 may be a component of the system or it may be located outside the system.

[0013] The acquisition unit 11 is a functional element that acquires multiple images of the actual workpiece (i.e., the real workpiece). The 3D model generation unit 12 is a functional module that generates a three-dimensional model (3D model) of the workpiece from the multiple images of the workpiece. The labeling unit 13 is a functional module that adds a ground truth label, which indicates the correct ground truth of the workpiece information, to the 3D model of the workpiece. The deformation unit 14 is a functional module that deforms the 3D model of the workpiece. The image generation unit 15 is a functional module that generates multiple workpiece images, each showing the workpiece from a different viewpoint, based on the 3D model of the workpiece. Each workpiece image represents a virtual workpiece. The image synthesis unit 16 is a functional module that generates one or more virtual random stack images based on the multiple workpiece images. Each virtual random stack image is associated with the respective ground truth labels of the multiple workpieces. The storage unit 17 is a functional module that stores the generated virtual random stack images. The learning unit 18 is a functional module that trains an inference model 30 based on one or more virtual random stack images. The inference unit 19 is a functional module that inputs real random stack images into the trained inference model 30 and infers work information about one or more real workpieces. The work execution unit 20 is a functional module that performs an operation on at least one of the one or more real workpieces based on the inferred work information. For example, the work execution unit 20 causes machine 2 to perform that operation.

[0014] System 1 can be implemented by any type of computer. The computer may be a general-purpose computer such as a personal computer or a business server, or it may be incorporated into a dedicated device that executes specific processing.

[0015] FIG. 2 is a diagram showing an example of the hardware configuration of computer 100 used for system 1. In this example, computer 100 includes main body 110, monitor 120, and input device 130.

[0016] Main body 110 is a device having circuit 160. Circuit 160 has processor 161, memory 162, storage 163, input / output port 164, and communication port 165. The number of each hardware component may be 1 or 2 or more. Storage 163 records programs for configuring each functional module of main body 110. Storage 163 is a computer-readable recording medium such as a hard disk, a nonvolatile semiconductor memory, a magnetic disk, or an optical disk. Memory 162 temporarily stores programs loaded from storage 163, calculation results of processor 161, and the like. Processor 161 configures each functional module by executing a program in cooperation with memory 162. Input / output port 164 performs input / output of electrical signals with monitor 120 or input device 130 in accordance with a command from processor 161. Input / output port 164 may perform input / output of electrical signals with other devices. Communication port 165 performs data communication with other devices via communication network N in accordance with a command from processor 161.

[0017] Monitor 120 is a device for outputting information from main body 110. Examples of monitor 120 include display devices such as various displays and speakers.

[0018] Input device 130 is a device for inputting information to main body 110. Examples of input device 130 include operation interfaces such as a keypad, a mouse, and an operation controller.

[0019] The monitor 120 and the input device 130 may be integrated as a touch panel. For example, like a tablet computer, the main body 110, the monitor 120, and the input device 130 may be integrated.

[0020] Each functional module of the system 1 is realized by causing the inference model generation program to be loaded onto the processor 161 or the memory 162 and causing the processor 161 to execute the program. The inference model generation program includes code for realizing each functional module of the system 1. The processor 161 operates the input / output port 164 and the communication port 165 according to the inference model generation program, and reads and writes data in the memory 162 or the storage 163.

[0021] The inference model generation program may be provided after being recorded on a non-temporary recording medium such as a CD-ROM, a DVD-ROM, or a semiconductor memory. Alternatively, the inference model generation program may be provided via a communication network as a data signal superimposed on a carrier wave.

[0022] [Operation of the System] Referring to FIG. 3, as an example of the inference model generation method according to the present disclosure, the operation of the system 1 in the learning phase will be described. FIG. 3 is a flowchart showing an example of the operation as a processing flow S1. That is, the system 1 executes the processing flow S1.

[0023] In step S11, the acquisition unit 11 acquires a plurality of captured images obtained by capturing the actual work (i.e., the real work) from different viewpoints. All of the plurality of captured images depict the same work, but the directions in which the work is captured are different from each other among the plurality of captured images. The acquisition unit 11 may directly receive each captured image from an imaging device such as a camera, may read each captured image from a predetermined storage device such as an image database, or may receive each captured image input by the user.

[0024] In step S12, the 3D model generation unit 12 generates a 3D model of the workpiece from multiple captured images of the workpiece. In one example, for each of the multiple captured images of the workpiece, the 3D model generation unit 12 estimates the viewpoint of the imaging device (i.e., the position and orientation of the imaging device) at the time the image was taken. For example, the 3D model generation unit 12 performs this estimation process using a method called COLMAP. The 3D model generation unit 12 generates a 3D model of the workpiece by image synthesis using a neural radiance field (NeRF) based on the multiple captured images of the workpiece. NeRF represents the scene as a function that takes a two-dimensional vector representing the position and viewpoint direction in three-dimensional space as input and outputs the emitted color and its density at that position. NeRF estimates the light and transparency (θ,φ) of each position (x,y,z) in three-dimensional space using a neural network. When using NeRF, the 3D model generation unit 12 inputs each of the multiple captured images into a neural network based on the viewpoint estimated from the captured image, and estimates a 5-dimensional (x, y, z, θ, φ) space. The 3D model generation unit 12 then synthesizes the estimation results of the 5-dimensional space to generate a 3D model composed of a 3D mesh. The 3D model generation unit 12 may be implemented using, for example, derivative technologies such as Nerfacto, Instant-NGP, TensoRF, or Splatfacto.

[0025] In step S13, the labeling unit 13 adds the correct label for work information related to the workpiece to the 3D model of the workpiece. The correct label refers to information that indicates the correct work information.

[0026] Work information can be various types of information about a workpiece. For example, work information may be information about the workpiece as a whole, such as its segmentation, location, or orientation, or it may be a class indicating the classification of the workpiece, or it may be the structure of the workpiece. The structure of a workpiece refers to information that defines the location, shape, or orientation of at least a part of the workpiece. For example, the structure is represented by key points and the connections between key points. Work information about the workpiece itself is associated with the workpiece. Alternatively, work information may be information about a part of the workpiece, such as its segmentation, location, or orientation, or it may be a subclass indicating the classification of the part. Work information about a part is associated with that part. Alternatively, work information may be a position set for a workpiece, such as a work position, which is the position of an operation performed on the workpiece. A work position may be a picking position (e.g., gripping position), painting position, welding position, cutting position, machining position, etc. Alternatively, work information may be information about an operation performed on a workpiece.

[0027] The positions set as work information (for example, the position of the workpiece itself, the position of parts of the workpiece, the structure of the workpiece, and the working position) are relative positions determined in relation to the workpiece. These relative positions are not absolute positions, but are only determined when the workpiece is placed in three-dimensional space. The regions set as work information (for example, the area occupied by the workpiece, and parts of the workpiece) are relative regions determined in relation to the workpiece. These relative regions are not absolute regions, but are only determined when the workpiece is placed in three-dimensional space.

[0028] In one example, the labeling unit 13 provides the user with a user interface for adding correct labels to a 3D model. The labeling unit 13 may also provide a user interface for the user to select one or more information items from a plurality of pre-prepared types of information items for work information. The plurality of information items are pre-prepared for at least two of the various types of work information described above. The labeling unit 13 selects one or more information items in response to user operations. For example, the labeling unit 13 selects information items related to at least one of the relative position and relative region determined relative to the work as one or more information items. The labeling unit 13 adds the correct labels for the work information corresponding to the selected one or more information items to the 3D model of the work. In another example, the labeling unit 13 may add the correct labels for the work information corresponding to a plurality of pre-prepared types of information items to the 3D model of the work, and then select one or more information items from that plurality of information items in response to user operations. In other words, the labeling unit 13 may perform the process of adding the correct label to the 3D model and the process of selecting information items in either order. In any case, the labeling unit 13 may also function as a selection unit, selecting one or more information items from a plurality of pre-prepared types of information items regarding work information about the work, in accordance with the user's operation.

[0029] The labeling unit 13 may provide the user with a user interface for inputting correct labels and add the correct labels entered by the user to the 3D model. For example, the user inputs a correct label for at least one of the following: the workpiece itself, the skeleton, one or more parts, and the work position, and the labeling unit 13 adds that correct label to the 3D model. If the machine 2 is a robot, the labeling unit 13 may provide the user with a user interface for registering information about the robot's end effectors, such as grippers and suction hands (e.g., specifications of the end effectors). In this case, the labeling unit 13 may automatically set the correct labels for work positions, such as picking positions, on the 3D model based on the information about the end effectors, or it may semi-automatically set the correct labels on the 3D model based on further predetermined additional inputs from the user.

[0030] Figure 4 shows an example of adding correct labels to a 3D model of a workpiece. The workpiece shown in this example is a T-tube. In this example, the labeling unit 13 adds part definitions (locations of parts) 211 to 213 corresponding to three parts of the workpiece, and the workpiece skeleton 220 as workpiece information to the 3D model 200 of the workpiece in response to user operation.

[0031] Returning to Figure 3, in step S14, the deformation unit 14 deforms the 3D model of the workpiece. For example, step S14 may be performed if the workpiece is a flexible object. Step S14 is optional; for example, it is omitted if the workpiece is highly rigid. The deformation unit 14 approximates the 3D model with multiple particles connected by elastic parameters and deforms the 3D model. This deformation process may be implemented using Obi, an asset of the game development platform Unity® that can simulate the movement of deformable objects. The deformation unit 14 may accept a 3D model deformed by user operation, or it may automatically deform the 3D model. Of the one or more ground truth labels attached to the 3D model, the ground truth labels that are affected by the deformation of the 3D model (e.g., ground truth labels for segmentation, part information, work position, etc.) are modified to follow the deformation. The deformation unit 14 may generate two or more deformed 3D models based on the 3D model generated by the 3D model generation unit 12 (hereinafter also referred to as the "original 3D model"). The deformation unit 14 may provide the original 3D model and one or more deformed 3D models for subsequent processing.

[0032] In step S15, the image generation unit 15 generates multiple work images based on a 3D model of the work, each showing the work from a different viewpoint. Each work image represents a virtual work obtained by projecting the 3D model. In one example, the image generation unit 15 generates one or more work images showing the work from a viewpoint different from any of the multiple captured images acquired in step S11. The image generation unit 15 may generate multiple work images from both the original 3D model and one or more deformed 3D models. The image generation unit 15 associates the correct labels attached to the corresponding 3D models of the work with the multiple work images. The image generation unit 15 may generate individual work images using various methods. In one example, the image generation unit 15 generates multiple work images while changing the viewpoint of the virtual camera relative to the 3D model. Alternatively, the image generation unit 15 may render the 3D model to obtain a mask image, generate a pseudo-image showing the scene including the work using NeRF, and extract the work in the pseudo-image as a work image using the mask image.

[0033] In step S16, the image synthesis unit 16 generates one or more virtual scattered images representing multiple scattered workpieces based on multiple workpiece images, each associated with a correct label. Typically, the image synthesis unit 16 generates multiple virtual scattered images. In one example, the image synthesis unit 16 reads a source image from a predetermined storage device that shows the location (basket, box, stand, etc.) where the multiple workpieces are scattered. The image synthesis unit 16 then pastes the multiple workpiece images onto the source image to generate a virtual scattered image. As described above, each workpiece image is associated with the correct label of the workpiece information of the corresponding workpiece. Therefore, the virtual scattered image is associated with multiple correct labels corresponding to the multiple workpiece images that have been pasted onto it. The image synthesis unit 16 may paste at least one of the multiple workpiece images onto the source image multiple times. The image synthesis unit 16 may paste both a workpiece image generated based on the original 3D model and a workpiece image generated based on the deformed 3D model onto a single source image to generate a single virtual scattered image. The image synthesis unit 16 may also paste one or more images of the actual workpiece onto the source image in addition to multiple workpiece images to generate a virtual random stack image. The image synthesis unit 16 stores the generated one or more virtual random stack images in the storage unit 17.

[0034] In step S17, the learning unit 18 generates an inference model 30 by performing machine learning using one or more virtual random images. Typically, the learning unit 18 trains the inference model 30 based on multiple virtual random images.

[0035] The inference model 30 is constructed, for example, using a neural network. The inference model 30 may be configured using RTMDet, an object detection technique that offers high accuracy and short inference time. Alternatively, the inference model 30 may be configured using RTMPose, which estimates the object's pose, in addition to RTMDet. The learning unit 18 may configure the inference model 30 to present work information for workpieces whose recognition score, obtained by object detection, meets a predetermined criterion, and not present work information for workpieces whose recognition score does not meet the criterion. In other words, the learning unit 18 may configure the inference model 30 to present work information only for workpieces with a relatively high detection accuracy. When the recognition score is set to a range from 0 to 1, the criterion for the recognition score is set to, for example, 0.8.

[0036] In one example, the learning unit 18 pre-maintains multiple heads corresponding to multiple anticipated information items. The learning unit 18 then connects one or more heads corresponding to one or more information items selected in step S13 to the output layer of the network constituting the inference model 30. A head is a part that performs higher-order processing such as classification and judgment based on the features extracted from the input data by the computational model network. The head outputs the inference result corresponding to the information item. As described above, the labeling unit 13 may select information items related to at least one of the relative position and relative region determined relative to the workpiece as one or more information items. In this case, the learning unit 18 connects the head corresponding to at least one of the information items related to the relative position and relative region to the output layer. For example, if the relative position is the position of the operation performed on the workpiece, the learning unit 18 connects a head that recognizes the position of the operation to the output layer as the head corresponding to the information item related to the relative position. If the relative position is the skeleton of the workpiece, the learning unit 18 connects a head that recognizes that skeleton to the output layer as the head corresponding to the information item related to the relative position. If the relative region is associated with one or more subclasses set for the workpiece, the learning unit 18 connects a head that recognizes one or more subclasses to the output layer as a head corresponding to the information item related to the relative region.

[0037] The learning unit 18 trains an inference model 30 based on one or more virtual scattered images (for example, multiple virtual scattered images) stored in the memory unit 17. As described above, the learning unit 18 can train an inference model 30 that includes a network of one or more connected heads. In one example, the learning unit 18 performs the following processing for each virtual scattered image. That is, the learning unit 18 inputs the scattered image into a predetermined machine learning model. The learning unit 18 updates the parameter set in the machine learning model by performing backpropagation based on the error between the output data estimated by the machine learning model and the correct label associated with the virtual scattered image. The learning unit 18 repeats this processing for each virtual scattered image until a predetermined termination condition is met to generate an inference model 30. The termination condition may be processing all data records (scattered images) of the training data. As another example, the termination condition may be repeatedly processing all data records (scattered images) of the training data a predetermined number of times. Alternatively, the termination condition may be that a predetermined metric (e.g., the error between the correct label and the predicted value in the evaluation data, the accuracy rate, etc.) no longer improves, that is, that the metric converges. Note that the inference model 30 is a computational model that is estimated to be optimal, and is not necessarily the "actually optimal computational model". The learning unit 18 stores the trained inference model 30 in a predetermined memory device. The trained inference model 30 is then used by the inference unit 19.

[0038] Figure 5 shows an example of a series of processes for generating a virtual random stack image. In this example, the acquisition unit 11 acquires multiple captured images 300 obtained by photographing the workpiece 9 from different viewpoints (step S11). The 3D model generation unit 12 generates a 3D model 310 of the workpiece 9 from the multiple captured images 300 (step S12). The labeling unit 13 adds a correct label 311 of workpiece information related to the workpiece 9 to the 3D model 310 (step S13). In this example, the correct label 311 indicates the correct picking position. The deformation unit 14 can deform the 3D model 310 to generate a new 3D model of the workpiece 9 (step S14). The image generation unit 15 generates multiple workpiece images 320 based on the 3D model 310, each showing the workpiece 9 from a different viewpoint (step S15). Each workpiece image 320 is associated with a correct label 311. The image synthesis unit 16 generates one or more virtual random stack images 330 representing multiple randomly stacked workpieces 9 based on multiple workpiece images 320, each associated with a correct label 311 (step S16). The learning unit 18 trains the inference model 30 based on these random stack images 330 (step S17).

[0039] Let's explain the processing flow S1 shown in Figure 3 again. System 1 may execute processing flow S1 for each of the multiple classes of workpieces. A class refers to information indicating the classification of a part or the whole of a workpiece. Multiple classes of workpieces refer to multiple workpieces whose parts or the whole of the workpiece are classified differently from each other. For example, multiple classes of workpieces may be multiple workpieces of different types (for example, multiple workpieces of different shapes). Alternatively, multiple classes of workpieces may be multiple workpieces of the same type but with different classes assigned to them. For example, if a certain workpiece has a first part and a second part, System 1 processes the workpiece with a class assigned to the first part and the workpiece with a class assigned to the second part as multiple classes of workpieces. As another example, System 1 processes the front side and the back side of a certain workpiece as multiple classes of workpieces. Below, we will explain the process of generating one or more random stack images based on multiple captured images of multiple classes of workpieces.

[0040] In step S11, the acquisition unit 11 acquires multiple images of each of the multiple classes of workpieces by photographing the actual workpiece from different viewpoints.

[0041] In step S12, the 3D model generation unit 12 generates a 3D model of each of the multiple classes of workpieces from multiple captured images of the workpiece.

[0042] In step S13, the labeling unit 13 adds the correct label for work information related to each of the multiple classes of workpieces to the 3D model of the workpiece. For example, the labeling unit 13 adds the correct label to the 3D model of each of the multiple classes of workpieces in response to user operations on the workpiece.

[0043] In step S14, the deformation unit 14 can deform a 3D model of at least one of the multiple classes of workpieces.

[0044] In step S15, the image generation unit 15 generates multiple work images for each of the multiple classes of work, based on the 3D model of the work, showing the work from different viewpoints. The image generation unit 15 may generate multiple work images for at least one of the multiple classes of work, from both the original 3D model and one or more modified 3D models. The image generation unit 15 associates each work image with the correct label attached to the corresponding 3D model of the work. As described above, the image generation unit 15 may generate individual work images using various methods.

[0045] In step S16, the image synthesis unit 16 generates one or more virtual random stack images (for example, multiple virtual random stack images) that represent randomly stacked workpieces of multiple classes, based on multiple workpiece images, each of which has a correct label associated with it. A "virtual random stack image that represents multiple classes of workpieces" means a virtual random stack image that contains one or more workpieces for each of the multiple classes of workpieces.

[0046] In step S17, the learning unit 18 generates an inference model 30 by performing machine learning using one or more virtual random stack images (for example, multiple virtual random stack images). This inference model 30 has the function of inferring work information for each of two or more types (two or more classes) of work represented by a random stack image. That is, the learning unit 18 trains the inference model 30 based on one or more virtual random stack images so that it can infer work information for each of the multiple classes of work. In one example, the learning unit 18 connects one or more heads corresponding to the work and one or more information items selected according to the user's operation to the output layer of the network that constitutes the inference model 30 for each of the multiple classes of work. The learning unit 18 also outputs a class indicating the type of work from the output layer and configures the inference model 30 to switch one or more heads for inferring work information according to the output class. The learning unit 18 then trains the inference model 30, which includes a network in which one or more heads are connected for each of the multiple classes of work. If there may be work for multiple classes but it is not necessary to switch heads depending on the class, the learning unit 18 can connect one or more heads common to multiple classes to the output layer of the network that constitutes the inference model 30.

[0047] The inference model 30, including the heads, will be described with reference to Figure 6. Figure 6 shows an example of the structure of the inference model 30. In this example, the inference model 30 is configured such that three heads 32-34, corresponding to three types of information items selected in response to user operations, are connected to the output layer 31a of the network 31. Head 32 is connected to the output layer 31a in parallel with the output layer 31a, and heads 33 and 34 are connected downstream of the output layer 31a. For example, head 32 infers the segmentation of parts of the workpiece, head 33 infers the skeleton of the workpiece, and head 34 infers the picking position of the workpiece. Such an inference model 30 may be realized by implementing the network 31 and head 32 using RTMDet, and heads 33 and 34 using RTMPose.

[0048] In this example, three heads are provided for each of the three types of information items, corresponding to three classes of workpieces. If the three classes are distinguished as classes Cx, Cy, and Cz, then head 32a infers the segmentation of parts of class Cx workpieces, head 32b infers the segmentation of parts of class Cy workpieces, and head 32c infers the segmentation of parts of class Cz workpieces. Head 33a infers the skeleton of class Cx workpieces, head 33b infers the skeleton of class Cy workpieces, and head 33c infers the skeleton of class Cz workpieces. Head 34a infers the picking position of class Cx workpieces, head 34b infers the picking position of class Cy workpieces, and head 34c infers the picking position of class Cz workpieces.

[0049] The output layer 31a outputs a class indicating the type of workpiece. The inference model 30 infers workpiece information using heads 32a, 33a, and 34a, depending on whether the output layer 31a outputs class Cx. If the output layer 31a outputs class Cy, the inference model 30 infers workpiece information using heads 32b, 33b, and 34b. If the output layer 31a outputs class Cz, the inference model 30 infers workpiece information using heads 32c, 33c, and 34c. In this way, the inference model 30 selects one of the three heads 32, one of the three heads 33, and one of the three heads 34, depending on the class output from the output layer 31a. By switching heads in this way, the inference model 30 can infer workpiece information (part segmentation, skeleton, and picking position) for workpieces of multiple classes. When the inference model 30 infers work information for a single class of work, or when it is not necessary to switch heads depending on the class indicating the type of work, the number of heads 32, 33, and 34 is all 1, and no head switching occurs.

[0050] Referring to Figure 7, the operation of System 1 in the inference phase (operation phase) will be explained as an example of the inference model generation method related to this disclosure. Figure 7 is a flowchart showing an example of this operation as processing flow S2. That is, System 1 executes processing flow S2.

[0051] In step S21, the inference unit 19 acquires a real-world bulk image showing multiple real-world workpieces. In one example, the inference unit 19 receives a real-world bulk image from an imaging device that photographs workpieces bulk-stacked in a real-world workspace.

[0052] In step S22, the inference unit 19 inputs a real-world random stack image into the trained inference model 30 and infers work information for one or more real-world workpieces. Depending on the settings in the training phase, the inference unit 19 can infer various types of information about the workpiece as work information, such as segmentation, relative position, relative region, working position, class, and subclass. As described above, in one example, the inference unit 19 (inference model 30) presents work information for workpieces whose recognition score obtained by object detection meets a predetermined criterion, and does not present work information for workpieces whose recognition score does not meet the criterion.

[0053] In step S23, the work execution unit 20 performs a predetermined task on at least one of the one or more real workpieces shown in the real-world bulk image, based on the inferred workpiece information. For example, the work execution unit 20 causes a machine 2, such as a robot, to perform the task. Examples of tasks include picking, painting, welding, cutting, and processing. If machine 2 is a robot, the work execution unit 20 may generate a robot path for performing the task and cause the robot to perform the task based on that path.

[0054] [Differentiation] The technology described herein has been explained in detail above based on various examples. However, the technology described herein is not limited to the examples given above. Various modifications are possible without departing from the gist of this disclosure.

[0055] The system may use a pre-prepared 3D model of a workpiece instead of generating it from multiple images of the workpiece. Therefore, the system does not need to include functional modules corresponding to the acquisition unit 11 and the 3D model generation unit 12. As another example, the system does not need to deform the 3D model, and therefore does not need to include a functional module corresponding to the deformation unit 14.

[0056] The system may generate multiple workpiece images showing the workpiece from different viewpoints using predetermined image processing that does not involve a 3D model of the workpiece. Therefore, the system does not need to include functional modules corresponding to the acquisition unit 11, the 3D model generation unit 12, and the deformation unit 14.

[0057] The trained inference model can be ported to another computer system. Therefore, another computer system may perform the inference phase (operation phase). In other words, the system does not need to have functional modules corresponding to the inference unit 19 and the work execution unit 20.

[0058] In System 1 described above, the correct label is attached to the 3D model of the workpiece, and that correct label is associated with each workpiece image. Alternatively, the labeling unit may directly attach the correct label to each workpiece image. Or, another computer system may directly attach the correct label to each workpiece image, in which case the system does not need to have a functional module equivalent to the labeling unit 13.

[0059] The system may train an inference model based on one or more work images generated based on a 3D model of the work (for example, one or more work images showing the work from a viewpoint different from any of the multiple images of the work captured), without generating a virtual stack of images. This inference model is a computational model that infers work information about the work from an image containing the work. In this example, the system does not need to have a functional module corresponding to the image synthesis unit 16.

[0060] The image used to train the inference model may be different from either the work image or the virtual random stack image described above, as long as it includes the workpiece. In such a modified version, the system may comprise at least a labeling unit and a learning unit. The labeling unit selects one or more information items from a plurality of pre-prepared information items for work information related to the workpiece, in response to user operation. The labeling unit may add the correct labels for work information corresponding to the plurality of information items, or the correct labels for work information corresponding to the selected one or more information items, to the 3D model of the workpiece. The learning unit connects one or more heads corresponding to the selected one or more information items to the output layer of the network that constitutes the inference model for inferring work information about the workpiece from the image including the workpiece. The learning unit then trains the inference model, which includes the network with one or more heads connected, using at least the correct labels for work information corresponding to the selected information items that have been added to the 3D model of the workpiece. For example, the learning unit trains the inference model using one or more images that show the workpiece corresponding to the 3D model and are associated with its correct label. In this example, the image generation unit may generate one or more work images representing the workpiece (for example, multiple work images showing the workpiece from different viewpoints) based on its 3D model for training the inference model. That is, the learning unit may train the inference model using the correct labels of the workpiece information corresponding to the selected information items attached to the 3D model of the workpiece, and one or more work images representing the workpiece.

[0061] The system's hardware configuration is not limited to a configuration in which each functional module is realized by program execution. For example, at least a portion of the above-mentioned group of functional modules may be composed of logic circuits specialized for that function, or they may be composed of an ASIC (Application Specific Integrated Circuit) that integrates such logic circuits.

[0062] The processing steps for a method executed by at least one processor are not limited to the examples above. For example, some of the steps or processes described above may be omitted, or each step may be performed in a different order. Also, any two or more of the steps described above may be combined, or some of the steps may be modified or deleted. Alternatively, other steps may be performed in addition to each of the steps described above.

[0063] When comparing the relative magnitudes of two numbers in a computer system or within a computer, either the two criteria "greater than or equal to" and "greater than" may be used, or either the two criteria "less than or equal to" and "less than" may be used.

[0064] [Note] As can be seen from the various examples above, this disclosure includes the following aspects:

[0065] (Note 1) An image generation unit that generates multiple work images showing the work from different viewpoints, An image synthesis unit that generates one or more virtual bulk images representing the bulk stacked workpieces based on the aforementioned bulk workpiece images, A learning unit that trains an inference model, which infers work information relating to one or more workpieces shown in the bulk image, based on the one or more virtual bulk images, A system equipped with these features. According to Appendix 1, a random stack image is generated using multiple work images showing the work from various viewpoints, and an inference model is trained using this random stack image. This method reduces the time required to generate the random stack image, and therefore, the training of the inference model using the random stack image can be easily carried out. In other words, an inference model that performs inference on randomly stacked work can be generated more easily. Furthermore, since a random stack image is obtained that shows a situation in which multiple workpieces exhibiting various external shapes are randomly stacked, i.e., a random stack image that shows a situation similar to a real-world scenario, it is expected that a highly accurate inference model can be generated by training using this random stack image.

[0066] (Note 2) The image generation unit generates the multiple work images for each of the multiple classes of work, The image synthesis unit generates one or more virtual bulk images representing the bulk stacked workpieces of the multiple classes, The learning unit trains the inference model based on the one or more virtual random stack images so that it can infer the work information for each of the multiple classes of work. The system described in Appendix 1. According to Appendix 2, since a random stacking image showing workpieces of multiple classes (e.g., multiple classes) is generated, it becomes possible to generate an inference model that can infer workpiece information even when workpieces of different classes are randomly stacked.

[0067] (Note 3) The image generation unit generates the plurality of work images based on the 3D model of the work. The system described in Appendix 1 or 2. According to Appendix 3, by using a 3D model of the workpiece, it is possible to easily generate multiple workpiece images showing the workpiece from different viewpoints.

[0068] (Note 4) An acquisition unit that acquires multiple images obtained by photographing the actual workpiece from different viewpoints, A 3D model generation unit generates a 3D model of the workpiece from the plurality of captured images of the workpiece, The system described in Appendix 3 further includes the following: According to Appendix 4, a 3D model of the workpiece is generated based on multiple captured images of the actual workpiece, thus obtaining a 3D model that closely resembles the actual workpiece. By using this 3D model, multiple workpiece images that more closely resemble the actual workpiece can be obtained, and therefore, a more realistic image of the stacked workpiece can be generated. For example, when using a method that converts simulated images into realistic images using techniques such as CycleGAN (Sim2Real), it is necessary to train the model using many training images to obtain the realistic image, which takes a long time to train. In addition, the training can be unstable, so it is not guaranteed that a highly accurate image will be obtained in the end. In contrast, according to Appendix 4, a highly accurate stacked workpiece image can be generated more stably and in a shorter amount of time. Since a highly accurate stacked workpiece image is used for training, it becomes possible to generate an inference model that can accurately infer workpiece information in a shorter amount of time.

[0069] (Note 5) The 3D model generation unit generates the 3D model of the workpiece by image synthesis using a neural radiance field based on the plurality of captured images of the workpiece. The system described in Appendix 4. According to Appendix 5, since a neural radiance field is used to generate a 3D model from multiple captured images, a highly accurate workpiece image similar to the actual workpiece can be obtained. Using this workpiece image, it becomes possible to generate a highly accurate random stack image.

[0070] (Note 6) The deformation section further comprises a deformation section that approximates the 3D model of the workpiece with a plurality of particles connected by elastic parameters, thereby deforming the 3D model. The image generation unit generates the plurality of work images based on the deformed 3D model. The system described in any one of the appendices 3 to 5. According to Appendix 6, since multiple work images are generated based on the deformed 3D model, it is possible to easily generate 3D models with various external shapes and obtain multiple work images showing various external shapes. By using these work images, it is possible to generate a random stack image that more closely resembles a real-world scenario, and a highly accurate inference model can be generated by training using this random stack image. For example, it becomes possible to generate an inference model that accurately infers work information for flexible or irregularly shaped objects whose external shape changes easily.

[0071] (Note 7) The labeling unit further includes a labeling unit that adds the correct label for the work information of the workpiece to the 3D model of the workpiece, The image generation unit associates the correct label attached to the 3D model of the workpiece with each of the multiple workpiece images of the workpiece. The image synthesis unit generates one or more virtual random stack images based on the plurality of work images to which the correct labels are associated. The learning unit trains the inference model based on the one or more virtual random stack images to which multiple ground truth labels are associated. The system described in any one of the appendices 3 to 6. According to Appendix 7, the correct label is attached to the 3D model, and that correct label is associated with the work image. Since it is no longer necessary to attach the correct label to each individual work image, the labeling time can be reduced. As a result, the time required to generate the random stack images used for training the inference model can be reduced. In one example, since the correct label is attached to the 3D model in accordance with its deformation, the time required to label 3D models with various external shapes can also be reduced.

[0072] (Note 8) The label attachment portion is, From the aforementioned work information, one or more information items are selected from several types of pre-prepared information items according to the user's operation. The aforementioned learning unit, One or more heads corresponding to the selected one or more information items are connected to the output layer of the network constituting the inference model. The inference model, which includes the network in which the one or more heads are connected, is trained. The system described in Appendix 7. According to Appendix 8, in response to the user's operation, the system adds correct labels to the 3D model and sets the head of the inference model, corresponding to the information items selected. This mechanism makes it easy to generate an inference model that infers work information according to the user's requests.

[0073] (Note 9) The labeling unit selects, as one or more information items, information items relating to at least one of the relative position and relative region determined relative to the workpiece, The learning unit connects the head corresponding to the information item relating to at least one of the relative position and the relative region to the output layer. The system described in Appendix 8. According to Appendix 9, an inference model can be easily generated that infers a position or region determined relative to the workpiece.

[0074] (Note 10) The aforementioned relative position is the position of the operation performed on the workpiece. The head corresponding to the information item relating to the relative position is a head that recognizes the position of the work. The system described in Appendix 9. According to Appendix 10, it is possible to easily generate an inference model that infers the work position, which is often used to process bulk-stacked workpieces.

[0075] (Note 11) The aforementioned relative position is the skeleton of the workpiece, The head corresponding to the information item relating to the relative position is a head that recognizes the skeleton. The system described in Appendix 9. According to Appendix 11, it is possible to easily generate an inference model that infers the structure of a workpiece, which is often used to process bulk-packed workpieces.

[0076] (Note 12) The aforementioned relative region is associated with one or more subclasses set for the workpiece, The head corresponding to the information item relating to the relative region is a head that recognizes one or more subclasses. The system described in Appendix 9. According to Appendix 12, it is possible to easily generate inference models that infer subclasses, which are often used to process bulk workpieces.

[0077] (Note 13) The aforementioned learning unit, For each of the multiple classes of workpieces, one or more heads corresponding to the workpiece and the selected one or more information items are connected to the output layer. The inference model is configured such that a class indicating the type of work is output from the output layer, and one or more heads for inferring the work information are switched according to the output class. The system described in any one of the appendices 8-12. According to Appendix 13, the inference model is configured to infer work information for workpieces of multiple classes (e.g., multiple classes). Therefore, it becomes possible to infer work information for each workpiece even when workpieces of different classes are piled up randomly. For example, it is possible to infer work information for each workpiece even when workpieces of multiple classes with different basic structural configurations are mixed together.

[0078] (Note 14) The learning unit configures the inference model so as not to present work information for workpieces whose recognition score obtained by object detection does not meet a predetermined standard. The system described in any one of the appendices 1-13. According to Appendix 14, work information is not presented for workpieces whose recognition score does not meet the criteria. Therefore, work information can be presented only for workpieces with a relatively high degree of recognition accuracy. For example, work information can be presented only for workpieces that are expected to be processed reliably.

[0079] (Note 15) The learned inference model is input to an inference unit that infers work information about one or more of the real workpieces by inputting real-world bulk images representing multiple real workpieces. A work execution unit that performs work on at least one of the one or more actual workpieces based on the inferred work information, A system further comprising any one of the appendices 1 to 14. According to Appendix 15, the inference model can be used to process a real-world image of a bulk pile, and then an operation can be performed on at least one workpiece depicted in the image.

[0080] (Note 16) An acquisition unit that acquires multiple images obtained by photographing the actual workpiece from different viewpoints, A 3D model generation unit generates a 3D model of the workpiece by image synthesis using a neural radiance field based on the multiple captured images of the workpiece, An image generation unit generates one or more workpiece images showing the workpiece from a viewpoint different from any of the multiple captured images of the workpiece, based on the 3D model of the workpiece. A learning unit that trains an inference model, which infers work information about a work from an image containing the work, based on one or more work images, A system equipped with these features. According to Appendix 16, a 3D model of the workpiece is generated using a neural radiance field based on multiple captured images of the actual workpiece, and a workpiece image is generated using this 3D model. Through this series of processes, a highly accurate workpiece image similar to the actual workpiece can be obtained. For example, when using a method that converts a simulated image into a realistic image using techniques such as CycleGAN (Sim2Real), it is necessary to train the model using many training images to obtain a realistic image, which takes a long time to train. In addition, since the training can be unstable, a highly accurate image is not always obtained in the end. In contrast, according to Appendix 15, a highly accurate workpiece image can be generated more quickly and stably. By using a highly accurate workpiece image for training, it becomes possible to easily generate an inference model that can accurately infer workpiece information.

[0081] (Note 17) A selection unit that selects one or more information items from a pre-prepared set of multiple information items related to the work, according to the user's operation. A learning unit that connects one or more heads corresponding to one or more selected information items to the output layer of a network that constitutes an inference model for inferring work information about a work from an image including the work, and trains the inference model including the network to which the one or more heads are connected, using at least the correct labels of the work information that are attached to a 3D model of the work and correspond to the selected information items, A system equipped with these features. According to Appendix 17, the head of the inference model is set in accordance with the information item selected by the user. The inference model is then trained using at least the correct labels of the work information corresponding to that information item. This mechanism makes it easy to generate an inference model that infers work information according to the user's requests.

[0082] (Note 18) The system further comprises an image generation unit that generates one or more workpiece images representing the workpiece based on the 3D model of the workpiece, The learning unit further trains the inference model using the one or more work images. The system described in Appendix 16. According to Appendix 18, since a work image is generated from a 3D model of the work, it is easy to generate a work image to use for training an inference model.

[0083] (Note 19) A method for generating an inference model, which is performed by a system having at least one processor, The steps include generating multiple work images showing the work from different perspectives, The steps include generating one or more virtual bulk images representing the bulk stacked workpieces based on the aforementioned bulk workpiece images, The steps include training an inference model that infers work information relating to one or more workpieces shown in the aforementioned bulk image based on the one or more virtual bulk images, A method for generating inference models that includes this. According to Appendix 19, a random stack image is generated using multiple work images showing the work from various viewpoints, and an inference model is trained using this random stack image. This method reduces the time required to generate the random stack image, and therefore, the training of the inference model using the random stack image can be easily carried out. In other words, an inference model that performs inference on randomly stacked work can be generated more easily. Furthermore, since a random stack image is obtained that shows a situation in which multiple workpieces exhibiting various external shapes are randomly stacked, i.e., a random stack image that shows a situation similar to a real-world scenario, it is expected that a highly accurate inference model can be generated by training using this random stack image.

[0084] (Note 20) The steps include generating multiple work images that show the work from different perspectives, The steps include generating one or more virtual bulk images representing the bulk stacked workpieces based on the aforementioned bulk workpiece images, The steps include training an inference model that infers work information relating to one or more workpieces shown in the aforementioned bulk image based on the one or more virtual bulk images, A program that generates inference models, which is executed by a computer. According to Appendix 20, a random stack image is generated using multiple work images showing the work from various viewpoints, and an inference model is trained using this random stack image. This method reduces the time required to generate the random stack image, and therefore, the training of the inference model using the random stack image can be easily carried out. In other words, an inference model that performs inference on randomly stacked work can be generated more easily. Furthermore, since a random stack image is obtained that shows a situation in which multiple workpieces exhibiting various external shapes are randomly stacked, i.e., a random stack image that shows a situation similar to a real-world scenario, it is expected that a highly accurate inference model can be generated by training using this random stack image. [Explanation of symbols]

[0085] 1...Inference model generation system, 2...Machine, 11...Acquisition unit, 12...3D model generation unit, 13...Labeling unit, 14...Deformation unit, 15...Image generation unit, 16...Image synthesis unit, 17...Storage unit, 18...Learning unit, 19...Inference unit, 20...Task execution unit, 30...Inference model, 31...Network, 31a...Output layer, 32-34...Head, 300...Captured image, 310...3D model, 311...Correct label, 320...Work image, 330...Virtual random stack image.

Claims

1. An image generation unit that generates multiple work images showing the work from different viewpoints, An image synthesis unit that generates one or more virtual bulk stack images representing the bulk stacked workpieces based on the aforementioned bulk stack images, A learning unit that trains an inference model, which infers work information relating to one or more workpieces shown in the bulk image, based on the one or more virtual bulk images, A system equipped with these features.

2. The image generation unit generates the multiple work images for each of the multiple classes of work, The image synthesis unit generates one or more virtual bulk images representing the bulk stacked workpieces of the multiple classes, The learning unit trains the inference model based on the one or more virtual random stack images so that it can infer the work information for each of the multiple classes of work. The system according to claim 1.

3. The image generation unit generates the plurality of work images based on the 3D model of the work. The system according to claim 1.

4. An acquisition unit that acquires multiple images obtained by photographing the actual workpiece from different viewpoints, A 3D model generation unit generates a 3D model of the workpiece from the plurality of captured images of the workpiece, The system according to claim 3, further comprising:

5. The 3D model generation unit generates the 3D model of the workpiece by image synthesis using a neural radiance field based on the plurality of captured images of the workpiece. The system according to claim 4.

6. The deformation section further comprises a deformation section that approximates the 3D model of the workpiece with a plurality of particles connected by elastic parameters, thereby deforming the 3D model. The image generation unit generates the plurality of work images based on the deformed 3D model. The system according to claim 3.

7. The labeling unit further includes a labeling unit that adds the correct label for the work information of the workpiece to the 3D model of the workpiece. The image generation unit associates the correct label attached to the 3D model of the workpiece with each of the multiple workpiece images of the workpiece. The image synthesis unit generates one or more virtual random stack images based on the plurality of work images to which the correct labels are associated. The learning unit trains the inference model based on the one or more virtual random stack images to which multiple correct labels are associated. The system according to any one of claims 3 to 6.

8. The labeling unit selects one or more information items from a plurality of pre-prepared information items for the work information in accordance with the user's operation. The aforementioned learning unit, One or more heads corresponding to the selected one or more information items are connected to the output layer of the network constituting the inference model. The inference model, which includes the network in which the one or more heads are connected, is trained. The system according to claim 7.

9. The labeling unit selects, as one or more information items, information items relating to at least one of the relative position and relative region determined relative to the workpiece, The learning unit connects the head corresponding to the information item relating to at least one of the relative position and the relative region to the output layer. The system according to claim 8.

10. The aforementioned relative position is the position of the operation performed on the workpiece. The head corresponding to the information item relating to the relative position is a head that recognizes the position of the work. The system according to claim 9.

11. The aforementioned relative position is the skeleton of the workpiece, The head corresponding to the information item relating to the relative position is a head that recognizes the skeleton. The system according to claim 9.

12. The aforementioned relative region is associated with one or more subclasses set for the workpiece, The head corresponding to the information item relating to the relative region is a head that recognizes one or more subclasses. The system according to claim 9.

13. The aforementioned learning unit, For each of the multiple classes of workpieces, one or more heads corresponding to the workpiece and the selected one or more information items are connected to the output layer. The inference model is configured such that a class indicating the type of work is output from the output layer, and one or more heads for inferring the work information are switched according to the output class. The system according to claim 8.

14. The learning unit configures the inference model so as not to present work information for workpieces whose recognition score obtained by object detection does not meet a predetermined standard. The system according to any one of claims 1 to 6.

15. The learned inference model is input to an inference unit that infers work information relating to one or more of the real workpieces by inputting real-world bulk images representing multiple real workpieces. A work execution unit that performs work on at least one of the one or more actual workpieces based on the inferred work information, The system according to any one of claims 1 to 6, further comprising:

16. An acquisition unit that acquires multiple images obtained by photographing the actual workpiece from different viewpoints, A 3D model generation unit generates a 3D model of the workpiece by image synthesis using a neural radiance field based on the plurality of captured images of the workpiece, An image generation unit generates one or more workpiece images showing the workpiece from a viewpoint different from any of the multiple captured images of the workpiece, based on the 3D model of the workpiece. A learning unit that trains an inference model, which infers work information about a work from an image containing the work, based on one or more work images, A system equipped with these features.

17. A selection unit that selects one or more information items from a pre-prepared set of multiple information items related to the work, according to the user's operation. A learning unit that connects one or more heads corresponding to one or more selected information items to the output layer of a network that constitutes an inference model for inferring work information about a work from an image including the work, and trains the inference model including the network to which the one or more heads are connected, using at least the correct labels of the work information that are attached to a 3D model of the work and correspond to the selected information items, A system equipped with these features.

18. The system further comprises an image generation unit that generates one or more workpiece images representing the workpiece based on the 3D model of the workpiece, The learning unit further uses the one or more work images to train the inference model. The system according to claim 16.

19. A method for generating an inference model, which is performed by a system having at least one processor, The steps include generating multiple work images showing the work from different perspectives, The steps include generating one or more virtual bulk images representing the bulk stacked workpieces based on the aforementioned bulk workpiece images, The steps include training an inference model that infers work information relating to one or more workpieces shown in the aforementioned bulk image based on the one or more virtual bulk images, A method for generating inference models that includes this.

20. The steps include generating multiple work images showing the work from different perspectives, The steps include generating one or more virtual bulk images representing the bulk stacked workpieces based on the aforementioned bulk workpiece images, The steps include training an inference model that infers work information relating to one or more workpieces shown in the aforementioned bulk image based on the one or more virtual bulk images, A program that generates inference models, which is executed by a computer.

Citation Information

Patent Citations

  • Information processing device, image recognition method, and image recognition program

    JP6822929B2