Robot control device, robot control system, trained model, and method for generating trained model
Patent Information
- Application Number
- JP2025516823
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2025-10-23
- Publication Date
- 2026-01-27
AI Technical Summary
Robot control systems face challenges in accurately selecting and holding work targets due to visibility issues, as existing technologies do not effectively account for object overlap and field of view limitations, leading to increased failure rates in object recognition and grasping.
A robot control system with a trained model that includes a feature extraction unit, mask generation unit, and visibility estimation unit, which acquires images, generates masks, and estimates visibility based on image characteristics, allowing the robot to select work targets with higher visibility scores, reducing failure rates by prioritizing objects that are fully visible and less likely to be hidden or cut off.
The system significantly reduces failure rates in object recognition and grasping by selecting work targets with higher visibility scores, improving the reliability of robot operations by considering the visibility of objects in the image.
Abstract
Description
Robot control device, robot control system, trained model, and trained model generation method Cross-reference to related applications
[0001] This application claims priority from People's Republic of China Patent Application No. 202310461786.6 (filed April 25, 2023), the entire disclosure of which is incorporated herein by reference.
[0002] The present disclosure relates to a robot control device, a robot control system, a trained model, and a method for generating a trained model.
[0003] 2. Description of the Related Art Conventionally, a bulk picking device equipped with an object recognition processing device has been known (see, for example, Patent Document 1).
[0004] JP 2010-120141 A
[0005] A robot control device according to an embodiment of the present disclosure includes an acquisition unit, an estimation unit, and a work object identification unit. The acquisition unit acquires images of multiple work objects for a robot. The estimation unit estimates a visibility of each of the multiple work objects based on the images. The work object identification unit selects a specific work object for the robot from the multiple work objects based on the visibility.
[0006] A robot control system according to an embodiment of the present disclosure includes the robot control device and a display device, wherein the display device displays a score indicating a basis for selecting the specific task object selected by the robot control device.
[0007] A trained model according to an embodiment of the present disclosure includes a feature extraction unit, a mask generation unit, and a visibility estimation unit. The feature extraction unit extracts features of each of a plurality of objects based on an image of the objects. The mask generation unit generates a plurality of masks representing each of the plurality of objects based on the features of each of the objects. The visibility estimation unit estimates the visibility of each of the plurality of objects based on the masks.
[0008] A method for generating a trained model according to one embodiment of the present disclosure includes performing training to generate a trained model that estimates the visibility of an object appearing in an inference image when the inference image is input, based on training data including a plurality of training images that show the object.
[0009] 7 is a block diagram showing an example of the configuration of a robot control system according to an embodiment. FIG. 4A is a schematic diagram showing an example of the configuration of a robot control system according to an embodiment. FIG. 4B is a block diagram showing an example of the configuration of a trained model that generates a mask and estimates the visibility of the mask. FIG. 4C is a diagram showing an example of an image to be input to the trained model. FIG. 4D is a diagram showing an example of a bounding box generated for each object in FIG. 4A. FIG. 4E is a diagram showing an example of a mask generated for each object in FIG. 4A. FIG. 4F is an enlarged view of the framed area A in FIG. 4A. FIG. 4G is a diagram explaining the size of a mask image used to generate a synthetic image or annotation data to be used as a training image. FIG. 4H is a diagram showing an example of an image depicting a work object. FIG. 4H is a diagram showing an image obtained by cutting out the framed area B in FIG. 7. FIG. 7I is a flowchart showing an example of the procedure of a trained model generation method. FIG. 7J is a flowchart showing an example of the procedure of a robot control method. FIG. 7J is a graph showing the area distribution of a mask. FIG. 7I is a diagram showing an example of a histogram of mask sizes.
[0010] The robot control system 1 according to an embodiment of the present disclosure can select a specific work target to be held by the robot based on the visibility of each object that is the work target.
[0011] The visibility of an object is the ratio of the area of the actual image of the object to the area of the object if the entire object were to appear in the image. For example, when another object overlaps an object and partially obscures it, the area of the object in the image may become smaller. In this case, the visibility of the object is the ratio of the area of the actual image that is not obscured by other objects to the area that would appear in the image if the object were not obscured by other objects. Also, when an object is located at the edge of the camera's field of view and is cut off from the image, the area of the object in the image may become smaller. In this case, the visibility of the object is the ratio of the area of the actual image that is not obscured to the area that would appear in the image if the object were not obscured by other objects.
[0012] The lower the visibility of an object in an image captured by a robot, the greater the area in which the object is hidden by other objects when viewed from the robot. In other words, the object is located further back than other objects when viewed from the robot. When a robot tries to grasp an object as a task target, it is more likely to fail to grasp an object located deep within a group of multiple overlapping objects than to grasp an object located in front.
[0013] Therefore, the present disclosure can reduce the number of failures of a robot to hold an object by specifying the object to be held by the robot while taking into consideration the visibility of the object. Specific embodiments will be described below.
[0014] (Configuration example of robot control system 1) As shown in Figures 1 and 2, a robot control system 1 according to an embodiment of the present disclosure includes a robot 10 and a robot control device 20. In the robot control system 1, the robot control device 20 controls the robot 10 so that the robot 10 holds a work object 8 and moves from a source 7 to a destination 6. In the example of Figure 2, the work object 8 is an industrial part. The source 7 is a tray that stores the industrial part. The destination 6 is a conveyor that transports the industrial part. The work object 8 is not limited to the object illustrated, and may be various other objects. The source 7 or destination 6 is not limited to the location illustrated, and may be various other locations.
[0015] <Robot 10> As shown in FIGS. 1 and 2, the robot 10 includes an arm 12, a hand 14, a camera 16, and an interface 18.
[0016] The arm 12 is configured to include joints and links. The arm 12 may be configured, for example, as a six-axis or seven-axis vertical articulated robot. The arm 12 may be configured as a three-axis or four-axis horizontal articulated robot or a SCARA robot. The arm 12 may be configured as a two-axis or three-axis Cartesian robot. The arm 12 may be configured as a parallel link robot or the like. The number of axes configuring the arm 12 is not limited to those exemplified.
[0017] The hand 14 is attached to the tip of the arm 12 or at a predetermined position. The hand 14 may include a suction hand configured to be able to suction the work object 8. The suction hand may have one or more suction portions. The hand 14 may include a gripping hand configured to be able to grip the work object 8. The gripping hand may have multiple fingers. The number of fingers of the gripping hand may be two or more. The fingers of the gripping hand may have one or more joints. The hand 14 may include a scooping hand configured to be able to scoop up the work object 8.
[0018] The robot 10 can control the position and posture of the hand 14 by operating the arm 12. The posture of the hand 14 may be represented by an angle specifying the direction in which the hand 14 acts on the work object 8. The posture of the hand 14 is not limited to an angle and may be represented in various other ways, such as a spatial vector. The direction in which the hand 14 acts on the work object 8 may be, for example, the direction in which the hand 14 approaches the work object 8 when it picks up and holds the work object 8, or the direction in which it picks up the work object 8. The direction in which the hand 14 acts on the work object 8 may be the direction in which the hand 14 approaches the work object 8 when it grasps the work object 8 with multiple fingers, or the direction in which it grasps the work object 8. The direction in which the hand 14 acts on the work object 8 is not limited to these examples and may be various other directions.
[0019] The robot 10 may further include a sensor that detects the state of the arm 12 including joints or links, or the state of the hand 14. The sensor may detect information related to the actual position or posture of the arm 12 or hand 14, or the velocity or acceleration of the arm 12 or hand 14, as the state of the arm 12 or hand 14. The sensor may detect a force acting on the arm 12 or hand 14. The sensor may detect a current flowing through a motor that drives a joint, or the torque of the motor. The sensor can detect information obtained as a result of the actual operation of the robot 10. The robot control device 20 can grasp the result of the actual operation of the robot 10 by obtaining the detection results of the sensors.
[0020] The camera 16 may be configured to capture RGB images and to acquire depth data. The robot 10 can control the position and orientation of the camera 16 by operating the arm 12. The orientation of the camera 16 may be expressed by an angle that specifies the direction in which the camera 16 captures images. The orientation of the camera 16 is not limited to an angle and may be expressed in various other ways, such as a space vector. The camera 16 outputs images captured at the position and orientation determined by the operation of the robot 10 to the robot control device 20.
[0021] The camera 16 may be configured to output RAW images, which may then be converted into RGB data by the robot control device 20. The camera 16 may also be, for example, an infrared camera. In this case, for example, when handling fruit as the work object 8, it is possible to determine when the work object 8 is ripe and select the work object 8, or when handling beverages as the work object 8, it is possible to select the work object 8 based on the amount of mineral components they contain.
[0022] The interface 18 may be configured to include a communication device that communicates with the robot control device 20 via a wired or wireless connection. The communication device may be configured to be able to communicate based on various communication standards, such as a local area network (LAN), a wide area network (WAN), RS-232C, or RS-485. The interface 18 may acquire information for controlling the arm 12, the hand 14, or the camera 16 from the robot control device 20. The interface 18 may output to the robot control device 20 the detection results of the state of the arm 12 or the hand 14, or images or depth data captured by the camera 16.
[0023] <Robot Control Device 20> As shown in FIG. 1, the robot control device 20 includes an acquisition unit 22, an estimation unit 24, a work object identification unit 26, and an interface 28.
[0024] The acquisition unit 22 acquires information output from the robot 10. The acquisition unit 22 acquires an image of the work object 8 captured by the camera 16 of the robot 10. The image of the work object 8 is assumed to include a plurality of work objects 8. The acquisition unit 22 may be configured to include a communication interface.
[0025] The estimation unit 24 estimates the visibility of each of the multiple work objects 8 based on images of the multiple work objects 8 acquired from the robot 10. Visibility is the ratio of the area of the portion of an object that is visible and not obscured by other objects to the area of the entire object if it were shown in the image. If the entire object is shown in the image, the visibility of the object is 100%. As will be described later, the estimation unit 24 may estimate the visibility using a trained model 30 (see FIG. 3 ), or may estimate the visibility based on various other algorithms.
[0026] The work object identification unit 26 selects a specific work object to be held by the hand 14 of the robot 10 from among the multiple work objects 8 based on the visibility estimated by the estimation unit 24. The specific work object is the work object 8 identified as the object to be held by the hand 14 of the robot 10 next.
[0027] The estimation unit 24 or the work object identification unit 26 may be configured to include at least one processor. The estimation unit 24 and the work object identification unit 26 may be realized as functions of a single processor. The estimation unit 24 and the work object identification unit 26 may each be realized as functions of different processors. The processor may execute a program that realizes the functions of the estimation unit 24 or the work object identification unit 26. The processor may be realized as a single integrated circuit. An integrated circuit is also called an IC (Integrated Circuit). The processor may be realized as multiple integrated circuits and discrete circuits that are connected to each other in a communicative manner. The processor may be configured to include a CPU (Central Processing Unit). The processor may be configured to include a DSP (Digital Signal Processor) or a GPU (Graphics Processing Unit). The processor may be realized based on various other known technologies.
[0028] The robot control device 20 may further include a storage unit. The storage unit may include an electromagnetic storage medium such as a magnetic disk, or may include a memory such as a semiconductor memory or a magnetic memory. The storage unit may be configured as a hard disk drive (HDD) or a solid state drive (SSD). The storage unit stores various information and programs executed by the processor. The storage unit may function as a work memory for the processor. At least a portion of the storage unit may be included in the processor. At least a portion of the storage unit may be configured as a storage device separate from the robot control device 20.
[0029] The interface 28 may include a communication device configured to be able to communicate via wire or wirelessly, and the communication device may be configured to be able to communicate based on various communication standards.
[0030] The interface 28 may have an input device and an output device. The input device can accept input from, for example, a user of the robot 10. The input device may include, for example, a touch panel or touch sensor, or a pointing device such as a mouse. The input device may include physical keys. The input device may include an audio input device such as a microphone. The input device is not limited to these examples and may include various other devices.
[0031] The output device may include a display device. The output device may display, for example, the work results of the robot 10, the estimation results of the estimation unit 24, or the identification results of the work object identification unit 26 to a user. The display device may include, for example, a liquid crystal display (LCD), an organic electroluminescence (EL) display, an inorganic electroluminescence (EL) display, or a plasma display panel (PDP). The display device is not limited to these displays and may include various other types of displays. The display device may include a light-emitting device such as an LED (light-emitting diode). The display device may include various other devices. The output device may include an audio output device such as a speaker that outputs auditory information such as sound. The output device is not limited to these examples and may include various other devices.
[0032] 1, one robot controller 20 is connected to one robot 10. One robot controller 20 may be connected to two or more robots 10. One robot controller 20 may control only one robot 10, or may control two or more robots 10. The number of robot controllers 20 and robots 10 is not limited to one, and may be two or more.
[0033] The robot controller 20 may be configured to include one or more servers. The robot controller 20 may be configured to cause multiple servers to execute parallel processing. The robot controller 20 does not need to be configured to include a physical housing, and may be configured based on virtualization technology such as a virtual machine or a container orchestration system. The robot controller 20 may be configured using cloud services. When the robot controller 20 is configured using cloud services, it may be configured by combining managed services. In other words, the functions of the robot controller 20 may be realized as cloud services.
[0034] The robot controller 20 may include at least one server group and at least one database group. The server group functions as the estimation unit 24 or the work object identification unit 26. The database group functions as a storage unit. The number of server groups may be one or two or more. When there is one server group, the functions realized by one server group include the functions realized by each server group. The server groups are connected to each other so as to be able to communicate with each other via wired or wireless communication. The number of database groups may be one or two or more. The number of database groups may be increased or decreased as appropriate based on the amount of data managed by the robot controller 20 and the availability requirements required of the robot controller 20. The database groups are connected to each server group so as to be able to communicate with each other via wired or wireless communication. The robot controller 20 may be connected to an external database. An information processing system may be configured that includes the robot controller 20 and an external database.
[0035] Although the robot control device 20 is depicted as a single configuration in FIG. 1 , multiple configurations can be operated as a single system as needed. In other words, the robot control device 20 is configured as a platform with variable capacity. By using multiple configurations as the robot control device 20, even if one configuration becomes inoperable due to an unforeseen event such as a natural disaster, the system can continue to operate using the other configurations. In this case, each of the multiple configurations is connected by a line, whether wired or wireless, and is configured to be able to communicate with each other. The multiple configurations may be configured across cloud services and on-premises environments.
[0036] The robot control device 20 is communicatively connected via a line, whether wired or wireless, to at least one component of the robot control system 1. The robot control device 20 and at least one component of the robot control system 1 are provided with an interface that uses a mutually standard protocol, enabling two-way communication.
[0037] (Example of operation of robot control system 1) In the robot control system 1, the robot control device 20 selects a specific work object to be identified as the work object 8 to be held based on an image of the work object 8 taken by the camera 16 of the robot 10, and controls the robot 10 so that the hand 14 of the robot 10 holds the specific work object.
[0038] Specifically, the acquisition unit 22 of the robot controller 20 acquires images of multiple work objects 8 of the robot 10. The estimation unit 24 of the robot controller 20 estimates the visibility of each of the multiple work objects 8 based on the acquired images. The work object identification unit 26 of the robot controller 20 selects a specific work object to be held by the hand 14 of the robot 10 from the multiple work objects 8 based on the estimated visibility of the work object 8. As described above, the visibility of an object is the ratio of the area of the portion of the object that is actually captured in the image to the area of the object as if the entire object were captured in the image.
[0039] The robot 10 is more likely to hold a work object 8 that is not obscured by other objects and is completely visible than a work object 8 that is obscured by other objects and only partially visible. Furthermore, the robot 10 is more likely to hold a work object 8 that is not overlapped by other objects than a work object 8 that is overlapped by other objects. Furthermore, the robot 10 is more likely to hold a work object 8 that is within the angle of view of the camera 16 of the robot 10 than a work object 8 that is located at the edge of the angle of view of the camera 16 and is cut off from the image. Therefore, the work object identification unit 26 may select a work object 8 with high visibility as the identified work object.
[0040] The estimation unit 24 may classify the work objects 8 by type or by individual from images of the work objects 8. The estimation unit 24 may estimate the classification degree of each of the classified work objects 8. In other words, the estimation unit 24 may be able to estimate the classification degree of each of multiple work objects 8 from images of the work objects 8. The classification degree is an index that represents the degree of classification accuracy of the work objects 8. The classification accuracy may be expressed as the degree of confidence or likelihood of the classification of the work objects 8.
[0041] The work object identification unit 26 may select a specific work object from among the multiple work objects 8 based on the visibility and classification level of each of the multiple work objects 8.
[0042] The estimation unit 24 may generate a mask for each of the multiple work objects 8 based on an image capturing the multiple work objects 8. The estimation unit 24 may estimate the visibility of each of the multiple work objects 8 based on the mask for each of the multiple work objects 8. The estimation unit 24 may also estimate the accuracy of mask generation. That is, the estimation unit 24 may be capable of estimating the accuracy of mask generation. The accuracy of mask generation corresponds to the degree of overlap between the mask generated by the estimation unit 24 and the correct mask. That is, the closer the mask generated by the estimation unit 24 is to the correct mask, the higher the accuracy of mask generation. The estimation unit 24 may calculate the matching rate between the generated mask and the correct mask as the accuracy of mask generation.
[0043] The work object identification unit 26 may select a specific work object from among the multiple work objects 8 based on the visibility of each of the multiple work objects 8 and the accuracy of mask generation.
[0044] The estimation unit 24 may estimate height information for each of the multiple work objects 8 from an image capturing the work objects 8. In other words, the estimation unit 24 may be capable of estimating height information for each of the multiple work objects 8 from an image capturing the work objects 8. The height information may be expressed as the distance from the camera 16 of the robot 10 to the work object 8 when viewed from the camera 16.
[0045] The estimation unit 24 may estimate height information based on the area of each of multiple work objects 8 included in an image capturing the work objects 8. The estimation unit 24 may estimate height information of the work objects 8 by assuming that a work object 8 with a larger area is located in the foreground. The acquisition unit 22 may acquire depth data from the camera 16 of the robot 10. The estimation unit 24 may estimate height information of the work object 8 based on the depth data.
[0046] The work object identification unit 26 may select a specific work object from among the multiple work objects 8 based on the visibility and height information of each of the multiple work objects 8. The robot 10 is more likely to hold a work object 8 that is located closer than a work object 8 that is located farther away. The work object identification unit 26 may select a work object 8 that is located closer than the multiple work objects 8 as the specific work object based on the height information.
[0047] Furthermore, the work object identification unit 26 may select a specific work object from among the work objects 8 that appear in the image and that are located toward the center of the image.
[0048] (Trained Model 30) As described above, the estimation unit 24 of the robot control device 20 may estimate the visibility of the work object 8 shown in the image using the trained model 30. The trained model 30 is, for example, AI (artificial intelligence). The trained model 30 is generated by performing predetermined machine learning on the pre-trained model. The predetermined machine learning may be, for example, deep learning. The pre-trained model is an AI model in a state in which the item to be inferred is not yet inferable. For example, when the trained model 30 estimates the visibility of the work object 8, the pre-trained model refers to an AI model that is not yet in a state in which it can estimate the visibility of the work object 8. For example, the pre-trained model includes an AI model that has not undergone training processing or an AI model that has undergone training processing capable of inferring things other than visibility. Note that, although an example in which the estimation unit 24 is AI is described in this disclosure, the estimation unit 24 may be implemented using a template matching technique or a combination of AI and template matching.
[0049] The estimation unit 24 inputs an image for inference to the trained model 30. The image for inference is an image from which the work object 8 to be held by the robot 10 is extracted and the visibility is estimated. The estimation unit 24 estimates the visibility of the work object 8 depicted in the image for inference. In other words, the trained model 30 estimates the visibility of the object depicted in the image for inference in response to the input image for inference.
[0050] 3 , the trained model 30 may include a feature extraction unit 31, a mask generation unit 32, a visibility estimation unit 33, a classification generation unit 34, and a generation accuracy estimation unit 35. The trained model 30 may not include the classification generation unit 34. The trained model 30 may not include the generation accuracy estimation unit 35.
[0051] The trained model 30 may be configured to include a CNN (Convolutional Neural Network) or R-CNN having multiple layers. Convolution processing based on predetermined weighting coefficients is performed in each layer of the CNN on information input to the trained model. The weighting coefficients are updated during training of the trained model. The trained model may be configured to include a fully connected layer. The trained model may be configured using VGG16 or ResNet50. The trained model may be configured as a transformer. The trained model 30 may be configured to perform pixel-level processing on input images. The trained model 30 may be configured to combine various backbones and heads. The trained model 30 may be configured by adding a new backbone or head to another AI model. The trained model is not limited to these examples and may be configured to include various other models.
[0052] The operation of each component of the trained model 30 will be described below. In the following description, it is assumed that the image input to the trained model 30 is an image showing recognition objects 41 to 44, as exemplified in FIG. 4A. It is assumed that the recognition objects 41 to 43 are parts that correspond to the work object 8. It is assumed that the recognition object 44 is a pencil that does not correspond to the work object 8.
[0053] 4B , the feature extraction unit 31 extracts features of objects in an image and converts them into regions of interest (ROIs). In other words, the feature extraction unit 31 extracts features of each of multiple objects based on an image containing the multiple objects. The feature extraction unit 31 outputs the regions of interest of the objects in the image to the mask generation unit 32 and the classification generation unit 34.
[0054] The feature extraction unit 31 extracts features of the recognition target 41 that corresponds to the work target 8 and converts them into a region of interest of the recognition target 41 surrounded by a bounding box 41F. The feature extraction unit 31 extracts features of the recognition target 42 that corresponds to the work target 8 and converts them into a region of interest of the recognition target 42 surrounded by a bounding box 42F. The feature extraction unit 31 extracts features of the recognition target 43 that corresponds to the work target 8 and converts them into a region of interest of the recognition target 43 surrounded by a bounding box 43F. The feature extraction unit 31 extracts features of the recognition target 44 that does not correspond to the work target 8 and converts them into a region of interest of the recognition target 44 surrounded by a bounding box 44F. In this embodiment, the bounding boxes 41F to 44F are generated by the classification generation unit 34, which will be described later. The feature extraction unit 31 may also generate the bounding boxes 41F to 44F.
[0055] 4C , based on the region of interest output from the feature extraction unit 31. In other words, the mask generation unit 32 generates a plurality of masks showing each of the plurality of objects based on the features of each of the plurality of objects.
[0056] The mask generation unit 32 outputs the generated mask to the trained model 30. The mask generation unit 32 also outputs the generated mask to the visibility estimation unit 33 and the generation accuracy estimation unit 35.
[0057] 4C, a mask is generated for each individual object, not for each type of object. That is, each of recognition targets 41 to 43 corresponding to work object 8 is distinguished as an individual part corresponding to work object 8. Mask 41M is a mask for recognition target 41 corresponding to work object 8. Mask 42M is a mask for recognition target 42 corresponding to work object 8. Mask 43M is a mask for recognition target 43 corresponding to work object 8. Mask 44M is a mask for recognition target 44 which does not correspond to work object 8.
[0058] <Visibility Estimation Unit 33> The visibility estimation unit 33 estimates the visibility of an object corresponding to the mask generated by the mask generation unit 32. In other words, the visibility estimation unit 33 estimates the visibility of each of a plurality of objects based on a plurality of masks.
[0059] The visibility estimation unit 33 outputs the visibility score to the trained model 30. The visibility score is an index representing the degree of visibility of an object corresponding to a mask. The visibility estimation unit 33 may output the visibility value of the object corresponding to the mask itself as the visibility score. The visibility estimation unit 33 may output a value according to the degree of visibility of the object corresponding to the mask as the visibility score. As described above, the visibility of an object is the ratio of the area of a portion of the object that is not obscured by other objects in an actual image to the area of the object that would appear in an image if the object were not obscured by other objects from a certain shooting viewpoint.
[0060] The visibility estimation unit 33 may estimate the visibility of the mask of each object based on the generation result of the mask of each object illustrated in Fig. 4C. As shown in Fig. 5, in the image of the recognition target 42, a portion of the recognition target 42 is hidden by the recognition targets 41 and 43. Specifically, the recognition target 42 includes a portion 42A hidden behind the recognition target 41 and a portion 42B hidden behind the recognition target 43. Therefore, the mask 42M of the recognition target 42 (see Fig. 4C) is not a mask that represents the entire view of the recognition target 42.
[0061] The visibility estimation unit 33 estimates the area of portions 42A and 42B of the recognition target 42 that are hidden by the other recognition targets 41 and 43. The visibility estimation unit 33 calculates the area of portions of the recognition target 42 that are not hidden by the other recognition targets 41 and 43 and appear in the image. The visibility estimation unit 33 calculates the total area of portions 42A and 42B of the recognition target 42 that are hidden by the other recognition targets 41 and 43 and the area of portions of the recognition target 42 that are not hidden by the other recognition targets 41 and 43 and appear in the image as the area of the recognition target 42 when the entire view is shown in the image. The visibility estimation unit 33 can estimate the visibility by dividing the area of portions of the recognition target 42 that are not hidden by the other recognition targets 41 and 43 and appear in the image by the area of the recognition target 42 when the entire view is shown in the image.
[0062] <Classification Generation Unit 34> The classification generation unit 34 determines into which class an object is classified based on the region of interest of the object output from the feature extraction unit 31. In other words, the classification generation unit 34 classifies the object into classes. The region of interest of the object is classified into a class that specifies the type of object, such as a part that corresponds to the work object 8, or a pencil that does not correspond to the work object 8. The region of interest of the object is classified into a class that specifies an attribute other than that of the object to be recognized, such as the background. In other words, the classification generation unit 34 classifies each of a plurality of objects into a plurality of classes based on the features of the object.
[0063] When an object may be classified into a plurality of classes, each of which identifies a specific type of object, the classification generating unit 34 may calculate the probability that the object will be classified into each class. The classification generating unit 34 may classify the object into the class with the highest probability. The classification generating unit 34 may calculate the classification accuracy of the object. The classification accuracy of the object may be expressed as the probability that the object will be classified into the class classified by the classification generating unit 34.
[0064] The classification generation unit 34 outputs a classification score corresponding to the above-mentioned classification degree as an output of the trained model 30. The classification score is an index representing the degree of certainty (certainty) of the classification of an object. In addition to the certainty, the classification generation unit 34 may output a numerical value (likelihood) representing the likelihood as the classification score. The classification generation unit 34 may output the certainty of the object itself as the classification score. The classification generation unit 34 may output a value according to the certainty of the object as the classification score.
[0065] The classification generator 34 may train a bounding box corresponding to the region of interest of each object in the image, and the classification generator 34 may then output the bounding boxes as the trained model 30.
[0066] <Generation accuracy estimation unit 35> The generation accuracy estimation unit 35 estimates the generation accuracy of the mask generated by the mask generation unit 32. In other words, the generation accuracy estimation unit 35 estimates the generation accuracy of each of the multiple masks based on the multiple masks. As described above, the generation accuracy of the mask corresponds to the degree of overlap between the mask generated by the mask generation unit 32 and the correct mask. In other words, the closer the mask generated by the mask generation unit 32 is to the correct mask, the higher the generation accuracy of the mask. The generation accuracy estimation unit 35 outputs a generation accuracy score as the trained model 30. The generation accuracy score is an index representing the level of mask generation accuracy. The generation accuracy estimation unit 35 may output the matching rate between the mask generated by the mask generation unit 32 and the correct mask as the generation accuracy score.
[0067] <Other Blocks> The trained model 30 may include an area calculation unit that calculates the area of each of multiple objects included in an image capturing the object. The area calculation unit may be capable of calculating the area of each of multiple objects in the image from the image capturing the object. The area calculation unit may be located after the mask generation unit 32 and calculate the area of the object based on a mask generated by the mask generation unit 32. The area calculation unit may be located after the feature extraction unit 31 in parallel with the mask generation unit 32 and calculate the area of the object based on a region of interest corresponding to the feature extracted by the feature extraction unit 31. The area calculation unit may calculate the area by, for example, counting the number of pixels in the generated mask.
[0068] The trained model 30 may include a height estimation unit that estimates height information for each of a plurality of objects from an image capturing the objects. The height estimation unit may be capable of estimating height information for each of a plurality of objects from an image capturing the objects. The height information may be expressed as a distance from a camera capturing the objects to the objects. The height estimation unit may be located after the area calculation unit and estimate the height information for the objects based on the area of the objects. The height estimation unit may be located after the mask generation unit 32 and estimate the height information for the objects based on a mask generated by the mask generation unit 32 and depth data. The height estimation unit may be located after the feature extraction unit 31 and estimate the height information for the objects based on a region of interest corresponding to a feature extracted by the feature extraction unit 31 and depth data.
[0069] <Generation of Trained Model 30> The trained model 30 is generated by performing learning based on training data including a plurality of training images depicting an object. The learning for generating the trained model 30 may be performed by an external device or by the robot control device 20. In this embodiment, the learning for generating the trained model 30 is performed by a learning device. The learning device may include an external device or the robot control device 20, etc. The learning device may include a processor, a storage unit, an interface, etc.
[0070] As the learning images, images of actual objects may be used. Alternatively, synthetic images may be used as the learning images. Note that a synthetic image is, for example, an image generated based on data that virtually models the appearance of an object. In this embodiment, it is assumed that synthetic images are used as the learning images.
[0071] The learning data may include annotation data associated with the learning image. When generating the visibility estimation unit 33, the learning device acquires, as annotation data, the proportion of the entire image of the object that is visible in the composite image. The learning device may calculate, based on the image of the object included in the composite image, the proportion of the entire image of the object that is visible in the composite image. The learning device may associate the calculated proportion with the composite image as annotation data.
[0072] The learning device may estimate the portion of the object that is obscured by other objects based on the image of the object included in the composite image, and calculate the proportion of the portion of the object that is obscured by other objects relative to the entire image of the object. The learning device may associate the calculated proportion with the composite image as annotation data.
[0073] If the object does not fit within the angle of view of the composite image, the image of the object may be cut off at the edge of the composite image. The learning device may estimate the portion of the object that extends beyond the angle of view of the composite image based on the image of the object included in the composite image, and calculate the proportion of the portion of the object that extends beyond the angle of view of the composite image relative to the entire image of the object. The learning device may associate the calculated proportion with the composite image as annotation data. The angle of view of the composite image corresponds to the angle of view of camera 16.
[0074] Portions of the object image that extend beyond the angle of view of the composite image can be identified by generating another image with a wider angle of view than the composite image. Specifically, as shown in FIG. 6 , the learning device may generate a base composite image extending within a range 51 indicated by a dashed rectangle, and then obtain an excerpted composite image cropped from the base composite image within a range 52 indicated by a dashed-dotted rectangle. The base composite image is assumed to be an image in which multiple objects modeled using 3-dimensional computer graphics (3DCG) technology based on, for example, computer-aided design (CAD) data of the objects are randomly overlapped within the specific range 51. The excerpted composite image is assumed to be an image in which multiple objects extending within the specific range 51 are captured by a camera 16 (virtual camera 53), and portions of the multiple objects are cropped from the base composite image within the angle of view (range 52) of the camera 16 that is narrower than the specific range 51. The base composite image is, for example, about two to three times larger than the excerpt composite image, and is large enough to fit all objects that extend beyond the angle of view of the excerpt composite image within the angle of view.
[0075] Specifically, suppose the image shown in FIG. 7 is generated as the base composite image. Furthermore, suppose the image shown in FIG. 8 is generated as the excerpt composite image, by cutting out the portion enclosed by the dashed-dotted rectangular frame shown in B in FIG. 7. Assume that recognition targets 61 to 65, which are industrial parts, are shown in FIGS. 7 and 8. In FIG. 8, the entire image of recognition target 61 is included in the excerpt composite image. The image of recognition target 62 includes a portion 62A that is cut out from the excerpt composite image and a portion 62B that is hidden by another object. The image of recognition target 63 includes a portion 63A that is cut out from the excerpt composite image. The image of recognition target 64 includes a portion 64A that is cut out from the excerpt composite image and a portion 64B that is hidden by another object. The image of recognition target 65 includes a portion 65A that is cut out from the excerpt composite image.
[0076] The learning device can acquire the area of the entire image of the recognition targets 63 and 65 from each first mask image of each object rendered so that the multiple objects are arranged similarly to the base composite image based on three-dimensional modeling data of the multiple objects. The learning device can also calculate the area of the portions 63A and 65A cut off from the excerpt composite image, for example, from each first mask image and each second mask image rendered so that the multiple objects are arranged similarly to the excerpt composite image. Specifically, first, each first mask image of each object is generated so as to reproduce the arrangement of the multiple objects in the base composite image. Each first mask image of each object is generated in an uncut state. Next, each second mask image of each object is generated so as to reproduce the arrangement of the multiple objects in the excerpt composite image. Each second mask image of each object is generated in an uncut state. Therefore, by calculating the area of each first mask image and each second mask image of each object, the learning device can calculate, for the recognition targets 63 and 65, the proportion of the object's entire image that extends beyond the angle of view of the excerpt composite image. Specifically, the learning device may calculate, based on each first mask image and each second mask image, the area of object X1, a portion of which appears in the excerpt composite image that extends beyond the angle of view of the excerpt composite image, and the area of object X1 in the base composite image, and calculate the proportion of the portion of object X1 that is cut off from the excerpt composite image. Conversely, the learning device may calculate the proportion of the entire image of object X1 that is not cut off from the excerpt composite image and is visible in the excerpt composite image. The learning device may associate the calculated proportions with the recognition targets 62 and 64 as annotation data.
[0077] The learning device can calculate the areas of portions 62A and 64A of recognition targets 62 and 64 that are cut off from the excerpt composite image in the same manner as for recognition targets 63 and 65. Furthermore, the learning device may estimate the areas of portions 62B and 64B of recognition targets 62 and 64 that are hidden by other objects, for example, based on three-dimensional modeling data of the multiple objects, from first mask images of each object rendered so that the multiple objects are arranged similarly to the base composite image and second mask images of each object rendered so that the multiple objects are arranged similarly to the excerpt composite image. Specifically, first, each first mask image of each object is generated to reproduce the arrangement of the multiple objects in the base composite image. The generated first mask images of each object are generated assuming that the objects are not hidden by other objects. Next, each second mask image of each object is generated to reproduce the arrangement of the multiple objects in the excerpt composite image. The generated second mask images of each object are generated with the portions overlapping with other objects missing. Therefore, by calculating and comparing the areas of the first mask image and the second mask image of each object, the learning device can calculate the proportion of the entire image of the object that is hidden by other objects for the recognition targets 62 and 64. Therefore, the learning device may estimate the proportion of the entire image of the object that extends beyond the angle of view of the extracted composite image and the proportion of the entire image of the object that is hidden by other objects for the recognition targets 62 and 64. The learning device may associate the estimated proportions with the recognition targets 62 and 64 as annotation data.
[0078] The learning device can calculate the areas of the portions 62B and 64B of the recognition targets 62 and 64 that are hidden by other objects in the excerpted composite image based on the first mask image and the second mask image, as described above. Specifically, the learning device may calculate the area of the object X2 in the base composite image and the area of the object X2 in an image obtained by removing the portion of the object X2 that is hidden by other objects in the excerpted composite image based on the first mask image and the second mask image, and calculate the proportion of the portion of the object X2 that is hidden by other objects in the entire image of the object X2. Conversely, the learning device may calculate the proportion of the portion of the object X2 that is visible and not hidden by other objects in the entire image of the object X2. The learning device may associate the calculated proportions with the recognition targets 62 and 64 as annotation data.
[0079] By calculating the area of the portions of a plurality of objects hidden by other objects based on three-dimensional modeling data of the objects, the accuracy of calculating the proportion of the entire image of the object that is hidden by other objects is improved. In other words, the accuracy of the annotation data is improved. By improving the accuracy of the annotation data, the accuracy of the visibility estimation by the trained model 30 generated by performing learning using the annotation data as training data is improved.
[0080] As described above, annotation data may be generated by calculating, for example, the percentage of the portion of the training image where the object is hidden by other objects or the percentage of the portion of the training image where the object is cut off, based on the base composite image and the excerpt composite image. Furthermore, by performing training using the generated annotation data together with the excerpt composite image as training data, the accuracy of visibility estimation by the generated trained model 30 is improved. The training device may calculate both the percentage of the portion of the object where the object is hidden by other objects and the percentage of the portion of the object where the object is cut off in the training image.
[0081] In the above description, the portions of an object hidden or cut off by other objects are calculated from the first mask image of each object rendered based on the three-dimensional modeling data of the objects so that the objects are arranged similarly to the base composite image, and the second mask image of each object rendered so that the objects are arranged similarly to the excerpt composite image. However, the present disclosure is not limited to this. That is, for example, instead of generating each first mask image, the portions of an object hidden or cut off by other objects may be calculated from the second mask image of each object rendered so that the objects are arranged similarly to the excerpt composite image. In this case, instead of the first mask images, multiple third mask images may be generated based on the three-dimensional modeling data of the multiple objects, assuming that the objects are arranged in the same orientation as each object in the excerpt composite image. In this case, the third mask images are generated so that the objects are not hidden or cut off by other objects. Then, the portions of an object hidden or cut off by other objects may be calculated based on the second mask image and the third mask image.
[0082] Although the above describes an example in which an image is generated and the portions of an object that are hidden or cut off by other objects are calculated, it is not necessary to generate image data. In this case, for example, only the second mask image is generated, and a simulation is performed based on three-dimensional modeling data of the multiple objects instead of the third mask image, in which the objects are arranged in the same orientation as each object in the extracted composite image, and the areas of the objects that are not hidden or cut off by other objects are calculated as the simulation results.
[0083] In the above example, each second mask image is generated with portions of the object hidden by other objects or portions cut off missing. Each second mask image may be generated with portions of the object hidden by other objects or portions cut off missing. The area of each second mask image with portions cut off may be calculated by simulation based on the overlapping relationship of each object.
[0084] The images used by the training device may include a base composite image or an excerpt composite image, an RGB image or a monochrome image, and a mask image. The training device may also calculate a percentage of the object's overall image that is visible in the excerpt composite image based on the number of pixels in the object's image.
[0085] The learning device may acquire, as the learning image, a mask image of the object generated based on an image of the object. The learning device may also acquire, as annotation data, the ratio of the range of the mask image of the object to the range of the entire mask of the object. The learning device may perform learning to generate a trained model 30 based on the mask image and the annotation data.
[0086] The learning device may acquire an image of the object and an overall mask of the object, generate a mask image based on the image of the object, and calculate the ratio of the area of the mask image to the overall mask as annotation data.
[0087] Although the above describes an example in which annotation data is generated by a learning device, annotation data can also be prepared by a user of the trained model 30. That is, the user may directly input the percentage of each object that appears in an image that is visible.
[0088] When generating the feature extraction unit 31 and the classification generation unit 34, the learning device acquires, as annotation data, correct label information indicating the identity of the objects appearing in the learning images. Then, the learning device inputs the learning images to the feature extraction unit 31 and the classification generation unit 34 and trains them so that the classification results obtained match the correct label information, thereby generating the feature extraction unit 31 and the classification generation unit 34. Note that the mask generation unit 32 may generate a mask by determining the outer edges of individual objects having the features extracted by the feature extraction unit 31.
[0089] When generating the generation accuracy estimation unit 35, the learning device acquires, as annotation data, correct mask information in which the contours extracted from the modeling data of each object appearing in the learning image are converted into mask data. Then, the learning image is input to the feature extraction unit 31 and the mask generation unit 32, and the resulting mask is trained to match the correct mask information, thereby generating the generation accuracy estimation unit 35.
[0090] The learning device may acquire an image of the work object 8 of the robot 10 as an image of the object to be used as learning data. That is, the learning data may include an image of the work object 8 of the robot 10 as an image of the object. The trained model 30 generated by performing learning using the image of the work object 8 as learning data may output the mask and visibility of the work object 8 to the robot control device 20. The robot control device 20 may control the robot 10 based on the mask and visibility of the work object 8 output from the trained model 30.
[0091] <Example of Operation of Robot Control Device 20 Based on Output of Trained Model 30> As described above, the trained model 30 outputs a mask from the mask generation unit 32 and outputs a visibility score from the visibility estimation unit 33. The robot control device 20 selects a work object 8 to be held by the robot 10 based on the visibility score output from the trained model 30. The robot control device 20 determines the position or posture at which the work object 8 is held by the hand 14 of the robot 10, based on the mask corresponding to the selected work object 8 output from the trained model 30.
[0092] A work object 8 with a high visibility score is likely to be located closer than other objects. By selecting a work object 8 with a high visibility score, the robot control device 20 can reduce failures of the robot 10 to hold the work object 8.
[0093] If the trained model 30 includes a classification generation unit 34, the classification generation unit 34 outputs a classification score. The robot control device 20 may select a work object 8 to be held by the robot 10 based on the visibility score and the classification score. The robot control device 20 may calculate the product of the visibility score and the classification score, and select an object with a larger calculated value as the work object 8. The higher the score calculated as the product of the visibility score and the classification score for a certain object, the better the classification of that object with respect to other objects and the larger the portion of the object that is visible in the overall view.
[0094] If the trained model 30 includes a generation accuracy estimation unit 35, the generation accuracy estimation unit 35 outputs a generation accuracy score. The robot control device 20 may select a work object 8 to be held by the robot 10 based on the visibility score and the generation accuracy score. The robot control device 20 may calculate the product of the visibility score and the generation accuracy score, and select an object with a larger calculated value as the work object 8. The higher the score calculated as the product of the visibility score and the generation accuracy score for a certain object, the better the recognition of the object and the more of the object that is visible in the overall view.
[0095] The robot control device 20 may select a work object 8 to be held by the robot 10 based on the visibility score, classification score, and generation accuracy score. The robot control device 20 may calculate the product of the visibility score, classification score, and generation accuracy score, and select an object with a larger calculated value as the work object 8. The higher the score calculated as the product of the visibility score, classification score, and generation accuracy score for a certain object, the better the classification of the object with respect to other objects, the better the recognition of the object, and the more of the object that is visible in the overall view.
[0096] The trained model 30 may be configured to output the visibility score, the classification score, and the generation accuracy score separately. The trained model 30 may be configured to output the product of the visibility score, the classification score, and the generation accuracy score. The trained model 30 may be configured to output at least one of the product of the visibility score and the classification score or the product of the visibility score and the generation accuracy score.
[0097] In addition, when the trained model 30 outputs evaluation scores such as a visibility score, a classification score, and a generation accuracy score, the robot control system 1 may further include a display device, and the display device may display the evaluation scores as the basis for selecting a specific work object. The evaluation scores may be displayed for each work object 8 appearing in the image. The evaluation scores may be displayed superimposed on the image of each work object 8. Furthermore, the evaluation score may be, for example, one of the visibility score, the classification score, and the generation accuracy score, or an integrated score that takes into account multiple scores from the visibility score, the classification score, and the generation accuracy score, depending on user input. In addition, the display device may display data output by the robot control device 20, the trained model 30, and the robot control system 1, in addition to the integrated score. The display device may have a configuration similar to that of a display device.
[0098] As described above, the trained model 30 of the present disclosure can determine the visibility of an object and can improve the performance or reliability of the object recognition process. The robot control device 20 equipped with the trained model 30 can reduce failures of the robot 10 to hold the work object 8 by selecting the work object 8 taking into consideration the visibility of the object.
[0099] (Example of Procedure of Method for Generating Trained Model 30) The learning device may execute a method for generating the trained model 30, including the steps of the flowchart illustrated in Fig. 9. The method for generating the trained model 30 may be realized as a program for generating the trained model 30, which is executed by a processor constituting the learning device. The program for generating the trained model 30 may be stored in a non-transitory computer-readable medium.
[0100] The learning device acquires training images (step S1). The learning device acquires each annotation data required to generate a trained model 30 for an object shown in the training images (step S2). The learning device generates the trained model 30 by performing training using the training data including the annotation data and training images acquired in step S2 (step S3). After executing the procedure of step S3, the learning device ends execution of the procedure of the flowchart in FIG. 9.
[0101] (Example of Procedure of Robot Control Method) The robot control device 20 may execute a robot control method exemplified in Fig. 10 to control the robot 10. The robot control method may be realized as a robot control program executed by a processor constituting the robot control device 20. The robot control program may be stored on a non-transitory computer-readable medium.
[0102] The robot controller 20 acquires an image of the work object 8 (step S11). The robot controller 20 generates a mask of the work object 8 based on the image of the work object 8 (step S12). The robot controller 20 estimates the visibility of the work object 8 (step S13). The robot controller 20 estimates the classification level of the work object 8 (step S14). The robot controller 20 identifies the work object 8 to be held by the hand 14 of the robot 10 based on the visibility of the work object 8 (step S15). The robot controller 20 controls the robot 10 to hold the specific work object identified in the procedure of step S15 by the hand 14 of the robot 10 (step S16). After executing the procedure of step S16, the robot controller 20 ends the execution of the procedure of the flowchart in FIG. 10.
[0103] (Summary) As described above, in the robot control system 1 according to this embodiment, the robot control device 20 calculates a score that takes into account the visibility of each of the multiple work objects 8 from an image capturing the multiple work objects 8. The robot control device 20 selects a specific work object based on the score that takes into account visibility, and controls the robot 10 to hold the specific work object with the hand 14 of the robot 10. By selecting a specific work object based on the score that takes into account visibility, the number of times the hand 14 of the robot 10 fails to hold the work object 8 is reduced. In other words, the score calculated for the work object 8 correlates with the probability of success in holding the work object 8. As a result, the usefulness of the score calculated for the work object 8 is improved.
[0104] Other Embodiments Other embodiments will be described below.
[0105] The trained model 30 may perform object recognition on a plurality of irregularly shaped objects as the work target 8. The trained model 30 may estimate the visibility of each of the plurality of irregularly shaped objects based on an image of the objects. As a result, it is possible to selectively grasp an object having a certain visibility and a certain area, for example, or to inspect or evaluate whether a plurality of irregularly shaped objects have a certain area.
[0106] The trained model 30 may generate multiple masks for each amorphous object. The trained model 30 may calculate the number of pixels in at least one of the multiple masks for each of the multiple objects. As shown as a graph in FIG. 11, the estimation unit 24 may calculate an identification number for identifying each mask or object and the number of pixels (number of pixels) in each mask, and display them on a display device or display apparatus. In the graph in FIG. 11, the vertical axis represents the number of pixels in k (kilo). For example, a pixel count of 100 represents 100,000 pixels (100k). The horizontal axis represents the number assigned to each mask or object. In this way, a user of the trained model 30 can understand the results of object recognition or inspection, etc.
[0107] The trained model 30 may calculate the area of an object when the entire image of the object is visible, based on the visibility of the object. For example, if the visibility of the object is 50%, the area of the object when the entire image of the object is visible is twice the number of pixels of the mask.
[0108] The trained model 30 may classify the size of each mask into three categories: small, medium, and large. The number of categories of mask size is not limited to three, and may be two, four, or more.
[0109] 12, a histogram of the size of each amorphous object or each object mask may be generated and displayed on a display device or display device, allowing a user of the trained model 30 to understand the results of object recognition, inspection, etc.
[0110] The sizes constituting the histogram may be estimated areas when the entire image of the object is visible based on the visibility of the object, or may be areas actually appearing in the image. Furthermore, when a histogram is generated using a mask for each object based on the area actually appearing in the image, the size of the mask for an object that is overlapped by other objects will be smaller than the size of the mask for an object that is completely visible without being overlapped by other objects. By extracting and statistically processing only the mask sizes of objects that are estimated to be completely visible without being overlapped by other objects, the distribution of the sizes of all objects, including those located below the overlap, can be statistically estimated.
[0111] In addition to the above, various other processes can be performed. For example, after generating a mask for each object, the trained model 30 may assign a ranking number to each mask based on the level of visibility, and display the mask on a display device or a display apparatus. As a result, a user of the trained model 30 can understand the results of object recognition or inspection.
[0112] The above has described an embodiment of the robot control system 1, but embodiments of the present disclosure can also be embodied as a method or program for implementing a system or device, as well as a storage medium on which a program is recorded (for example, an optical disk, a magneto-optical disk, a CD-ROM, a CD-R, a CD-RW, a magnetic tape, a hard disk, or a memory card).
[0113] Furthermore, the implementation form of the program is not limited to application programs such as object code compiled by a compiler or program code executed by an interpreter, but may also be in the form of a program module incorporated into an operating system. Furthermore, the program may or may not be configured so that all processing is performed solely by the CPU on the control board. The program may also be configured so that part or all of it is executed by another processing unit mounted on an expansion board or expansion unit added to the board as needed.
[0114] Although the embodiments of the present disclosure have been described based on the drawings and examples, it should be noted that those skilled in the art could make various modifications or alterations based on the present disclosure. Therefore, it should be noted that these modifications or alterations are included in the scope of the present disclosure. For example, the functions included in each component can be rearranged so as not to cause logical inconsistencies, and multiple components can be combined or divided into one.
[0115] All of the features described in this disclosure and / or all steps of all of the disclosed methods or processes may be combined in any combination except combinations in which these features are mutually exclusive. Furthermore, each feature described in this disclosure may be replaced by an alternative feature serving the same, equivalent, or similar purpose, unless expressly denied. Thus, unless expressly denied, each disclosed feature is only one example of a generic series of identical or equivalent features.
[0116] Furthermore, embodiments of the present disclosure are not limited to the specific configurations of any of the above-described embodiments, but rather extend to any novel feature or combination thereof described herein, or any novel method or process step or combination thereof described herein.
[0117] 1 Robot control system 6 Destination 7 Source 8 Work object 10 Robot (12: Arm, 14: Hand, 16: Camera, 18: Interface) 20 Robot control device (22: Acquisition unit, 24: Estimation unit, 26: Work object identification unit, 28: Interface) 30 Trained model (31: Feature extraction unit, 32: Mask generation unit, 33: Visibility estimation unit, 34: Classification generation unit, 35: Generation accuracy estimation unit) 41 to 44 Recognition object (41F to 44F: Bounding box, 40M to 44M: Mask, 42A, 42B: Part) 51, 52 Range 53 Virtual camera 61 to 65 Recognition object (62A, 62B, 63A, 64A, 64B, 65A: Part)
Claims
1. an acquisition unit that acquires images of a plurality of work targets of the robot; an estimation unit that estimates the visibility of each of the plurality of work objects based on the image; a work object specifying unit that selects a specific work object for the robot from among the plurality of work objects based on the visibility; A robot control device comprising:
2. the estimation unit is capable of estimating a classification level of each of the plurality of work objects from the image; the work object identification unit selects the identified work object based on the visibility and the classification level. The robot control device according to claim 1 .
3. the estimation unit is capable of estimating height information of the plurality of work objects from the image, the work object identification unit selects the identified work object based on the visibility and the height information. The robot control device according to claim 1 .
4. the estimation unit estimates the height information based on an area of each of the plurality of work objects included in the image. The robot control device according to claim 3 .
5. A robot control device and a display device according to any one of claims 1 to 4, The display device displays a score indicating the basis for selecting the specific task object selected by the robot control device.
6. a feature extraction unit that extracts features of each of a plurality of objects based on an image of the plurality of objects; a mask generation unit that generates a plurality of masks representing each of the plurality of objects based on the characteristics of each of the plurality of objects; a visibility estimation unit that estimates a visibility of each of the plurality of objects based on the plurality of masks; A trained model with
7. The trained model of claim 6 , further comprising a classification generator that classifies each of the plurality of objects based on the features.
8. The trained model according to claim 6 or 7, further comprising a generation accuracy estimation unit that estimates generation accuracy of each of the plurality of masks based on the plurality of masks.
9. A method for generating a trained model, which performs training to generate a trained model that estimates the visibility of an object appearing in an inference image when the inference image is input, based on training data including a plurality of training images that show the object.
10. acquiring a synthetic image as the learning image; acquiring, as annotation data, a proportion of the entire image of the object that is visible in the composite image; The trained model generation method according to claim 9 , further comprising: performing learning to generate the trained model based on the synthetic image and the annotation data.
11. The method for generating a trained model according to claim 10, wherein the annotation data is calculated based on a proportion of the object that is hidden by another object.
12. The method for generating a trained model according to claim 10, wherein the annotation data is calculated based on a proportion of a portion of the object that extends beyond an angle of view of the composite image.
13. The method for generating a trained model according to claim 10, wherein the annotation data is calculated based on the number of pixels in a portion where the object is visible in the synthetic image.
14. acquiring a mask image of the object as the learning image; acquiring, as annotation data, a ratio of a range of a mask image of the object to a range of a correct mask of the object; The trained model generation method according to claim 9 , further comprising: performing learning to generate the trained model based on the mask image and the annotation data.
15. the learning data includes an image of a work target of the robot as the image of the object; The trained model outputs the mask and the visibility of the work object to a robot control device that controls the robot based on the mask and the visibility of the work object. A method for generating a trained model according to any one of claims 9 to 14.