Trained model generation method, trained model generation device, trained model, and holding mode estimation device

The learned model generation method improves robot grasping accuracy by using multiple inference models to classify and determine optimal grip positions and forces, addressing the inefficiencies in holding arbitrary-shaped objects.

JP7717169B2Active Publication Date: 2025-08-01KYOCERA CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023544012
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-08-27
Filing Date
2022-08-26
Publication Date
2025-08-01
Estimated Expiration
2042-08-26

AI Technical Summary

Technical Problem

Existing methods for robot hands to hold objects of arbitrary shapes lack accuracy in estimating the appropriate holding mode, leading to inefficiencies in object manipulation.

Method used

A learned model generation method incorporating a class inference model, a first holding mode inference model, and a second holding mode inference model to improve the estimation accuracy of holding modes by classifying and determining the optimal grip positions and forces based on captured images.

Benefits of technology

Enhances the precision of robot grasping by accurately estimating and executing the appropriate holding modes, improving the robot's ability to manipulate objects effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007717169000001
    Figure 0007717169000001
  • Figure 0007717169000002
    Figure 0007717169000002
  • Figure 0007717169000003
    Figure 0007717169000003
Patent Text Reader

Abstract

This trained model includes a class inference model, a first maintenance state inference model, and a second maintenance state inference model. The class inference model estimates a classification result, in which an object being maintained is classified into a prescribed maintenance division, on the basis of an image for estimation in which the object being maintained is photographed. The first maintenance state inference model estimates a first maintenance state of the object being maintained on the basis of the classification result and the image for estimation. The second maintenance state inference model estimates a second maintenance state of the object being maintained on the basis of the first maintenance state and the image for estimation. This trained model generation method includes generating a trained model by learning, as data for learning, an image for learning in which an object being learned that corresponds to the object being maintained is photographed.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This application claims the priority of Japanese Patent Application No. 2021-139421 (filed on August 27, 2021), and the entire disclosure of the application is incorporated herein by reference for that purpose.

Technical Field

[0002] The present disclosure relates to a learned model generation method, a learned model generation apparatus, a learned model, and a holding mode estimation apparatus.

Background Art

[0003] Conventionally, a method of appropriately holding an object with an arbitrary shape by a robot hand is known (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

[0005] In a learned model generation method according to an embodiment of the present disclosure, the learned model includes a class inference model, a first holding mode inference model, and a second holding mode inference model. The class inference model estimates a classification result obtained by classifying the object to be held into a predetermined holding category based on an estimation target image in which the object to be held by the robot is captured. The first holding mode inference model estimates a first holding mode of the object to be held based on the classification result and the estimation target image. The second holding mode inference model estimates a second holding mode of the object to be held based on the first holding mode and the estimation target image. The learned model generation method includes generating by learning a learning target image in which a learning target object corresponding to the object to be held is captured as learning data.

[0006] In a learned model generation device according to an embodiment of the present disclosure, the learned model includes a class inference model, a first holding mode inference model, and a second holding mode inference model. The class inference model estimates a classification result obtained by classifying a holding object into a predetermined holding category based on an estimation image in which the holding object held by the robot is captured. The first holding mode inference model estimates a first holding mode of the holding object based on the classification result and the estimation image. The second holding mode inference model estimates a second holding mode of the holding object based on the first holding mode and the estimation image. The learned model generation device generates the learned model by learning, as learning data, a learning image in which a learning object corresponding to the holding object is captured.

[0007] A learned model according to an embodiment of the present disclosure includes a class inference model, a first holding mode inference model, and a second holding mode inference model. The class inference model estimates a classification result obtained by classifying a holding object into a predetermined holding category based on an estimation image in which the holding object held by the robot is captured. The first holding mode inference model estimates a first holding mode of the holding object based on the classification result and the estimation image. The second holding mode inference model estimates a second holding mode of the holding object based on the first holding mode and the estimation image.

[0008] A holding mode estimation device according to an embodiment of the present disclosure estimates a mode in which a robot holds an object using a learned model. The learned model includes a class inference model, a first holding mode inference model, and a second holding mode inference model. The class inference model estimates a classification result obtained by classifying a holding object into a predetermined holding category based on an estimation image in which the holding object held by the robot is captured. The first holding mode inference model estimates a first holding mode of the holding object based on the classification result and the estimation image. The second holding mode inference model estimates a second holding mode of the holding object based on the first holding mode and the estimation image.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8A

Figure 8B

Figure 8C

Figure 9A

Figure 9B

Figure 9C

Figure 10

Figure 11

Embodiments for Carrying Out the Invention

[0010] In a learned model that inputs an image of an object and outputs the holding mode of the object, improvement in the estimation accuracy of the holding mode of the object is required. According to the learned model generation method, learned model generation device, learned model, and holding mode estimation device according to an embodiment of the present disclosure, the estimation accuracy of the holding mode of the object can be improved.

[0011] (Configuration example of robot control system 100) As shown in FIGS. 1 and 2, a robot control system 100 according to an embodiment of the present disclosure includes a robot 2, a robot control device 110, a learned model generation device 10, and a camera 4. The robot 2 holds a holding object 80 with an end effector 2B and executes work. The holding object 80 is also simply referred to as an object. The robot control device 110 controls the robot 2. The learned model generation device 10 generates a learned model used to determine the mode when the robot 2 holds the holding object 80. The mode in which the robot 2 holds the holding object 80 is also simply referred to as the holding mode. The holding mode includes the position where the robot 2 contacts the holding object 80 or the force applied by the robot 2 to the holding object 80, etc.

[0012] As shown in FIG. 2, in the present embodiment, the robot 2 holds the holding object 80 at the work start point 6. That is, the robot control device 110 controls the robot 2 to hold the holding object 80 at the work start point 6. The robot 2 may move the holding object 80 from the work start point 6 to the work target point 7. The holding object 80 is also referred to as a work target. The robot 2 operates inside the operation range 5.

[0013] <Robot 2> The robot 2 includes an arm 2A and an end effector 2B. The arm 2A may be configured as, for example, a 6-axis or 7-axis vertical articulated robot. The arm 2A may also be configured as a 3-axis or 4-axis horizontal articulated robot or a scalar robot. The arm 2A may be configured as a 2-axis or 3-axis orthogonal robot. The arm 2A may be configured as a parallel link robot or the like. The number of axes constituting the arm 2A is not limited to those exemplified. In other words, the robot 2 has an arm 2A connected by a plurality of joints and operates by driving the joints.

[0014] The end effector 2B may include, for example, a gripping hand configured to be able to grip the object 80 to be held. The gripping hand may have a plurality of fingers. The number of fingers of the gripping hand may be two or more. The fingers of the gripping hand may have one or more joints. The end effector 2B may also include a suction hand configured to be able to suck and hold the object 80 to be held. The end effector 2B may also include a scooping hand configured to be able to scoop and hold the object 80 to be held. The end effector 2B is also referred to as a holding part that holds the object 80 to be held. The end effector 2B is not limited to these examples and may be configured to be able to perform various other operations. In the configuration illustrated in FIG. 1, it is assumed that the end effector 2B includes a gripping hand.

[0015] The robot 2 can control the position of the end effector 2B by operating the arm 2A. The end effector 2B may have an axis that serves as a reference for the direction in which it acts on the object 80 to be held. When the end effector 2B has an axis, the robot 2 can control the direction of the axis of the end effector 2B by operating the arm 2A. The robot 2 controls the start and end of the operation in which the end effector 2B acts on the object 80 to be held. The robot 2 can move or process the object 80 to be held by controlling the operation of the end effector 2B while controlling the position of the end effector 2B or the direction of the axis of the end effector 2B. In the configuration illustrated in FIG. 1, the robot 2 causes the end effector 2B to hold the object 80 to be held at the work start point 6 and moves the end effector 2B to the work target point 7. The robot 2 releases the object 80 to be held from the end effector 2B at the work target point 7. By doing so, the robot 2 can move the object 80 to be held from the work start point 6 to the work target point 7.

[0016] <Sensor> The robot control system 100 further includes a sensor. The sensor detects physical information of the robot 2. The physical information of the robot 2 may include information regarding the actual position or orientation of each component of the robot 2, or the speed or acceleration of each component of the robot 2. The physical information of the robot 2 may include information regarding the force acting on each component of the robot 2. The physical information of the robot 2 may include information regarding the current flowing through the motor that drives each component of the robot 2 or the torque of the motor. The physical information of the robot represents the result of the actual operation of the robot 2. That is, the robot control system 100 can grasp the result of the actual operation of the robot 2 by acquiring the physical information of the robot 2.

[0017] The sensor may include a force sensor or a tactile sensor that detects, as physical information of the robot 2, a force, a distributed pressure, or a slip acting on the robot 2. The sensor may include a motion sensor that detects, as physical information of the robot 2, the position or orientation of the robot 2, or the speed or acceleration thereof. The sensor may include a current sensor that detects, as physical information of the robot 2, the current flowing through the motor that drives the robot 2. The sensor may include a torque sensor that detects, as physical information of the robot 2, the torque of the motor that drives the robot 2.

[0018] The sensor may be installed at a joint of the robot 2 or at a joint drive unit that drives the joint. The sensor may also be installed on the arm 2A or the end effector 2B of the robot 2.

[0019] The sensor outputs the detected physical information of the robot 2 to the robot control device 110. The sensor detects and outputs the physical information of the robot 2 at a predetermined timing. The sensor outputs the physical information of the robot 2 as time-series data.

[0020] <Camera 4> In the configuration example shown in FIG. 1, assume that the robot control system 100 includes a camera 4 attached to the end effector 2B of the robot 2. The camera 4 may, for example, photograph the object 80 to be held from the end effector 2B toward the object 80 to be held. That is, the camera 4 photographs the object 80 to be held from the direction in which the end effector 2B holds the object 80 to be held. The image obtained by photographing the object 80 to be held from the direction in which the end effector 2B holds the object 80 to be held is also referred to as a held object image (image for estimation). The image for estimation is an image in which the object 80 to be held for estimating the held object appears. Further, the camera 4 may be provided with a depth sensor and configured to be able to acquire depth data of the object 80 to be held. Note that the depth data is data regarding the distance in each direction within the angular range of the field of view of the depth sensor. More specifically, the depth data can also be said to be information regarding the distance from the camera 4 to the measurement point. The image captured by the camera 4 may include monochrome luminance information or may include luminance information of each color represented by RGB (Red, Green and Blue) or the like. The number of cameras 4 is not limited to one and may be two or more.

[0021] Note that the camera 4 is not limited to the configuration attached to the end effector 2B and may be provided at any position where the object 80 to be held can be photographed. In a configuration attached to a structure other than the end effector 2B, the above-described held object image may be synthesized based on the image captured by the camera 4 attached to the structure. The held object image may be synthesized by performing image conversion based on the relative position and relative orientation of the end effector 2B with respect to the attachment position and orientation of the camera 4. Alternatively, the held object image may be generated from CAD and drawing data.

[0022] <Trained model generation device 10> As shown in FIG. 1, the learned model generation device 10 includes a control unit 12 and an interface (I / F) 14. The interface 14 acquires information or data from an external device or outputs information or data to an external device. The interface 14 acquires learning data. The interface 14 acquires, as learning data, an image (learning image) obtained by photographing a holding object 80 (learning object) with the camera 4 and class information of the holding object 80. The learning image is an image in which the learning object appears. The control unit 12 learns based on the learning data and generates a learned model used to determine the holding mode of the holding object 80 in the robot control device 110. The control unit 12 outputs the generated learned model to the robot control device 110 via the interface 14.

[0023] The control unit 12 may be configured to include at least one processor to provide control and processing capabilities for executing various functions. The processor may execute a program for realizing various functions of the control unit 12. The processor may be realized as a single integrated circuit. The integrated circuit is also referred to as an IC (Integrated Circuit). The processor may be realized as a plurality of communicably connected integrated circuits and discrete circuits. The processor may be realized based on various other known technologies.

[0024] The control unit 12 may include a storage unit. The storage unit may include an electromagnetic storage medium such as a magnetic disk, or may include a memory such as a semiconductor memory or a magnetic memory. The storage unit stores various information. The storage unit stores programs and the like executed by the control unit 12. The storage unit may be configured as a non-temporary readable medium. The storage unit may function as a work memory of the control unit 12. At least a part of the storage unit may be configured separately from the control unit 12.

[0025] Interface 14 may be configured to include a communication device that can communicate either wired or wirelessly. The communication device may be configured to communicate by a communication method based on various communication standards. The communication device can be configured by known communication technologies.

[0026] Interface 14 may also be configured to include an input device that receives inputs such as information or data from a user. The input device may be configured to include, for example, a touch panel or touch sensor, or a pointing device such as a mouse. The input device may be configured to include physical keys. The input device may be configured to include a voice input device such as a microphone. Interface 14 may be configured to be connectable to an external input device. Interface 14 may be configured to obtain information input to the external input device from the external input device.

[0027] The interface 14 may be configured to include an output device that outputs information, data, etc. to the user. The output device may include, for example, a display device that displays information, data, etc. to the user. The display device may be configured to include, for example, an LCD (Liquid Crystal Display), an organic EL (Electro-Luminescence) display, an inorganic EL display, or a PDP (Plasma Display Panel), etc. The display device is not limited to these displays and may be configured to include various other types of displays. The display device may be configured to include a light-emitting device such as an LED (Light Emission Diode) or an LD (Laser Diode). The display device may be configured to include various other devices. The output device may include, for example, an audio output device such as a speaker that outputs auditory information such as sound. The output device may also include a vibration device that vibrates to give tactile information to the user. The output device is not limited to these examples and may include various other devices. The interface 14 may be configured to be connectable to an external output device. The interface 14 may output information to the external output device so that the external output device outputs the information to the user. The interface 14 may be configured to be connectable to the display device 20 described later as an external output device.

[0028] <Robot control device 110> The robot control device 110 determines a holding mode using the learned model generated by the learned model generation device 10, and controls the robot 2 so that the robot 2 holds the object 80 in the determined holding mode.

[0029] The robot control device 110 may be configured to include at least one processor to provide control and processing capabilities for executing various functions. Each component of the robot control device 110 may be configured to include at least one processor. A plurality of components among the components of the robot control device 110 may be realized by one processor. The entire robot control device 110 may be realized by one processor. The processor can execute a program that realizes various functions of the robot control device 110. The processor may be configured identically or similarly to the processor used in the learned model generation device 10.

[0030] The robot control device 110 may include a storage unit. The storage unit may be configured identically or similarly to the storage unit used in the learned model generation device 10.

[0031] (Operation example of the learned model generation device 10) The control unit 12 of the learned model generation device 10 learns, as learning data, an image obtained by photographing the object to be held 80 or an image generated from CAD data of the object to be held 80 and the class information of the object to be held 80, and generates a learned model for estimating the holding mode of the object to be held 80. The learning data may include so-called teacher data used in supervised learning. The learning data may include data generated by the device itself that executes learning and used in so-called unsupervised learning. An image obtained by photographing the object to be held 80 or an image generated as an image of the object to be held 80 is collectively referred to as an object image. As shown in FIG. 3, the model includes a class inference model 40, a first holding mode inference model 50, and a second holding mode inference model 60. The control unit 12 updates the model being learned by learning using the object image as learning data, and generates a learned model.

[0032] The class inference model 40 receives an object image as input. The class inference model 40 estimates the holding category to which the object to be held 80 included in the input object image should be classified. That is, the class inference model 40 classifies the object into any of a plurality of possible holding categories. The plurality of holding categories that can be estimated by the class inference model 40 are categories that can be output by the class inference model 40. The holding category is a category representing a difference in the shape of the object to be held 80 and is also referred to as a class.

[0033] The class inference model 40 classifies the object to be held 80 included in the input object image to be held into a predetermined class. The class inference model 40 outputs a classification result of classifying the object to be held 80 into a predetermined class based on the object image. In other words, the class inference model 40 estimates the class to which the object to be held 80 belongs when classified into classes based on the input object image.

[0034] The class inference model 40 outputs the estimated result of the class as class information. The classes that can be estimated by the class inference model 40 may be determined based on the shape of the object to be held 80. The classes that can be estimated by the class inference model 40 may be determined based on various characteristics such as the surface state or hardness of the object to be held 80. In the present embodiment, it is assumed that the number of classes that can be estimated by the class inference model 40 is four. The four classes are respectively referred to as the first class, the second class, the third class, and the fourth class. The number of classes may be three or less or five or more.

[0035] The first holding mode inference model fifty receives as input an object image and class information output from the class inference model 40. The first holding mode inference model 50 estimates the holding mode (first holding mode) of the object to be held 80 based on the input object image and class information, and outputs the estimated result of the holding mode.

[0036] The second holding mode inference model 60 receives, as inputs, the object image and the estimation result output from the first holding mode inference model 50. The second holding mode inference model 60 estimates the holding mode (second holding mode) of the object to be held 80 based on the input object image and the estimation result of the first holding mode inference model 50, and outputs the estimation result of the holding mode.

[0037] The class inference model 40, the first holding mode inference model 50, and the second holding mode inference model 60 may each be configured as, for example, a CNN (Convolution Neural Network) having a plurality of layers. The layers of the class inference model 40 are represented as processing layers 42. The layers of the first holding mode inference model 50 are represented as processing layers 52. The layers of the second holding mode inference model 60 are represented as processing layers 62. For the information input to the class inference model 40, the first holding mode inference model 50, and the second holding mode inference model 60, convolution processing based on predetermined weighting coefficients is executed in each layer of the CNN. In the learning of the class inference model 40, the first holding mode inference model 50, and the second holding mode inference model 60, the weighting coefficients are updated. The class inference model 40, the first holding mode inference model 50, and the second holding mode inference model 60 may be configured by VGG16 or ResNet50. The class inference model 40, the first holding mode inference model 50, and the second holding mode inference model 60 are not limited to these examples and may be configured as various other models.

[0038] The first holding mode inference model 50 includes a multiplier 54 that receives an input of class information. The multiplier 54 multiplies the output of the previous processing layer 52 by a weighting coefficient based on the class information and outputs the result to the subsequent processing layer 52.

[0039] As shown in FIG. 4, the first holding mode inference model 50 may include a first Non-G model 50N and a first Grasp model 50G. The first Non-G model 50N includes a processing layer 52N and a multiplier 54N. The first Grasp model 50G includes a processing layer 52G and a multiplier 54G. The second holding mode inference model 60 may include a second Non-G model 60N and a second Grasp model 60G. The second Non-G model 60N includes a processing layer 62N. The second Grasp model 60G includes a processing layer 62G.

[0040] The first Grasp model 50G and the second Grasp model 60G output the estimation result of the holding mode in the Grasp output mode. The Grasp output corresponds to the outputs of the first holding mode inference model 50 and the second holding mode inference model 60. Assume that the Grasp output is information representing the estimation result of the holding mode of the object 80 to be held. The Grasp output may be, for example, information representing the position to be contacted when holding the object 80 to be held. The Grasp output may be represented by, for example, an image.

[0041] The first Non-G model 50N and the second Non-G model 60N output the estimation result of the holding mode in the Non-G output mode. Assume that the Non-G output is information explaining the reason why the estimation result of the holding position 82 (see FIGS. 8A, 8B, 8C, etc.) of the object 80 to be held is derived. The Non-G output may be, for example, information representing the position that should not be contacted when holding the object 80 to be held. The Non-G output may be represented by, for example, an image. The first Non-G model 50N or the second Non-G model 60N is also referred to as a third holding mode inference model.

[0042] As described above, in the first holding mode inference model 50, the multiplier 54 multiplies the output of the processing layer 52 by the weighting coefficient based on the class information. Also, in the first Non-G model 50N and the first Grasp model 50G, the multipliers 54N and 54G multiply the outputs of the processing layers 52N and 52G by the weighting coefficient based on the class information. When the number of classes is four, the first holding mode inference model 50, or the first Non-G model 50N or the first Grasp model 50G branches the input object image to the processing layers 52 or 52N or 52G corresponding to each of the four classes, as shown in FIG. 5. The multiplier 54 or 54N or 54G includes a first multiplier 541, a second multiplier 542, a third multiplier 543, and a fourth multiplier 544 corresponding to each of the four classes. The first multiplier 541, the second multiplier 542, the third multiplier 543, and the fourth multiplier 544 multiply the outputs of the processing layers 52 or 52N or 52G corresponding to each of the four classes by the weighting coefficient. The first holding mode inference model 50, the first Non-G model 50N, and the first Grasp model 50G include an adder 56 that adds the outputs of the first multiplier 541, the second multiplier 542, the third multiplier 543, and the fourth multiplier 544, respectively. The output of the adder 56 is input to the processing layer 52 or 52N or 52G. The first holding mode inference model 50, the first Non-G model 50N, and the first Grasp model 50G output the output of the processing layer 52 or 52N or 52G located after the adder 56 as the estimation result of the holding mode.

[0043] The weighting coefficients by which each of the first multiplier 541, the second multiplier 542, the third multiplier 543, and the fourth multiplier 544 multiplies the output of the processing layer 52 or 52N or 52G are determined based on the class information input to the multiplier 54 or 54N or 54G. The class information represents which class among the four classes the object 80 to be held is classified into. When the object 80 to be held is classified into the first class, it is assumed that the weighting coefficient of the first multiplier 541 becomes 1, and the weighting coefficients of the second multiplier 542, the third multiplier 543, and the fourth multiplier 544 become 0. In this case, the output of the processing layer 52 or 52N or 52G corresponding to the first class is output from the adder 56.

[0044] The class information output by the class inference model 40 may be represented as the probabilities that the object to be held 80 is classified into each of the four classes. For example, the probabilities that the object to be held 80 is classified into the first class, the second class, the third class, and the fourth class may be represented by X1, X2, X3, and X4, respectively. In this case, the adder 56 multiplies the output of the processing layer 52 or 52N or 52G corresponding to the first class by X1, multiplies the output of the processing layer 52 or 52N or 52G corresponding to the second class by X2, multiplies the output of the processing layer 52 or 52N or 52G corresponding to the third class by X3, and multiplies the output of the processing layer 52 or 52N or 52G corresponding to the fourth class by X4, and then adds and outputs the results.

[0045] The control unit 12 generates the class inference model 40, the first holding mode inference model 50, and the second holding mode inference model 60 as learned models by learning the object image and the class information of the object to be held 80 as learning data. The class information of the object to be held 80 is given to the class inference model 40 as learning data.

[0046] The control unit 12 requires correct data in order to generate a learned model. For example, the control unit 12 requires learning data associated with information indicating which class the image used as learning data is correctly classified into in order to generate the class inference model 40. Also, the control unit 12 requires learning data associated with information indicating the correct position as the gripping position of the object shown in the image used as learning data in order to generate the first holding mode inference model 50 and the second holding mode inference model 60.

[0047] In the robot control system 1 according to the present embodiment, the learned model used may include a model that can learn the class or gripping position of the object to be held 80 in the object image during learning, or may include a model that cannot learn the class or gripping position of the object to be held 80. Even when the learned model cannot learn the class or gripping position of the object to be held 80, the control unit 12 may estimate the class and gripping position of the object to be held 80 by using a learned model learned based on the class and gripping position of a learning object corresponding to the object to be held 80, such as a learning object having a shape or dimensions similar to those of the object to be held 80. Therefore, in the robot control system 1 according to the present embodiment, the control unit 12 may or may not learn the object to be held 80 itself in order to generate a learned model.

[0048] <Estimation of Holding Mode by Learned Model> The robot control device 110 determines the holding mode of the object to be held 80 based on a learned model. The learned model is configured to receive, as an input, an image of the object to be held 80 and output an estimation result of the holding mode of the object to be held 80. The robot control device 110 acquires an image of the object to be held 80 taken by the camera 4, inputs it to the learned model, and acquires an estimation result of the holding mode of the object to be held 80 from the learned model. The robot control device 110 determines the holding mode based on the estimation result of the holding mode of the object to be held 80, and controls the robot 2 so that the object to be held 80 is held by the robot 2 in the determined holding mode.

[0049] As shown in FIG. 6, the robot control device 110 configures the learned model such that the class information from the class inference model 40 is input to the first holding mode inference model 50. Further, the robot control device 110 configures the learned model such that the output of the first holding mode inference model 50 is input to the second holding mode inference model 60.

[0050] The robot control device 110 inputs the object image into the class inference model 40 and causes the class inference model 40 to output class information. The robot control device 110 inputs the object image into the first holding mode inference model 50, and inputs the class information from the class inference model 40 into the first holding mode inference model 50, and causes the first holding mode inference model 50 to output an estimation result of the holding mode of the object to be held 80. Further, the robot control device 110 inputs the object image into the second holding mode inference model 60, and inputs the estimation result of the holding mode by the first holding mode inference model 50, and causes the second holding mode inference model 60 to output an estimation result of the holding mode of the object to be held 80. The robot control device 110 determines the holding mode of the object to be held 80 based on the estimation result of the holding mode by the second holding mode inference model 60.

[0051] The class information output from the class inference model 40 represents which class among the first to fourth classes the object to be held 80 is classified into based on the input object image. Specifically, it is assumed that the class inference model 40 is configured to output "1000" as class information when it estimates that the holding category (class) into which the object to be held 80 should be classified is the first class. It is assumed that the class inference model 40 is configured to output "0100" as class information when it estimates that the holding category (class) into which the object to be held 80 should be classified is the second class. It is assumed that the class inference model 40 is configured to output "0010" as class information when it estimates that the holding category (class) into which the object to be held 80 should be classified is the third class. It is assumed that the class inference model 40 is configured to output "0001" as class information when it estimates that the holding category (class) into which the object to be held 80 should be classified is the fourth class.

[0052] As shown in FIG. 7, the robot control device 110 may directly input the class information into the first holding mode inference model 50 by applying only the first holding mode inference model 50 and the second holding mode inference model 60 as the learned models.

[0053] In this embodiment, it is assumed that a shape image 81 representing the shape of the object 80 to be held, exemplified in FIGS. 8A, 8B, and 8C, is input to the learned model as the object image. In this case, it is assumed that the shape image 81 is classified into each class based on the shape of the object 80 to be held. When the shape of the object 80 to be held is an O shape, it is assumed that the shape image 81 is classified into the first class. When the shape of the object 80 to be held is an I shape, it is assumed that the shape image 81 is classified into the second class. When the shape of the object 80 to be held is a J shape, it is assumed that the shape image 81 is classified into the third class. It can be said that the J shape is a shape combining the I shape and the O shape. When the shape of the object 80 to be held is another shape, it is assumed that the shape image 81 is classified into the fourth class.

[0054] When the class inference model 40 estimates that the shape image 81 is classified into the first class, it outputs "1000" as class information. When the first holding mode inference model 50 acquires "1000" as class information, it estimates the holding position 82 inside the O-shaped object 80 in the shape image 81. When the class inference model 40 estimates that the shape image 81 is classified into the second class, it outputs "0100" as class information. When the first holding mode inference model 50 acquires "0100" as class information, it estimates the positions on both sides near the center of the I shape in the shape image 81 as the holding position 82. When the class inference model 40 estimates that the shape image 81 is classified into the third class, it outputs "0010" as class information. When the first holding mode inference model 50 acquires "0010" as class information, it estimates the positions on both sides near the ends of the J shape in the shape image 81 as the holding position 82. In other words, the first holding mode inference model 50 estimates the positions on both sides near the ends far from the O shape in the I shape as the holding position 82 among the shapes combining the I shape and the O shape.

[0055] An example in which the robot control device 110 estimates the holding position 82 when the object to be held 80 is a screw will be described. As shown in FIGS. 9A, 9B, and 9C, the robot control device 110 inputs an image of a screw as the object image into the learned model with the screw as the object to be held 80. The screw has a screw head 83 and a screw shaft 84. When the learned model tentatively classifies the shape image 81 of the screw into the first class, it cannot specify the position corresponding to the inside of the screw and cannot estimate the holding position 82. Therefore, the holding position 82 is not displayed in the shape image 81 classified into the first class. When the learned model tentatively classifies the shape image 81 of the screw into the second class, it estimates both sides near the center of the screw shaft 84 as the holding position 82. When the learned model tentatively classifies the shape image 81 of the screw into the third class, it estimates the positions on both sides near the end on the side far from the screw head 83 on the screw shaft 84 as the holding position 82.

[0056] <Parentheses> As described above, the learned model generation device 10 according to the present embodiment can generate a learned model including a class inference model 40, a first holding mode inference model 50, and a second holding mode inference model 60. The first holding mode inference model 50 estimates the holding mode based on the class information. The second holding mode inference model 60 estimates the holding mode without relying on the class information. By doing so, the robot control device 110 can complement the estimation result based on the class information and the estimation result not based on the class information to determine the holding mode. As a result, the estimation accuracy of the holding mode of the object to be held 80 can be improved.

[0057] The learned model generation device 10 may execute a learned model generation method including a procedure of generating a learned model including a class inference model 40, a first holding mode inference model 50, and a second holding mode inference model 60 by learning.

[0058] The robot control device 110 may estimate the holding mode using the learned model. The robot control device 110 is also referred to as a holding mode estimation device.

[0059] (Other Embodiments) The following describes other embodiments.

[0060] <Examples of Other Holding Modes>[ In the embodiments described above, a configuration has been described in which the learned model estimates the holding position 82 as a holding mode. The learned model can estimate not only the holding position 82 but also other modes as holding modes.

[0061] The learned model may estimate, as a holding mode, for example, the force applied by the robot 2 to hold the object 80. In this case, the first holding mode inference model 50 and the second holding mode inference model 60 estimate and output the force applied to hold the object 80.

[0062] Further, the learned model may estimate, as a holding mode, for example, the type of hand by which the robot 2 holds the object 80. In this case, the first holding mode inference model 50 and the second holding mode inference model 60 estimate and output the type of hand used to hold the object 80.

[0063] <Examples of Non-G Output and Grasp Output>[ Assume that the shape image 81 illustrated in FIG. 10 is input to the learned model as an object image. In this case, the learned model outputs information shown as an image in the table of FIG. 11, for example. In the table of FIG. 11, the information output by the first holding mode inference model 50 is described in the row marked "First Model". The information output by the second holding mode inference model 60 is described in the row marked "Second Model". Also, the information of the Non-G output is described in the column marked "Non-G". The information of the Grasp output is described in the column marked "Grasp".

[0064] The information described in the cell located in the row of "First Model" and the column of "Non-G" corresponds to the Non-G output output by the first Non-G model 50N included in the first holding mode inference model 50. The information described in the cell located in the row of "Second Model" and the column of "Non-G" corresponds to the Non-G output output by the second Non-G model 60N included in the second holding mode inference model 60. It is assumed that the Non-G output indicates that the darker the color, the less it should be contacted when holding, and the lighter the color, the more it can be contacted when holding. The meaning of the color of the Non-G output is not limited to this. The Non-G output is not limited to an image and may be represented as numerical information.

[0065] The information described in the cell located in the row of "First Model" and the column of "Grasp" corresponds to the Grasp output output by the first Grasp model 50G included in the first holding mode inference model 50. The information described in the cell located in the row of "Second Model" and the column of "Grasp" corresponds to the Grasp output output by the second Grasp model 60G included in the second holding mode inference model 60. In FIG. 11, the shape of the object 80 to be held in the Grasp output is displayed by a virtual line of a two-dot chain line. Also, the holding position 82 is displayed by a solid circle. It is assumed that the Grasp output indicates that the darker the color, the more suitable it is as the holding position 82, and the lighter the color, the less suitable it is as the holding position 82. The meaning of the color of the Grasp output is not limited to this. The Grasp output is not limited to an image and may be represented as numerical information.

[0066] As described above, the embodiments of the learned model generation device 10 and the robot control system 100 have been described. As embodiments of the present disclosure, in addition to a method or program for implementing a system or device, an embodiment as a storage medium (for example, an optical disk, a magneto-optical disk, a CD-ROM, a CD-R, a CD-RW, a magnetic tape, a hard disk, or a memory card, etc.) on which the program is recorded is also possible.

[0067] In addition, the implementation form of the program is not limited to application programs such as object code compiled by a compiler and program code executed by an interpreter, and may be in the form of program modules incorporated into an operating system. Furthermore, the program may or may not be configured such that all processing is performed only on the CPU on the control board. The program may be configured such that part or all of it is performed by another processing unit implemented on an expansion board or expansion unit added to the board as needed.

[0068] Although the embodiments according to the present disclosure have been described based on the various drawings and examples, it should be noted that those skilled in the art can make various modifications or alterations based on the present disclosure. Therefore, it should be noted that these modifications or alterations are included in the scope of the present disclosure. For example, the functions etc. included in each component etc. can be rearranged so as not to be logically contradictory, and it is possible to combine a plurality of components etc. into one or divide them.

[0069] Furthermore, all of the constituent elements described in the present disclosure, and / or all of the disclosed methods, or all of the steps of the processes, can be combined in any combination except combinations where these features are mutually exclusive. Also, each of the features described in the present disclosure can be replaced with an alternative feature that serves for the same purpose, an equivalent purpose, or a similar purpose, unless explicitly negated. Therefore, unless explicitly negated, each of the disclosed features is merely an example of a comprehensive series of identical or equivalent features.

[0070] Moreover, the embodiments according to the present disclosure are not limited to any specific configurations of the above-described embodiments. The embodiments according to the present disclosure can be extended to all of the novel features described in the present disclosure, or combinations thereof, or all of the novel methods, or steps of the processes, or combinations thereof.

[0071] In the present disclosure, descriptions such as "first" and "second" are identifiers for distinguishing the relevant configurations. The configurations distinguished by the descriptions such as "first" and "second" in the present disclosure can have their numbers in the configuration exchanged. For example, the first class can have the identifiers "first" and "second" exchanged with the second class. The exchange of identifiers is performed simultaneously. Even after the exchange of identifiers, the configurations are still distinguishable. The identifiers may be deleted. The configurations with the identifiers deleted are distinguished by symbols. Based only on the descriptions of the identifiers such as "first" and "second" in the present disclosure, the order of the configurations should not be interpreted, nor should it be used as the basis for the existence of an identifier with a smaller number.

Explanation of symbols

[0072] 10 Trained model generation device (12: control unit, 14: interface) 20 Display device 40 Class inference model (42: processing layer) 50 First holding mode inference model (52: processing layer, 54: multiplier (541 to 544: first to fourth multipliers), 56: adder) 50N First Non-G model (52N: processing layer, 54N: multiplier) 50G First Grasp model (52G: processing layer, 54G: multiplier) 60 Second holding mode inference model (62: processing layer) 60N Second Non-G model (62N: processing layer) 60G Second Grasp model (62G: processing layer) 80 Object to be held (83: screw head, 84: screw shaft) 81 Shape image 82 Holding position 100 Robot control system (2: robot, 2A: arm, 2B: end effector, 3: sensor, 4: camera, 5: influence range of robot, 6: work start point, 7: work target point, 110: robot control device)

Claims

1. A method for generating a learned model, comprising: a processor generates a learned model including a class inference model for inferring a classification result obtained by classifying a target object to be held, which appears in an estimation image in which the target object held by a robot appears, into a predetermined holding category, a first holding mode inference model for inferring a first holding mode of the target object based on the classification result and the estimation image, and a second holding mode inference model for inferring a second holding mode of the target object based on the first holding mode and the estimation image, by learning using a learning image in which a learning target object corresponding to the target object appears as learning data.

2. The method for generating a learned model according to claim 1, wherein the learning data further includes information regarding a class to which the target object is classified.

3. The learned model has a third holding mode inference model for inferring a mode in which the robot should not hold the target object based on the classification result input from the class inference model and an image of the target object, The method for generating a learned model according to claim 1 or 2, wherein the second holding mode inference model further infers and outputs a mode in which the robot holds the target object based on the inference result of the third holding mode inference model.

4. The first holding mode inference model outputs a position to be contacted when holding the target object, The method for generating a learned model according to claim 3, wherein the third holding mode inference model outputs a position that should not be contacted when holding the target object.

5. A learned model generation device comprising a processor that generates a learned model including a class inference model for inferring a classification result obtained by classifying a target object to be held, which appears in an estimation image in which the target object held by a robot appears, into a predetermined holding category, a first holding mode inference model for inferring a first holding mode of the target object based on the classification result and the estimation image, and a second holding mode inference model for inferring a second holding mode of the target object based on the first holding mode and the estimation image, by learning using a learning image in which a learning target object corresponding to the target object appears as learning data.

6. A learned model for inferring a holding mode of a target object held by a robot, A learned model that causes a computer to function so as to estimate a classification result obtained by classifying a holding target into a predetermined holding category based on an estimated image in which the holding target appears, estimate a first holding mode of the holding target based on the classification result and the estimated image, and estimate a second holding mode of the holding target based on the first holding mode and the estimated image.

7. A processor that estimates a mode in which the robot holds an object using a learned model including a class inference model that estimates a classification result obtained by classifying the object to be held by the robot into a predetermined holding category based on an estimated image in which the object appears, a first holding mode inference model that estimates a first holding mode of the object to be held based on the classification result and the estimated image, and a second holding mode inference model that estimates a second holding mode of the object to be held based on the first holding mode and the estimated image.

Citation Information

Patent Citations

  • Method for gripping object of optional shape by robot

    JP2005169564A

  • Information processing device, information processing method, and program

    JP2020021212A

  • Object detection device, object gripping system, object detection method, and object detection program

    JP2020197978A

  • Object handling device, and control method

    JP2021003805A

  • Learning data generation method

    JP2021070122A