Robot holding mode determination device, holding mode determination method, and robot control system

The robot holding mode determination device allows users to select and adjust holding modes through a display interface, enhancing user control and convenience in robot object handling by refining trained models.

JP2025147230AInactive Publication Date: 2025-10-06KYOCERA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025132686
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-08-19
Filing Date
2025-08-07
Publication Date
2025-10-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing robot systems lack user control over the output of trained models determining holding manners of objects, necessitating improved convenience in object handling.

Method used

A robot holding mode determination device and method that includes a control unit and interface, allowing users to select and adjust holding behaviors through a display device, using a trained model to estimate and refine holding modes based on object classification.

Benefits of technology

Enhances user control over robot object handling by enabling selection and adjustment of holding modes, improving convenience and accuracy in handling various objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025147230000001_ABST
    Figure 2025147230000001_ABST
Patent Text Reader

Abstract

To provide a robot holding mode determination device, a holding mode determination method, and a robot control system capable of improving convenience of a trained model that outputs an object holding mode.SOLUTION: A robot holding mode determination device includes a control part and an interface. The control part outputs at least one holding mode of a target object of a robot estimated by classifying the target object into at least one of a plurality of holding sections that may be estimated, to a display device together with an image of the target object. The interface acquires selection of a user with respect to the at least one holding mode via a display device.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This application claims priority to Japanese Patent Application No. 2021-134391 (filed August 19, 2021), the entire disclosure of which is incorporated herein by reference. [Technical Field]

[0002] The present disclosure relates to a holding mode determination device, a holding mode determination method, and a robot control system. [Background technology]

[0003] Conventionally, a method for appropriately holding an object of any shape by a robot hand is known (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2005-169564 Summary of the Invention

[0005] A holding behavior determination device for a robot according to an embodiment of the present disclosure includes a control unit and an interface. The control unit outputs at least one holding behavior of the robot for an object, estimated by classifying the object into at least one of a plurality of holding categories that can be estimated, to a display device together with an image of the object. The interface acquires a user's selection for the at least one holding behavior via the display device.

[0006] A method for determining a holding mode of a robot according to an embodiment of the present disclosure includes outputting, to a display device, at least one holding mode of an object of the robot estimated by classifying the object into at least one of a plurality of holding categories that can be estimated, together with an image of the object, and acquiring, via the display device, a user's selection of the at least one holding mode.

[0007] A robot control system according to an embodiment of the present disclosure includes a robot, a robot control device that controls the robot, a holding behavior determination device that determines a behavior of the robot holding an object, and a display device. The holding behavior determination device outputs at least one holding behavior of the object of the robot estimated by classifying the object into at least one of a plurality of possible holding categories to the display device, together with an image of the object. The display device displays the image of the object and the at least one holding behavior of the object. The display device accepts input of a user's selection for the at least one holding behavior and outputs the selection to the holding behavior determination device. The holding behavior determination device acquires the user's selection for the at least one holding behavior from the display device. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a block diagram illustrating an example of the configuration of a robot control system according to an embodiment. [Figure 2] FIG. 1 is a schematic diagram illustrating an example of the configuration of a robot control system according to an embodiment. [Figure 3] FIG. 10 is a diagram illustrating an example of a model during training. [Figure 4] FIG. 10 is a diagram illustrating an example of a retention mode inference model. [Figure 5] FIG. 10 is a diagram illustrating an example configuration for determining a retention mode using a trained model. [Figure 6] FIG. 10 is a diagram illustrating an example of a configuration for determining a holding mode by a holding mode inference model. [Figure 7A]FIG. 10 is a diagram showing an example of a gripping position when an object is classified into a first class. [Figure 7B] FIG. 10 is a diagram showing an example of a gripping position when the object is classified into a second class. [Figure 7C] FIG. 10 is a diagram showing an example of a gripping position when the object is classified into a third class. [Figure 8A] FIG. 10 is a diagram showing an example of a grip position determined when the object is classified into a first class. [Figure 8B] FIG. 10 is a diagram showing an example of a grip position determined when the object is classified into a second class. [Figure 8C] FIG. 10 is a diagram showing an example of a grip position determined when the target object is classified into a third class. [Figure 9] 10 is a flowchart illustrating an example of a procedure for a holding mode determination method according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] In a trained model that inputs an image of an object and outputs the holding manner of the object, a user cannot control the output of the trained model and therefore cannot control the holding manner of the object. There is a need to improve the convenience of trained models that output holding manners of objects. According to a robot holding manner determination device, a holding manner determination method, and a robot control system according to an embodiment of the present disclosure, the convenience of trained models that output holding manners of objects can be improved.

[0010] (Configuration example of robot control system 100) As shown in FIG. 1 , a robot control system 100 according to an embodiment of the present disclosure includes a robot 2, a robot control device 110, a holding mode determination device 10, a display device 20, and a camera 4. The robot 2 performs a task by holding a holding object 80 with an end effector 2B. The holding object 80 is also simply referred to as an object. The robot control device 110 controls the robot 2. The holding mode determination device 10 determines a mode in which the robot 2 holds the holding object 80 and outputs the determined mode to the robot control device 110. The mode in which the robot 2 holds the holding object 80 is also simply referred to as a holding mode. The holding mode includes the position at which the robot 2 contacts the holding object 80, the force the robot 2 applies to the holding object 80, etc.

[0011] As shown in FIG. 2 , in this embodiment, the robot 2 holds the holding object 80 at the work start point 6. That is, the robot control device 110 controls the robot 2 so that the holding object 80 is held at the work start point 6. The robot 2 may move the holding object 80 from the work start point 6 to the work target point 7. The holding object 80 is also referred to as the work object. The robot 2 operates within the operating range 5.

[0012] <Robot 2> The robot 2 includes an arm 2A and an end effector 2B. The arm 2A may be configured as, for example, a six- or seven-axis vertical articulated robot. The arm 2A may be configured as a three- or four-axis horizontal articulated robot or a SCARA robot. The arm 2A may be configured as a two- or three-axis Cartesian robot. The arm 2A may be configured as a parallel link robot or the like. The number of axes configuring the arm 2A is not limited to those illustrated. In other words, the robot 2 has an arm 2A connected by multiple joints, and operates by driving the joints.

[0013] The end effector 2B may include, for example, a gripping hand configured to grip the object to be held 80. The gripping hand may have multiple fingers. The number of fingers on the gripping hand may be two or more. The fingers on the gripping hand may have one or more joints. The end effector 2B may include a suction hand configured to suction and hold the object to be held 80. The end effector 2B may include a scooping hand configured to scoop and hold the object to be held 80. The end effector 2B is also referred to as a holding unit that holds the object to be held 80. The end effector 2B is not limited to these examples and may be configured to perform various other operations. In the configuration illustrated in FIG. 1, the end effector 2B includes a gripping hand.

[0014] The robot 2 can control the position of the end effector 2B by operating the arm 2A. The end effector 2B may have an axis that serves as a reference for the direction in which it acts on the held object 80. If the end effector 2B has an axis, the robot 2 can control the direction of the axis of the end effector 2B by operating the arm 2A. The robot 2 controls the start and end of the operation of the end effector 2B acting on the held object 80. The robot 2 can move or process the held object 80 by controlling the operation of the end effector 2B while controlling the position of the end effector 2B or the direction of the axis of the end effector 2B. In the configuration illustrated in FIG. 1 , the robot 2 has the end effector 2B hold the held object 80 at the work start point 6 and moves the end effector 2B to the work target point 7. The robot 2 has the end effector 2B release the held object 80 at the work target point 7. In this way, the robot 2 can move the object to be held 80 from the work start point 6 to the work target point 7.

[0015] <Sensor> The robot control system 100 further includes a sensor. The sensor detects physical information of the robot 2. The physical information of the robot 2 may include information about the actual position or posture of each component of the robot 2, or information about the speed or acceleration of each component of the robot 2. The physical information of the robot 2 may include information about the force acting on each component of the robot 2. The physical information of the robot 2 may include information about the current flowing through the motor that drives each component of the robot 2, or information about the torque of the motor. The physical information of the robot 2 represents the results of the actual operation of the robot 2. In other words, the robot control system 100 can grasp the results of the actual operation of the robot 2 by acquiring the physical information of the robot 2.

[0016] The sensors may include force sensors or tactile sensors that detect, as physical information of the robot 2, forces, distributed pressure, slippage, or the like acting on the robot 2. The sensors may include, as physical information of the robot 2, motion sensors that detect the position or posture, or the speed or acceleration of the robot 2. The sensors may include, as physical information of the robot 2, current sensors that detect the current flowing through the motor that drives the robot 2. The sensors may include, as physical information of the robot 2, torque sensors that detect the torque of the motor that drives the robot 2.

[0017] The sensor may be installed in a joint of the robot 2 or in a joint drive unit that drives the joint. The sensor may be installed in the arm 2A or the end effector 2B of the robot 2.

[0018] The sensor outputs the detected physical information of the robot 2 to the robot control device 110. The sensor detects and outputs the physical information of the robot 2 at a predetermined timing. The sensor outputs the physical information of the robot 2 as time-series data.

[0019] <Camera 4> In the exemplary configuration shown in FIG. 1, the robot control system 100 includes a camera 4 attached to the end effector 2B of the robot 2. The camera 4 captures an image of the holding object 80, for example, from the end effector 2B toward the holding object 80. That is, the camera 4 captures an image of the holding object 80 from the direction in which the end effector 2B holds the holding object 80. An image of the holding object 80 captured from the direction in which the end effector 2B holds the holding object 80 is also referred to as a holding object image. The camera 4 may also include a depth sensor and be configured to acquire depth data of the holding object 80. The depth data is data related to distances in different directions within the angle of view of the depth sensor. More specifically, the depth data can be considered information related to the distance from the camera 4 to a measurement point. The image captured by the camera 4 may include monochrome brightness information or brightness information for each color expressed by RGB (Red, Green, and Blue), etc. The number of cameras 4 is not limited to one, and may be two or more.

[0020] The camera 4 is not limited to being attached to the end effector 2B, but may be provided at any position where it can capture an image of the held object 80. In a configuration where the camera 4 is attached to a structure other than the end effector 2B, the above-mentioned image of the held object may be synthesized based on an image captured by the camera 4 attached to the structure. The image of the held object may be synthesized by image conversion based on the relative position and orientation of the end effector 2B with respect to the attachment position and orientation of the camera 4. Alternatively, the image of the held object may be generated from CAD and drawing data.

[0021] <Holding mode determination device 10> 1, the holding mode determination device 10 includes a control unit 12 and an interface (I / F) 14. The interface 14 acquires information or data relating to a holding object 80, etc. from an external device and outputs the information or data to the external device. The interface 14 also acquires an image of the holding object 80 from the camera 4. The interface 14 outputs the information or data to the display device 20, and causes the control unit of the display device 20, which will be described later, to display the information or data. The control unit 12 determines the holding mode of the holding object 8 by the robot 2 based on the information or data acquired by the interface 14, and outputs the determined information or data to the robot control device 110 via the interface 14.

[0022] The control unit 12 may include at least one processor to provide control and processing capabilities for executing various functions. The processor may execute programs that implement the various functions of the control unit 12. The processor may be implemented as a single integrated circuit. An integrated circuit is also called an IC (Integrated Circuit). The processor may be implemented as multiple integrated circuits and discrete circuits that are communicatively connected. The processor may also be implemented based on various other known technologies.

[0023] The control unit 12 may include a memory unit. The memory unit may include an electromagnetic storage medium such as a magnetic disk, or may include a memory such as a semiconductor memory or a magnetic memory. The memory unit stores various information. The memory unit stores programs to be executed by the control unit 12, etc. The memory unit may be configured as a non-transitory readable medium. The memory unit may function as a work memory for the control unit 12. At least a part of the memory unit may be configured as a separate entity from the control unit 12.

[0024] The interface 14 may include a communication device configured to be capable of wired or wireless communication. The communication device may be configured to be capable of communication using a communication method based on various communication standards. The communication device may be configured using known communication technology.

[0025] The interface 14 may be configured to include an input device that accepts input of information, data, etc. from a user. The input device may be configured to include, for example, a touch panel or touch sensor, or a pointing device such as a mouse. The input device may be configured to include physical keys. The input device may be configured to include an audio input device such as a microphone. The interface 14 may be configured to be connectable to an external input device. The interface 14 may be configured to acquire information input to the external input device from the external input device.

[0026] The interface 14 may be configured to include an output device that outputs information, data, etc. to the user. The output device may include, for example, an audio output device such as a speaker that outputs auditory information such as sound. The output device may include a vibration device that vibrates to provide tactile information to the user. The output device is not limited to these examples and may include various other devices. The interface 14 may be configured to be connectable to an external output device. The interface 14 may output information to an external output device so that the external output device outputs information to the user. The interface 14 may be configured to be connectable to a display device 20, described below, as an external output device.

[0027] <Display device 20> The display device 20 displays information, data, etc. to a user. The display device 20 may include a control unit that controls the display. The display device 20 may be configured to include, for example, an LCD (Liquid Crystal Display), an organic EL (Electro-Luminescence) display, an inorganic EL display, or a PDP (Plasma Display Panel). The display device 20 is not limited to these displays and may be configured to include displays of various other types. The display device 20 may be configured to include a light-emitting device such as an LED (Light Emission Diode) or an LD (Laser Diode). The display device 20 may be configured to include various other devices.

[0028] The display device 20 may also be configured to include an input device that accepts input of information, data, etc. from a user. The input device may be configured to include, for example, a touch panel or touch sensor, or a pointing device such as a mouse. The input device may be configured to include physical keys. The input device may be configured to include an audio input device such as a microphone. The input device may be connected to the interface 14, for example.

[0029] <Robot control device 110> The robot control device 110 acquires information specifying the holding mode from the holding mode determination device 10 and controls the robot 2 so that the robot 2 holds the holding target 80 in the holding mode determined by the holding mode determination device 10.

[0030] The robot control device 110 may be configured to include at least one processor to provide control and processing power for executing various functions. Each component of the robot control device 110 may be configured to include at least one processor. A plurality of components of the robot control device 110 may be implemented by a single processor. The entire robot control device 110 may be implemented by a single processor. The processor may execute programs that implement various functions of the robot control device 110. The processor may be configured the same as or similar to the processor used in the holding mode determination device 10.

[0031] The robot control device 110 may include a memory unit. The memory unit may be configured the same as or similar to the memory unit used in the holding mode determination device 10.

[0032] The robot control device 110 may include the holding mode determination device 10. The robot control device 110 and the holding mode determination device 10 may be configured as separate entities.

[0033] (Operation example of holding mode determination device 10) The control unit 12 of the holding mode determination device 10 determines the holding mode of the holding object 80 based on the trained model. The trained model is configured to receive an image of the holding object 80 as input, and output an estimation result of the holding mode of the holding object 80. The control unit 12 generates the trained model. The control unit 12 acquires an image of the holding object 80 from the camera 4, inputs the image to the trained model, and acquires an estimation result of the holding mode of the holding object 80 from the trained model. The control unit 12 determines the holding mode based on the estimation result of the holding mode of the holding object 80, and outputs it to the robot control device 110. The robot control device 110 controls the robot 2 so that the robot 2 holds the holding object 80 in the holding mode acquired from the holding mode determination device 10.

[0034] <Generating a trained model> The control unit 12 learns images of the held object 80 or images generated from CAD data of the held object 80, etc., as learning data, and generates a trained model that estimates the holding state of the held object 80. The learning data may include teacher data used in so-called supervised learning. The learning data may include data generated by the device performing the learning itself, used in so-called unsupervised learning. Images of the held object 80 or images generated as images of the held object 80 are collectively referred to as object images. As shown in FIG. 3, the inference model includes a class inference model 40 and a holding state inference model 50. The control unit 12 updates the model being trained by learning using the object images as learning data, and generates a trained model.

[0035] The class inference model 40 receives an object image as input. The class inference model 40 estimates a retention category into which the retained object 80 should be classified based on the input object image. In other words, the class inference model 40 classifies the object into one of a plurality of retention categories that can be estimated. The plurality of retention categories that can be estimated by the class inference model 40 are categories that can be output by the class inference model 40. A retention category is a category that represents differences in the shape of the retained object 80, and is also referred to as a class.

[0036] The class inference model 40 classifies the input target image into a predetermined class. The class inference model 40 outputs a classification result in which the retained object 80 is classified into a predetermined class based on the target image. In other words, the class inference model 40 estimates the class to which the retained object 80 belongs when the retained object 80 is classified into a class.

[0037] The class inference model 40 outputs the class estimation results as class information. The classes that can be estimated by the class inference model 40 may be determined based on the shape of the held object 80. The classes that can be estimated by the class inference model 40 may be determined based on various characteristics of the held object 80, such as the surface condition, material, or hardness. In this embodiment, the number of classes that can be estimated by the class inference model 40 is assumed to be four. The four classes are referred to as a first class, a second class, a third class, and a fourth class, respectively. The number of classes may be three or less, or may be five or more.

[0038] The holding mode inference model 50 receives as input the object image and the class information output from the class inference model 40. The holding mode inference model 50 estimates the holding mode of the held object 80 based on the input object image and class information, and outputs the estimated holding mode result.

[0039] The class inference model 40 and the retention behavior inference model 50 are each configured, for example, as a convolution neural network (CNN) having multiple layers. A layer of the class inference model 40 is represented as a processing layer 42. A layer of the retention behavior inference model 50 is represented as a processing layer 52. Convolution processing based on predetermined weighting coefficients is performed in each layer of the CNN on information input to the class inference model 40 and the retention behavior inference model 50. The weighting coefficients are updated during training of the class inference model 40 and the retention behavior inference model 50. The class inference model 40 and the retention behavior inference model 50 may be configured using VGG16 or ResNet50. The class inference model 40 and the retention behavior inference model 50 are not limited to these examples and may be configured as various other models.

[0040] The retention mode inference model 50 includes a multiplier 54 that receives input of class information. When the number of classes is four, the retention mode inference model 50 branches the retention objects 80 included in the input object image to processing layers 52 corresponding to each of the four classes, as shown in FIG. 4 . The multiplier 54 includes a first multiplier 541, a second multiplier 542, a third multiplier 543, and a fourth multiplier 544, each corresponding to one of the four classes. The first multiplier 541, the second multiplier 542, the third multiplier 543, and the fourth multiplier 544 multiply the output of the processing layer 52 corresponding to each of the four classes by a weighting coefficient. The retention mode inference model 50 includes an adder 56 that adds the outputs of the first multiplier 541, the second multiplier 542, the third multiplier 543, and the fourth multiplier 544. The output of the adder 56 is input to the processing layer 52. The retention mode inference model 50 outputs the output of the processing layer 52 located after the adder 56 as the retention mode estimation result.

[0041] The weighting coefficients by which the first multiplier 541, the second multiplier 542, the third multiplier 543, and the fourth multiplier 544 respectively multiply the output of the processing layer 52 are determined based on the class information input to the multiplier 54. The class information indicates which of the four classes the held object 80 has been classified into. If the held object 80 is classified into the first class, the weighting coefficient of the first multiplier 541 will be 1, and the weighting coefficients of the second multiplier 542, the third multiplier 543, and the fourth multiplier 544 will be 0. In this case, the output of the processing layer 52 corresponding to the first class is output from the adder 56.

[0042] The class information may be represented as the probability that the retention object 80 will be classified into each of the four classes. For example, the probabilities that the retention object 80 will be classified into the first class, second class, third class, and fourth class may be represented by X1, X2, X3, and X4, respectively. In this case, the adder 56 adds and outputs the output obtained by multiplying the output of the processing layer 52 corresponding to the first class by X1, the output obtained by multiplying the output of the processing layer 52 corresponding to the second class by X2, the output obtained by multiplying the output of the processing layer 52 corresponding to the third class by X3, and the output obtained by multiplying the output of the processing layer 52 corresponding to the fourth class by X4.

[0043] The control unit 12 generates a class inference model 40 and a holding manner inference model 50 as trained models by learning object images as training data. The control unit 12 requires correct answer data to generate the trained models. For example, to generate the class inference model 40, the control unit 12 requires training data associated with information indicating which class the image used as training data is correctly classified into. Furthermore, to generate the holding manner inference model 50, the control unit 12 requires training data associated with information indicating the correct gripping position of the object shown in the image used as training data.

[0044] The trained model used in the robot control system 1 according to this embodiment may include a model that has learned the class or grip position of the object in the object image during training, or may include a model that has not learned the class or grip position. Even if the trained model has not learned the class or grip position, the control unit 12 may estimate the class and grip position of the object by referring to the class and grip position of a trained object similar to the object. Therefore, in the robot control system 1 according to this embodiment, the control unit 12 may or may not train the object itself to generate the trained model.

[0045] <Estimation of retention patterns using a trained model> As shown in FIG. 5 , the control unit 12 configures the trained model so that class information from the class inference model 40 is input to the holding manner inference model 50 via the conversion unit 60. That is, the trained model includes the conversion unit 60 located between the class inference model 40 and the holding manner inference model 50. The control unit 12 inputs an object image to the class inference model 40 and causes the class inference model 40 to output class information. The control unit 12 inputs the object image to the holding manner inference model 50, and inputs class information from the class inference model 40 to the holding manner inference model 50 via the conversion unit 60, and causes the holding manner inference model 50 to output an estimation result of the holding manner of the held object 80.

[0046] The class information output from the class inference model 40 indicates which of the first to fourth classes the retained object 80 included in the input object image should be classified into. Specifically, the class inference model 40 is configured to output "1000" as the class information when it estimates that the retained category (class) into which the retained object 80 should be classified is the first class. The class inference model 40 is configured to output "0100" as the class information when it estimates that the retained category (class) into which the retained object 80 should be classified is the second class. The class inference model 40 is configured to output "0010" as the class information when it estimates that the retained category (class) into which the retained object 80 should be classified is the third class. The class inference model 40 is configured to output "0001" as the class information when it estimates that the retained category (class) into which the retained object 80 should be classified is the fourth class.

[0047] The conversion unit 60 converts the class information from the class inference model 40 and inputs it to the multiplier 54 of the retention mode inference model 50. For example, when the class information input from the class inference model 40 is "1000", the conversion unit 60 may be configured to convert it into another character string such as "0100" and output the converted class information to the retention mode inference model 50. The conversion unit 60 may also be configured to output the class information "1000" input from the class inference model 40 as is to the retention mode inference model 50. The conversion unit 60 may have a table that specifies rules for converting class information and may be configured to convert the class information based on the table. The conversion unit 60 may be configured to convert the class information using a matrix, a mathematical formula, or the like.

[0048] As shown in FIG. 6, the control unit 12 may apply only the retention pattern inference model 50 as a trained model and input class information directly to the retention pattern inference model 50.

[0049] In this embodiment, a shape image 81 representing the shape of a retention object 80 exemplified in FIGS. 7A, 7B, and 7C is input to the trained model as an object image. In this case, the retention object 80 included in the shape image 81 is classified into classes based on the shape of the retention object 80. If the shape of the retention object 80 is an O-shape, the retention object 80 included in the shape image 81 is classified into a first class. If the shape of the retention object 80 is an I-shape, the retention object 80 included in the shape image 81 is classified into a second class. If the shape of the retention object 80 is a J-shape, the retention object 80 included in the shape image 81 is classified into a third class. A J-shape can also be said to be a shape combining an I-shape and an O-shape. If the shape of the retention object 80 is another shape, the retention object 80 included in the shape image 81 is classified into a fourth class.

[0050] When the class inference model 40 estimates that the holding object 80 included in the shape image 81 is classified into the first class, it outputs "1000" as the class information. When the holding mode inference model 50 acquires "1000" as the class information, it estimates the holding position 82 inside the O-shaped holding object 80 in the shape image 81. When the class inference model 40 estimates that the holding object 80 included in the shape image 81 is classified into the second class, it outputs "0100" as the class information. When the holding mode inference model 50 acquires "0100" as the class information, it estimates positions on both sides near the center of the I-shape in the shape image 81 as the holding positions 82. When the class inference model 40 estimates that the holding object 80 included in the shape image 81 is classified into the third class, it outputs "0010" as the class information. The holding mode inference model 50 estimates the holding position 82 of the holding object 80 included in the shape image 81 according to the class into which the holding object 80 included in the shape image 81 is classified. For example, when the holding mode inference model 50 acquires "0010" as class information, it estimates positions on both sides near the end of the J-shape in the shape image 81 as holding positions 82. In other words, the holding mode inference model 50 estimates positions on both sides near the end of the I-shape farther from the O-shape in a shape that combines an I-shape and an O-shape as holding positions 82.

[0051] <Setting of conversion rules for the conversion unit 60> The control unit 12 can set rules for converting class information in the conversion unit 60. The control unit 12 receives input from the user via the interface 14 and sets rules based on the user input.

[0052] The control unit 12 inputs an object image into the trained model and causes the trained model to output an estimated holding state when the holding object 80 is classified into each of a plurality of classes. Specifically, as shown in FIGS. 8A, 8B, and 8C, the control unit 12 inputs an image of a screw as the holding object 80 as the object image into the trained model. The screw has a screw head 83 and a screw shaft 84. If the trained model classifies a screw as the holding object 80 included in a screw shape image 81 into the first class, the trained model cannot identify a position corresponding to the inside of the screw and cannot estimate the holding position 82. Therefore, the holding position 82 is not displayed in the shape image 81 including the screw classified into the first class. If the trained model classifies a screw as the holding object 80 included in a screw shape image 81 into the second class, the trained model estimates both sides near the center of the screw shaft 84 as the holding positions 82. If the trained model classifies the screw as the holding object 80 contained in the screw shape image 81 as the third class, it estimates the positions on both sides of the end of the screw shaft 84 farthest from the screw head 83 as holding positions 82.

[0053] The control unit 12 presents to the user the result of estimating the holding position 82 in the shape image 81 when the trained model classifies the holding object 80 included in the shape image 81 into each class. Specifically, the control unit 12 displays an image in which the holding position 82 is superimposed on the shape image 81, as illustrated in FIGS. 8A, 8B, and 8C, on the display device 20 as a candidate holding position 82 for holding the holding object 80, thereby allowing the user to visually recognize the image. The user can check the holding positions 82 (candidate holding positions 82) when the holding object 80 included in the shape image 81 is classified into each class and determine which holding position 82 is suitable as a position for holding a screw as the holding object 80. The control unit 12 allows the user to input information specifying which holding position 82 the user has determined to be suitable among the candidate holding positions 82 via the I / F 14. The control unit 12 acquires the class corresponding to the holding position 82 determined to be suitable by the user based on the user's input to the I / F 14.

[0054] When the control unit 12 causes the robot 2 to hold a screw as the holding object 80, the control unit 12 causes the holding pattern inference model 50 to estimate a holding position 82 of a class that the user has determined to be appropriate. Specifically, the control unit 12 can control the holding position 82 that the holding pattern inference model 50 estimates by controlling the class information input to the holding pattern inference model 50.

[0055] For example, when a holding object 80 included in an object image is classified into a class corresponding to a screw (e.g., the second class), the control unit 12 can cause the holding manner inference model 50 to estimate the holding position 82 as an inference result assuming that the holding object 80 has been classified into the third class. In this case, the control unit 12 inputs "0010" indicating the third class to the holding manner inference model 50 as class information to which the holding object 80 belongs. Even if the class information output by the class inference model 40 is originally "0100" indicating the second class, the control unit 12 inputs "0010" as class information to the holding manner inference model 50. In other words, although the holding manner of the holding object 80 classified into the second class would normally be inferred as a holding manner corresponding to the second class, the control unit 12 causes the holding manner inference model 50 to estimate the holding manner assuming that the holding object 80 has been classified into the third class. The control unit 12 sets a conversion rule in the conversion unit 60 so that when the class inference model 40 outputs "0100" as class information of the held object 80, "0010" is input to the holding pattern inference model 50 as class information.

[0056] The control unit 12 can cause the holding pattern inference model 50 to make an inference by converting the class into which the held object 80 is classified into another class and outputting the inference result. For example, even if the held object 80 is classified into the second class, the control unit 12 can cause the holding pattern inference model 50 to make an inference such that the holding position 82 of the held object 80 is classified into the third class and output as the inference result. Even if the held object 80 is classified into the third class, the control unit 12 can cause the holding pattern inference model 50 to make an inference such that the holding position 82 of the held object 80 is classified into the second class and output as the inference result. In this case, the control unit 12 sets a conversion rule in the conversion unit 60 so that when the class information from the class inference model 40 is "0100", it is converted to "0010" and input to the holding pattern inference model 50, and when the class information from the class inference model 40 is "0010", it is converted to "0100" and input to the holding pattern inference model 50.

[0057] As described above, by configuring the control unit 12 to be able to set the conversion rule, it is possible to reduce the need for the user to select a holding mode each time when processing multiple objects consecutively, for example. In this case, for example, once a holding mode is determined by the user, the conversion rule may be set so that the holding modes of the multiple objects become the determined holding mode.

[0058] Furthermore, the class inference model 40 may be set with a class that has no retention status as a retention classification. In this case, the retention status inference model 50 may infer that a certain class has no retention status when it is input. In this case, for example, when processing multiple objects continuously, the control unit 12 may determine that an object classified into a class other than the class expected by the user is a foreign object. Then, the control unit 12 can control the class of an object determined to be a foreign object to be converted into a class that has no retention status according to a set conversion rule, thereby preventing the foreign object from being retained.

[0059] <Example of procedure for determining retention mode> The control unit 12 of the holding mode determination device 10 may execute a holding mode determination method including the steps of the flowchart illustrated in Fig. 9. The holding mode determination method may be realized as a holding mode determination program executed by a processor constituting the control unit 12. The holding mode determination program may be stored in a non-transitory computer-readable medium.

[0060] The control unit 12 inputs an object image into the trained model (step S1).

[0061] The control unit 12 classifies the retained object 80 included in the object image into classes using the class inference model 40 (step S2). Specifically, the class inference model 40 estimates the class into which the retained object 80 is classified, and outputs the estimation result as class information.

[0062] The control unit 12 estimates the holding state of the held object 80 when the class inferred by the class inference model 40 (e.g., the third class) or another class (e.g., the first, second, or fourth class) is estimated, and generates the estimation result as a holding state candidate and displays it on the display device 20 (step S3). Specifically, the control unit 12 displays an image in which the holding position 82 is superimposed on the object image. At this time, the display may be made in a manner that allows the class estimated by the class inference model 40 to be identified.

[0063] The control unit 12 allows the user to select a holding mode from candidates (step S4). Specifically, the control unit 12 allows the user to input which holding mode to select from the holding mode candidates via the I / F 14, and acquires information specifying the holding mode selected by the user. In other words, if the user thinks that the holding mode for the class estimated by the class inference model 40 is not appropriate for the holding object 80, the control unit 12 allows the user to select another holding mode.

[0064] The control unit 12 causes the trained model to estimate a holding manner so that the holding object 80 is held in the holding manner selected by the user, and outputs the estimated holding manner to the robot control device 110 (step S5). After executing the procedure of step S5, the control unit 12 ends the execution of the procedure of the flowchart in FIG.

[0065] <Summary> As described above, the holding mode determination device 10 according to this embodiment allows the user to select a holding mode estimated by the trained model, and can control the robot 2 in the holding mode selected by the user. In this way, the user can control the holding mode of the robot 2 even if the user cannot control the holding mode estimation result by the trained model. As a result, the convenience of the trained model that outputs the holding mode of the holding object 80 is improved.

[0066] (Other embodiments) Other embodiments are described below.

[0067] In the above-described embodiment, the configuration has been described in which the holding mode determination device 10 determines the holding position 82 as the holding mode. The holding mode determination device 10 can determine other modes as the holding mode in addition to the holding position 82.

[0068] The holding mode determination device 10 may determine, as the holding mode, for example, the force that the robot 2 applies to hold the held object 80. In this case, the holding mode inference model 50 estimates and outputs the force that is applied to hold the held object 80.

[0069] Furthermore, the holding mode determination device 10 may determine, as the holding mode, for example, the type of hand with which the robot 2 holds the held object 80. In this case, the holding mode inference model 50 estimates and outputs the type of hand used to hold the held object 80.

[0070] The holding mode determination device 10 may superimpose the holding mode estimated by the holding mode inference model 50 on the object image and display it on the display device 20. Furthermore, when the robot 2 has at least two fingers that hold the held object 80, the holding mode determination device 10 may display the position where the held object 80 is held by the fingers as the holding position 82. Furthermore, the holding mode determination device 10 may display, on the display device 20, a position within a predetermined range from the position where the held object 80 is held.

[0071] The above has described embodiments of the holding mode determination device 10 and the robot control system 100. However, embodiments of the present disclosure can also be embodied as a method or program for implementing the system or device, or as a storage medium on which a program is recorded (for example, an optical disk, a magneto-optical disk, a CD-ROM, a CD-R, a CD-RW, a magnetic tape, a hard disk, or a memory card, etc.).

[0072] Furthermore, the implementation form of the program is not limited to application programs such as object code compiled by a compiler or program code executed by an interpreter, but may also be in the form of a program module incorporated into an operating system. Furthermore, the program may or may not be configured so that all processing is performed solely by the CPU on the control board. The program may also be configured so that part or all of it is executed by another processing unit mounted on an expansion board or expansion unit added to the board as needed.

[0073] Although the embodiments of the present disclosure have been described based on the drawings and examples, it should be noted that those skilled in the art could make various modifications or alterations based on the present disclosure. Therefore, it should be noted that these modifications or alterations are included in the scope of the present disclosure. For example, the functions included in each component can be rearranged so as not to cause logical inconsistencies, and multiple components can be combined or divided into one.

[0074] All of the features described in this disclosure and / or all steps of all of the disclosed methods or processes may be combined in any combination except combinations in which these features are mutually exclusive. Furthermore, each feature described in this disclosure may be replaced by an alternative feature serving the same, equivalent, or similar purpose, unless expressly denied. Thus, unless expressly denied, each disclosed feature is only one example of a generic series of identical or equivalent features.

[0075] Furthermore, embodiments of the present disclosure are not limited to the specific configurations of any of the above-described embodiments, but rather extend to any novel feature or combination thereof described herein, or any novel method or process step or combination thereof described herein.

[0076] In this disclosure, descriptions such as "first" and "second" are identifiers for distinguishing the configuration. In this disclosure, configurations distinguished by descriptions such as "first" and "second" can have their numbers exchanged. For example, the first class can exchange the identifiers "first" and "second" with the second class. The exchange of identifiers is performed simultaneously. The configurations remain distinguished even after the exchange of identifiers. Identifiers may be deleted. A configuration from which an identifier has been deleted is distinguished by a symbol. The descriptions of identifiers such as "first" and "second" in this disclosure should not be used solely to interpret the order of the configurations or to justify the existence of an identifier with a smaller number. [Explanation of symbols]

[0077] 10 holding mode determination device (12: control unit, 14: interface) 20 Display device 40 Class Inference Model (42: Processing Layer) 50 retention mode inference model (52: processing layer, 54: multiplier (541 to 544: first to fourth multipliers), 56: adder) 60 Conversion unit 80 Holding object (83: screw head, 84: screw shaft) 81 Shape Images 82 Holding position 100 Robot control system (2: Robot, 2A: Arm, 2B: End effector, 3: Sensor, 4: Camera, 5: Robot influence range, 6: Work start point, 7: Work target point, 110: Robot control device)

Claims

1. a control unit that outputs, to a display device, at least one holding state of the object by the robot estimated by classifying the object into at least one of a plurality of holding categories that can be estimated, together with an image of the object; an interface for acquiring a user's selection for the at least one holding manner via the display device; A robot holding mode determination device comprising:

2. 2. The robot holding mode determination device according to claim 1, wherein the interface outputs the holding mode selected by the user to a robot control device that controls the operation of the robot so that the robot holds the object in the holding mode selected by the user.

3. 3. The robot holding mode determination device according to claim 1, wherein the plurality of holding categories that can be estimated are holding categories that can be output by a class inference model that outputs a classification result that classifies the object into a predetermined class based on an image of the object.

4. 4. The robot holding mode determination device according to claim 3, wherein when the classification results are input to a holding mode inference model that outputs the holding mode of the object based on an input holding classification, the control unit outputs the holding mode for each of the classification results to the display device.

5. The robot holding manner determination device according to claim 4 , wherein the holding manner inference model determines a position at which the robot holds the object as a manner in which the robot holds the object.

6. 6. The robot holding mode determination device according to claim 4, wherein the holding mode inference model determines a force to be applied by the robot to hold the object as a mode in which the robot holds the object.

7. 7. The robot holding mode determination device according to claim 4, wherein the holding mode inference model determines a type of hand with which the robot holds the object as a mode with which the robot holds the object.

8. the robot has a holding unit that holds the object based on an image of the object acquired from a camera and an output of a trained model including the class inference model and the holding manner inference model; 8. The robot holding mode determination device according to claim 4, wherein the control unit outputs the user's selection to the robot so that the robot holds the object in the holding mode output by a holding mode inference model based on the class corresponding to the user's selection.

9. the holding portion has at least two fingers for holding the object; 9. The robot holding mode determination device according to claim 8, wherein the control unit outputs, to the display device, a position where the object is pinched by the fingers, or a position within a predetermined range from the position where the object is pinched, as the holding mode.

10. The robot holding mode determination device according to claim 1 , wherein the control unit outputs to the display device an image of the object superimposed with a mode of holding the object.

11. The control unit sets a conversion rule for converting a classification result of a holding category output from a class inference model that classifies an object into at least one of a plurality of holding categories that can be estimated into another holding category.

12. The robot holding mode determination device according to claim 1 , wherein the control unit determines whether a foreign object exists based on a result of classifying the object into at least one of a plurality of possible holding categories.

13. outputting, to a display device, at least one holding mode of the object by the robot estimated by classifying the object into at least one of a plurality of holding categories that can be estimated, together with an image of the object; obtaining a user selection for the at least one holding manner via the display device; and A method for determining a holding mode of a robot, comprising:

14. a robot; a robot control device that controls the robot; a holding mode determination device that determines a mode in which the robot holds an object; and a display device; the holding mode determination device outputs, to the display device, at least one holding mode of the object of the robot estimated by classifying the object into at least one of a plurality of holding categories that can be estimated, together with an image of the object; the display device displays an image of the object and at least one holding mode of the object, receives a user's input of a selection of the at least one holding mode, and outputs the selection to the holding mode determination device; the holding mode determination device acquires, from the display device, the user's selection of the at least one holding mode; Robot control system.

Citation Information

Patent Citations

  • Method for gripping object of optional shape by robot

    JP2005169564A