Learning method, learning device, processing device, and program
A trained model using Mask Scoring R-CNN improves robotic object recognition and handling by accurately identifying and gripping hard portions of objects with combined soft and hard shapes, addressing challenges in robotic assembly tasks.
Patent Information
- Application Number
- PCT/JP2025/023749
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-02
- Filing Date
- 2025-07-01
- Publication Date
- 2026-01-08
AI Technical Summary
Existing technologies face challenges in accurately recognizing and handling objects with varying shapes, particularly when soft and hard portions are combined, such as nozzles with hoses, which complicates robotic assembly tasks.
A learning method and device that utilize a trained model, like Mask Scoring R-CNN, to recognize hard portions of objects using color and depth images, enabling precise robotic handling by identifying and manipulating these portions.
Enhances the accuracy and reliability of robotic object recognition and assembly by effectively distinguishing and gripping the hard portions of objects, even when soft portions like hoses are present, improving task success rates.
Smart Images

Figure JP2025023749_08012026_PF_FP_ABST
Abstract
Description
Learning method, learning device, processing device, and program
[0001] The present disclosure relates to a technology for recognizing objects.
[0002] Patent Document 1 describes a technology for recognizing an object.
[0003] JP 2010-120414 A
[0004] A learning method, a learning device, a processing device, and a program are disclosed. In one embodiment, the learning method generates a trained model that recognizes a recognition target portion that is a part of a work object based on an inference image that shows the work object to be worked by a robot, by training a pre-trained model based on training data that includes the training image and annotation data related to the recognition target portion.
[0005] In one embodiment, a learning device executes the above learning method.
[0006] In one embodiment, the processing device includes a recognition unit that recognizes a recognition target portion, which is part of a work object, based on an image showing the work object that is the target of work by the robot.
[0007] In one embodiment, the program is a program for causing a computer device to execute the above learning method.
[0008] In one embodiment, the program is a program for causing a computer device to function as the processing device.
[0009] FIG. 1 is a schematic diagram showing an example of a robot system. FIG. 2 is a schematic diagram showing an example of a work target. FIG. 3 is a schematic diagram showing an example of the configuration of a processing device. FIG. 4 is a schematic diagram showing an example of the configuration of a control unit. FIG. 5 is a flowchart showing an example of the operation of the robot system. FIG. 6 is a schematic diagram showing an example of the operation of the robot system. FIG. 7 is a schematic diagram showing an example of a color image. FIG. 8 is a schematic diagram showing an example of a binary image including a mask of a recognition target portion. FIG. 9 is a schematic diagram showing an example of a color image with a mask superimposed. FIG. 10 is a schematic diagram showing an example of 3D matching. FIG. 11 is a schematic diagram showing an example of the operation of the robot system. FIG. 12 is a schematic diagram showing an example of the operation of the robot system. FIG. 13 is a schematic diagram showing an example of the operation of the robot system. FIG. 14 is a schematic diagram showing an example of the operation of the robot system. FIG. 15 is a schematic diagram showing an example of a color image. FIG. 16 is a schematic diagram showing an example of a color image. FIG. 17 is a schematic diagram showing an example of a binary image showing only a hose. FIG. 18 is a schematic diagram showing an example of a binary image showing only a hose. FIG. 19 is a schematic diagram showing an example of the operation of the robot system. FIG. 20 is a schematic diagram showing an example of the operation of the robot system. Fig. 21 is a schematic diagram showing an example of the operation of the robot system. Fig. 22 is a schematic diagram for explaining an example of a learning method. Fig. 23 is a schematic diagram showing an example of a learning image. Fig. 24 is a schematic diagram showing an example of a learning image. Fig. 25 is a schematic diagram showing an example of a learning image.
[0010] Fig. 1 is a schematic diagram showing an example of a robot system 50. As shown in Fig. 1, the robot system 50 includes, for example, a robot 10, a processing device 1, a camera 15, and a camera 16.
[0011] The robot 10 is, for example, an arm-type robot and includes an arm 11 and an end effector 12 connected to the arm 11. The end effector 12 is, for example, a holding mechanism capable of holding an object. The end effector 12 may be, for example, a hand including a plurality of fingers 13 capable of grasping an object. Each finger 13 is, for example, a linearly extending, thin member. In the example of FIG. 1 , the end effector 12 includes two fingers 13, but the number of fingers 13 may be three or more. The end effector 12 can hold an object by pinching it between the plurality of fingers 13. The end effector 12 may also be an adsorption mechanism including an adsorption portion that adsorbs an object.
[0012] The arm 11 is fixed on a base 18. The arm 11 includes, for example, a plurality of joints. The posture of the arm 11 changes as the amount of rotation of at least one of the plurality of joints changes. The change in posture of the arm 11 changes the position and posture of the end effector 12. The change in posture of the arm 11 also changes the position and posture of an object held by the end effector 12.
[0013] For example, the robot 10 can move an object from a first location to a second location different from the first location. The robot 10 can also perform a task of combining an object with another object. Specifically, the robot 10 can combine the container body 25 of a spray bottle with the nozzle 20 of the spray bottle.
[0014] In this example, a plurality of nozzles 20 are stacked in a random pile in the first tray 30. A plurality of container bodies 25 are arranged in a line in the second tray 35. The robot 10 holds the nozzles 20 in the first tray 30 one by one and assembles the held nozzles 20 with the container bodies 25 in the second tray 35. Note that the plurality of container bodies 25 do not necessarily have to be arranged in the second tray 35, and may be arranged in a line on a work table.
[0015] The nozzle 20 and the container body 25 are each a work object of the robot 10. In other words, the nozzle 20 and the container body 25 are each an object on which the robot 10 performs work. The work object is a holdable object that can be held by the end effector 12, and can also be said to be a graspable object that can be grasped by the multiple fingers 13 of the end effector 12.
[0016] FIG. 2 is a schematic diagram showing an example of a nozzle 20. The nozzle 20 includes, for example, a head 21, a hose 24, and a connection portion 22 connecting the head 21 and the hose 24. The connection portion 22 has an elongated shape and extends straight from the head 21. The head 21 has a multi-stepped outer shape. The diameter of the head 21 is larger than the diameter of the connection portion 22. The connection portion 22 is connected to the largest diameter portion of the multi-stepped outer shape of the head 21 so as to extend from that portion. Specifically, the connection portion 22 is connected to the concave underside of the head 21 and extends from that underside beyond the head 21. As a result, the outer shapes of the head 21 and the connection portion 22 have a shape with a step between them. As a result, the connection portion 22 is the largest step in the multi-step structure of the hard portion 23 described below.
[0017] If the work object has a step, when a first region with a smaller diameter and a second region with a larger diameter are recognized with the step as a boundary, the holding position for the robot 10 to hold the work object may be set within the first region. Also, if the work object has multiple steps, the holding position may be set within the first region in the part with the larger step.
[0018] The head 21 and the connection portion 22 are relatively hard and have a fixed shape. Therefore, there is little variation in the shapes of the head 21 and the connection portion 22 among the multiple nozzles 20. On the other hand, the hose 24 is relatively soft and is easily bent by an external force. More specifically, the head 21 is a harder part than the connection portion 22. The connection portion 22 is a harder part than the hose 24. Furthermore, the overall length of the head 21 and the connection portion 22 is shorter than that of the hose 24. As a result, the head 21 and the connection portion 22 have a fixed shape compared to the hose 24.
[0019] The shape of the hose 24 varies among the multiple nozzles 20. For example, there are some nozzles 20 in which the hose 24 extends straight from the connection portion 22, and there are other nozzles 20 in which the hose 24 is bent. The degree to which the hose 24 is bent also varies among the multiple nozzles 20. Therefore, it can be said that the shape of the hose 24 is irregular. Hereinafter, the head 21 and the connection portion 22, which are harder than the hose 24, may be collectively referred to as the hard portion 23. In contrast to the hard portion 23, the hose 24 can also be said to be a soft portion.
[0020] The robot 10 holds the hard portion 23 of the nozzle 20 in the first tray 30, for example, with the multiple fingers 13 of the end effector 12. More specifically, the robot 10 holds the connection portion 22 of the hard portion 23. Then, the robot 10 changes the posture of the arm 11 to move the held nozzle 20 above the second tray 35. The robot 10 then changes the posture of the arm 11 to insert the hose 24 of the held nozzle 20 into the container body 25 through the opening 25a at the top end of the container body 25 in the second tray 35. At this time, the robot 10 inserts the hose 24 into the container body 25 from its tip. After inserting the hose 24 into the container body 25, the robot 10 opens the multiple fingers 13 to release its hold on the connection portion 22 of the nozzle 20. This allows the nozzle 20 and the container body 25 to be combined. Thereafter, the robot 10 holds the next nozzle 20 in the first tray 30 and inserts the hose 24 of the held nozzle 20 into another container body 25 in the second tray 35. Thereafter, the robot 10 operates in the same manner. Note that the portion held by the end effector 12 may be the head 21 of the hard portion 23 or the end of the hose 24 on the connection portion 22 side.
[0021] The camera 15 is fixed to, for example, the end effector 12. Therefore, the imaging range of the camera 15 changes depending on the position and posture of the end effector 12. It can also be said that the imaging range of the camera 15 changes depending on the posture of the arm 11. The camera 15 fixed to the end effector 12 can capture, for example, the tip side of the finger portion 13 of the end effector 12. The optical axis direction of the camera 15 is set, for example, parallel to the length direction of the finger portion 13. For example, when the end effector 12 is positioned above the first tray 30 such that the length direction of the finger portion 13 is aligned perpendicular to the opening surface at the top end of the first tray 30 and the tip of the finger portion 13 faces toward the first tray 30, the camera 15 can capture the state inside the first tray 30. In other words, the camera 15 can capture an image of the multiple nozzles 20 randomly stacked in the first tray 30.
[0022] The camera 15 is, for example, a three-dimensional camera. A camera image 150 (see FIG. 3 ) generated by the camera 15 includes, for example, a color image 151 and a depth image 152. The camera 15 outputs the generated camera image 150 to the processing device 1. The depth image 152 may be generated using a stereo system, a projector system, a combination of the stereo system and the projector system, or another system. The depth image 152 is, for example, a grayscale image. The color image 151 may be, for example, an RGB image. The color image 151 can also be considered a captured image that shows the state of the camera 15's capture range. Hereinafter, the camera 15 may be referred to as an effector camera 15.
[0023] The camera 16 is fixed, for example, on the stand 18 to which the arm 11 is fixed. The imaging range of the camera 16 is fixed. The camera 16 is capable of imaging, for example, the area above the first tray 30 from the side. The optical axis direction of the camera 16 is set, for example, parallel to the horizontal plane. The camera 16 is capable of imaging along the horizontal direction. It can also be said that the optical axis direction of the camera 16 is set parallel to the opening plane of the first tray 30. Note that the camera 16 does not need to be fixed on the stand 18, and may be fixed to other members such as a wall, a pillar, or a frame.
[0024] The camera 16 is, for example, a three-dimensional camera. A camera image 160 (see FIG. 3 ) generated by the camera 16 includes, for example, a color image 161 and a depth image 162. The camera 16 outputs the generated camera image 160 to the processing device 1. The depth image 162 may be generated using a stereo method, a projector method, a combination of the stereo method and the projector method, or another method. The depth image 162 is, for example, a grayscale image. The color image 161 may be, for example, an RGB image. Hereinafter, the camera 16 may be referred to as a fixed camera 16.
[0025] The processing device 1 can recognize each nozzle 20 in the first tray 30 based on a color image 151 that shows the state inside the first tray 30 and is generated by the effector camera 15. For example, the processing device 1 can recognize a recognition target portion that constitutes part of the nozzle 20. Then, the processing device 1 controls the robot 10 based on the recognition result of the recognition target portion, and causes the robot 10 to hold the recognition target portion. The processing device 1 in this example functions as a control device that controls the robot 10.
[0026] In this example, the hard portion 23 that constitutes part of the nozzle 20 is the recognition target portion. In other words, the head 21 and the connection portion 22 are the recognition target portions. The processing device 1 can recognize the hard portion 23 of the nozzle 20. When performing work, the robot 10 holds a predetermined portion of the hard portion 23. Specifically, the robot 10 holds a predetermined portion of the connection portion 22 of the hard portion 23. Hereinafter, the hard portion 23 may be referred to as the recognition target portion 23.
[0027] Fig. 3 is a schematic diagram showing an example of the configuration of the processing device 1. As shown in Fig. 3, the processing device 1 includes, for example, a control unit 2, a storage unit 3, and interfaces 4, 5, and 6. The processing device 1 can also be considered, for example, a computer device. The processing device 1 can also be considered, for example, a processing circuit.
[0028] The interface 5 is capable of communicating with the effector camera 15. The control unit 2 can acquire the camera image 150 generated by the effector camera 15 through the interface 5. The interface 5 can also be called, for example, an interface circuit, a communication unit, or a communication circuit. The interface 5 may communicate with the effector camera 15 via wired or wireless communication.
[0029] The interface 6 is capable of communicating with the fixed camera 16. The control unit 2 can acquire the camera image 160 generated by the fixed camera 16 through the interface 6. The interface 6 can also be referred to as, for example, an interface circuit, a communication unit, or a communication circuit. The interface 6 may communicate with the fixed camera 16 via wired communication or wireless communication.
[0030] The interface 4 is capable of communicating with the robot 10. The control unit 2 is capable of controlling the robot 10 through the interface 4. The interface 4 may also be referred to as, for example, an interface circuit, a communication unit, or a communication circuit. The interface 4 may communicate with the robot 10 via wired or wireless communication.
[0031] The control unit 2 can generally manage the operation of the processing device 1 by controlling the other components of the processing device 1. The control unit 2 can also be referred to as a control circuit, for example. The control unit 2 includes at least one processor to provide control and processing power for performing various functions, as described in more detail below.
[0032] According to various embodiments, the at least one processor may be implemented as a single integrated circuit (IC) or as multiple communicatively connected integrated circuits ICs and / or discrete circuits. The at least one processor may be implemented according to various known techniques.
[0033] In one embodiment, a processor includes one or more circuits or units configured to perform one or more data computational procedures or processes, for example, by executing instructions stored in associated memory. In other embodiments, a processor may be firmware (e.g., discrete logic components) configured to perform one or more data computational procedures or processes.
[0034] According to various embodiments, the processor may include one or more processors, controllers, microprocessors, microcontrollers, application specific integrated circuits (ASICs), digital signal processors, programmable logic devices, field programmable gate arrays, or any combination of these devices or configurations, or other known devices and configurations, to perform the functions described below.
[0035] The control unit 2 may include, for example, a CPU (Central Processing Unit) as a processor. The storage unit 3 may include a non-transitory recording medium readable by the CPU of the control unit 2, such as a ROM (Read Only Memory) and a RAM (Random Access Memory). The storage unit 3 stores, for example, a program 3a for controlling the processing device 1. Various functions of the control unit 2 are realized, for example, by the CPU of the control unit 2 executing the program 3a in the storage unit 3.
[0036] The configuration of the processing device 1 is not limited to the above example. For example, the control unit 2 may include multiple CPUs. The control unit 2 may also include at least one DSP (Digital Signal Processor). All or some of the functions of the control unit 2 may be realized by a hardware circuit that does not require software to realize the function. The storage unit 3 may also include a computer-readable non-transitory recording medium other than ROM and RAM. The storage unit 3 may also include, for example, a small hard disk drive or SSD (Solid State Drive).
[0037] The processing device 1 may also include a display unit controlled by the control unit 2. The display unit may be, for example, a liquid crystal display, an organic electroluminescence (EL) display, or a plasma display. The display unit of the processing device 1 may, for example, display at least one of a color image 151 and a depth image 152 generated by the camera 15, or at least one of a color image 161 and a depth image 162 generated by the camera 16. The display unit of the processing device 1 may also display various recognition results of the work object.
[0038] The processing device 1 may also include an input unit that accepts input from a user. The input unit may include, for example, a mouse and a keyboard. The input unit may also include a touch sensor that accepts touch operations by the user. The input unit may also include a microphone that accepts voice input by the user.
[0039] The processing device 1 may also be configured with a plurality of computer devices. The processing device 1 may also be a cloud server. In this case, the interface 4 of the processing device 1 may communicate with the robot 10 through a network including the Internet.
[0040] The robot system 50 may also be provided with an overhead camera that captures video of the working environment of the robot 10. In this case, the video (e.g., color image) generated by the overhead camera may be displayed on the display unit of the processing device 1.
[0041] <Configuration example of control unit of processing device> Fig. 4 is a schematic diagram showing an example of the configuration of the control unit 2. As shown in Fig. 4, the control unit 2 includes, for example, a robot control unit 120, a recognition unit 121, a selection unit 122, an attitude identification unit 123, a holding position identification unit 124, a holdability determination unit 125, and a tip position identification unit 126. The robot control unit 120, the recognition unit 121, the selection unit 122, the attitude identification unit 123, the holding position identification unit 124, the holdability determination unit 125, and the tip position identification unit 126 are each a functional block generated when the control unit 2 executes the program 3a in the storage unit 3.
[0042] The robot control unit 120 can control the robot 10 through the interface 4. The recognition unit 121 recognizes the recognition target portion 23 of each nozzle 20 in the first tray 30 based on a color image 151 showing the state inside the first tray 30.
[0043] Based on the recognition result by the recognition unit 121, the selection unit 122 selects the recognition target portion 23 to be held by the robot 10 (in other words, the recognition target portion 23 of the nozzle 20 that will cause the robot 10 to perform a task) from the recognition target portions 23 of the multiple nozzles 20 in the first tray 30. Hereinafter, the recognition target portion 23 selected by the selection unit 122 may be referred to as the selected recognition target portion 23.
[0044] The selection of the recognition target portion 23 can be seen as the selection of the nozzle 20 having the recognition target portion 23. Therefore, it can also be said that the selection unit 122 selects the nozzle 20 to be held by the robot 10 (in other words, the nozzle 20 to be made to perform a task by the robot 10) from the plurality of nozzles 20 in the first tray 30 based on the recognition result by the recognition unit 121. Hereinafter, the nozzle 20 having the recognition target portion 23 selected by the selection unit 122, that is, the nozzle 20 selected by the selection unit 122, may be referred to as the selected nozzle 20.
[0045] The posture identification unit 123 identifies the posture of the selected recognition target portion 23 selected by the selection unit 122 based on the recognition result from the recognition unit 121 and a depth image 152 obtained by photographing the inside of the first tray 30 with the camera 15.
[0046] The holding position specifying unit 124 specifies positions (also referred to as holding positions) of the plurality of fingers 13 of the robot 10 when the plurality of fingers 13 hold predetermined locations of the connection portion 22 included in the selected recognition target portion 23, based on the orientation of the selected recognition target portion 23 specified by the orientation specifying unit 123. In this example, the locations of the connection portion 22 of the nozzle 20 that the plurality of fingers 13 hold are determined in advance.
[0047] The holdability determination unit 125 determines whether the multiple fingers 13 can hold the connection portion 22 of the selected recognition target portion 23 at the holding position identified by the holding position identification unit 124, based on the depth image 152 obtained by the camera 15 capturing the state inside the first tray 30.
[0048] When the holdability determination unit 125 determines that the multiple finger portions 13 are capable of holding the selected recognition target portion 23, the robot control unit 120 controls the robot 10 so that the multiple finger portions 13 hold predetermined locations of the connection portion 22 of the selected recognition target portion 23.
[0049] When the selection recognition target portion 23 is held by the multiple fingers 13, in other words, when the selection nozzle 20 is held by the multiple fingers 13, the robot control unit 120 controls the robot 10 to move the work object to a location where at least the tip of the hose 24 can be simultaneously photographed by both the camera 15 and the camera 16. Then, the tip position identification unit 126 identifies the position of the tip of the hose 24 of the selection nozzle 20 held by the multiple fingers 13, based on a color image 151 of the hose 24 of the selection nozzle 20 obtained by the camera 15 and a color image 161 of the hose 24 of the selection nozzle 20 obtained by the camera 16. Based on the position of the tip of the hose 24 of the selection nozzle 20 identified by the tip position identification unit 126, the robot control unit 120 controls the robot 10 so that the hose 24 is inserted, tip-first, into the container body 25 in the second tray 35. Note that the cameras 15 and 16 do not have to photograph simultaneously, and may photograph at different times. In this case, the robot control unit 120 may control the robot 10 to sequentially move the work object to a position where each camera can take an image.
[0050] Note that all or some of the functions of the robot control unit 120 may be realized by hardware circuits that do not require software to realize the functions. The same applies to the recognition unit 121, the selection unit 122, the orientation identification unit 123, the holding position identification unit 124, the holdability determination unit 125, and the tip position identification unit 126. The operation of each component of the control unit 2 will be described in more detail later.
[0051] <Regarding the Recognition Unit> The recognition unit 121 is configured, for example, by a trained model, which is a training model that has been trained. The training model that configures the recognition unit 121 is configured, for example, by a neural network that realizes instance segmentation. The training model is configured, for example, by Mask Scoring R-CNN. R-CNN is an abbreviation for Region Based Convolutional Neural Networks. The training model can also be referred to, for example, as a machine learning model. The trained model is generated by training a pre-training model, which is a training model before learning. Hereinafter, the recognition unit 121 may be referred to as the trained model 121.
[0052] A color image 151 showing the state inside the first tray 30 is input to the trained model 121. The trained model 121 recognizes the recognition target portion 23 of each nozzle 20 inside the first tray 30 based on the input color image 151. The color image 151 can be considered an image for inference. The trained model 121 then outputs the recognition result of the recognition target portion 23 of each nozzle 20 inside the first tray 30.
[0053] The recognition result output by the trained model 121 includes, for example, a mask of the recognition target portion 23 of each nozzle 20 in the first tray 30. Specifically, the recognition result output by the trained model 121 includes position information and shape information of the mask of the recognition target portion 23 of each nozzle 20 in the first tray 30. The position and shape of the mask of the recognition target portion 23 represent the position and shape of the recognition target portion 23. It can also be said that the trained model 121 recognizes the position and shape of the recognition target portion 23.
[0054] Furthermore, the recognition result output by the trained model 121 includes, for example, a classification score that indicates the accuracy of classification of the recognition target portion 23 for each nozzle 20 in the first tray 30. The classification score can be said to indicate the degree of certainty in the classification of the recognition target portion 23, the likelihood of the classification of the recognition target portion 23, or the likelihood of the recognition target portion 23. The classification score is, for example, a numerical value greater than or equal to 0 and less than or equal to 1. The larger the classification score, the higher the accuracy of classification of the recognition target portion 23. The classification score can also be said to indicate the probability that the mask generated by the trained model 121 is a mask of the recognition target portion 23. The score can also be said to be an evaluation value.
[0055] Furthermore, the recognition result output by the trained model 121 includes, for example, a mask score indicating the accuracy of generation of each mask generated by the trained model 121. The mask score is, for example, a numerical value greater than or equal to 0 and less than or equal to 1. The larger the mask score, the higher the accuracy of mask generation. The mask score can be said to indicate the degree of confidence in the position information and shape information of the mask generated by the trained model 121, or the likelihood of the position information and shape information of the mask, or the certainty of the position information and shape information of the mask.
[0056] Furthermore, the recognition result output by the trained model 121 includes a visibility score indicating the proportion of the portion of the recognition target portion 23 that appears in the color image 151 relative to the entire image of the recognition target portion 23 for each nozzle 20 in the first tray 30. The visibility score is, for example, a numerical value greater than or equal to 0 and less than or equal to 1. The larger the visibility score, the greater the proportion of the portion of the recognition target portion 23 that appears in the color image 151 relative to the entire image of the recognition target portion 23.
[0057] Here, the recognition target portion 23 of the explanation target in the first tray 30, in other words, the recognition target portion 23 of interest in the first tray 30, is referred to as the "attention recognition target portion 23".
[0058] For example, consider a case where another object (such as a hose 24) overlaps the attention recognition target portion 23, and part of the attention recognition target portion 23 is not captured in the color image 151. In this case, the visibility score of the attention recognition target portion 23 will be small. Also consider a case where part of the attention recognition target portion 23 is located outside the shooting range of the camera 15, and part of the attention recognition target portion 23 is not captured in the color image 151. In this case as well, the visibility score of the attention recognition target portion 23 will be small.
[0059] On the other hand, if the entire target recognition portion 23 is located within the shooting range of the camera 15 and no other objects overlap the target recognition portion 23, the visibility score of the target recognition portion 23 will be the maximum value of 1.
[0060] The area of the region in the color image 151 in which the attention recognition target portion 23 actually appears is referred to as the visible area of the attention recognition target portion 23. Furthermore, the area of the region in the color image 151 in which the attention recognition target portion 23 appears, assuming that the entire attention recognition target portion 23 appears in the color image 151, is referred to as the total area of the attention recognition target portion 23. For example, a value obtained by dividing the visible area of the attention recognition target portion 23 by the total area of the attention recognition target portion 23 may be used as the visibility score of the attention recognition target portion 23. The visible area of the attention recognition target portion 23 can also be said to be the area of a mask of the attention recognition target portion 23.
[0061] When a part of the attention recognition target portion 23 is located outside the shooting range of the camera 15, the total area is the area of the region in the color image 151 in which the attention recognition target portion 23 appears if it is assumed that the shooting range of the camera 15 is enlarged so that the part is included in the shooting range. Also, when the entire attention recognition target portion 23 is located within the shooting range of the camera 15 but another object overlaps the attention recognition target portion 23, the total area is the area of the region in the color image 151 in which the attention recognition target portion 23 appears if it is assumed that the other object does not overlap the attention recognition target portion 23. When the entire attention recognition target portion 23 is located within the shooting range of the camera 15 and no other object overlaps the attention recognition target portion 23, the visible area and the total area of the attention recognition target portion 23 are the same.
[0062] The visibility score of the attention recognition target portion 23 can also be said to indicate the proportion of the attention recognition target portion 23 that is actually visible from the effector camera 15 to the entire attention recognition target portion 23 that is visible from the effector camera 15, assuming that the entire attention recognition target portion 23 is visible from the effector camera 15.
[0063] If the camera 15 fixed to the end effector 12 is considered to be the eyes of the robot 10, the color image 151 obtained by the camera 15 can be considered to reflect what the robot 10 sees with its eyes. Therefore, the visibility score of the target recognition portion 23 can be said to represent the visibility (in other words, the degree of visibility) of the target recognition portion 23 when viewed from the robot 10. Therefore, the larger the visibility score of the target recognition portion 23, the easier it is for the robot 10 to see the target recognition portion 23, and as a result, the easier it is for the robot 10 to hold the target recognition portion 23. The visibility score of the target recognition portion 23 can also be said to represent the degree of visibility when the robot 10 views the target recognition portion 23.
[0064] As described above, the recognition unit 121, in other words, the trained model 121, outputs a mask, classification score, mask score, and visibility score for each of the recognition target portions 23 of the multiple nozzles 20 in the first tray 30.
[0065] 5 is a flowchart showing an example of the operation of the robot system 50 when the robot 10 holds the nozzle 20 in the first tray 30 and inserts the hose 24 of the held nozzle 20 into the container body in the second tray 35. While the robot 10 is performing the operation, for example, a video generated by the overhead camera described above may be displayed on the display unit of the processing device 1. This allows the user of the processing device 1 to check the video displayed on the display unit and see how the robot 10 is performing the operation.
[0066] As shown in Fig. 5, in step s1, the robot control unit 120 controls the arm 11 to move the end effector 12 so that the effector camera 15 can capture images of the interior of the first tray 30. Fig. 6 is a schematic diagram showing an example of the state in which the end effector 12 has moved so that the effector camera 15 can capture images of the interior of the first tray 30. Compared to Fig. 1, Fig. 6 omits the illustration of the arm 11, the platform 18, and the processing device 1. The same applies to later-described figures showing the robot system 50.
[0067] Next, in step s2, the control unit 2 causes the effector camera 15 to take a photograph. The effector camera 15 takes a photograph of the state inside the first tray 30 from directly above the first tray 30. The effector camera 15 generates a color image 151 (also referred to as a first color image 151a) showing the state inside the first tray 30. The first color image 151a may be displayed on the display unit of the processing device 1.
[0068] 7 is a schematic diagram showing an example of the first color image 151 a. As shown in FIG. 6, the first color image 151 a shows the first tray 30 and a plurality of nozzles 20 randomly stacked in the first tray 30.
[0069] Next, the recognition unit 121 (in other words, the trained model 121) recognizes the recognition target portion 23 of each nozzle 20 in the first tray 30 based on the first color image 151a. It can also be said that the recognition unit 121 detects or searches for whether the recognition target portion 23 is captured in the first color image 151a. For the recognition target portion 23 of each nozzle 20, the recognition unit 121 outputs a mask, a classification score, a mask score, and a visibility score for the recognition target portion 23.
[0070] 8 is a schematic diagram showing an example of a binary image 100 including a mask 110 of the target recognition portion 23. The binary image 100 is an image of the same size as the first color image 151a. The position of the mask 110 in the binary image 100 is the same as the position of the area in the first color image 151a where the target recognition portion 23 appears. The outer shape of the mask 110 included in the binary image 100 is the same as the outer shape of the area in the first color image 151a where the target recognition portion 23 appears. The control unit 2 may generate the binary image 100 and display it on the display unit of the processing device 1.
[0071] 9 is a schematic diagram showing an example of a state in which a plurality of masks 110 obtained by the recognition unit 121 are superimposed on a first color image 151a. In the example of FIG. 9, the masks 110 of the recognition target portions 23 are superimposed on the areas of the first color image 151a in which the recognition target portions 23 appear. In the example of FIG. 9, the masks 110 are shown with sandy hatching, but may be shown in a color different from that of the recognition target portions 23. The control unit 2 may generate the first color image 151a shown in FIG. 9 on which the masks 110 are superimposed, and display it on the display unit of the processing device 1.
[0072] After step s3, in step s4, the selection unit 122 selects a recognition target portion 23 to be held by the robot 10 from the recognition target portions 23 of the multiple nozzles 20 in the first tray 30 based on the recognition result of the recognition unit 121. In other words, the selection unit 122 selects a nozzle 20 to be held by the robot 10 from the multiple nozzles 20 in the first tray 30. It can also be said that the selection unit 122 selects a recognition target portion 23 to be held by the robot 10 from the multiple recognition target portions 23 recognized by the recognition unit 121.
[0073] The selection unit 122 selects a recognition target portion 23 to be held by the robot 10 from the recognition target portions 23 of the multiple nozzles 20 in the first tray 30 based on, for example, the classification score, mask score, and visibility score output from the recognition unit 121. For example, the selection unit 122 calculates a holdability score for each recognition target portion 23, indicating the possibility that the robot 10 will be able to hold the recognition target portion 23, based on the classification score, mask score, and visibility score of the recognition target portion 23. The selection unit 122 then selects, from the recognition target portions 23 of the multiple nozzles 20 in the first tray 30, the recognition target portion 23 with the highest holdability score as the recognition target portion 23 to be held by the robot 10. In other words, the selection unit 122 selects, from the multiple nozzles 20 in the first tray 30, the nozzle 20 having the recognition target portion 23 with the highest holdability score as the nozzle 20 to be held by the robot 10.
[0074] When calculating the retainability score of the recognition target portion 23, the selection unit 122, for example, multiplies the classification score, the mask score, and the visibility score individually by a weighting coefficient. Then, the selection unit 122 determines the value obtained by adding together the classification score, the mask score, and the visibility score multiplied by the weighting coefficient as the retainability score. Alternatively, the selection unit 122 may determine the value obtained by multiplying the classification score, the mask score, and the visibility score as the retainability score.
[0075] The selection unit 122 selects the nozzle 20 (in other words, the work object) that the robot 10 will work on while taking into consideration the visibility score, thereby increasing the likelihood that a nozzle 20 that is easy for the robot 10 to work on will be selected. This increases the likelihood that the robot 10 will be successful in working on the selected nozzle 20. In this example, the robot 10 is more likely to be successful in holding the selected nozzle 20.
[0076] When the display unit of the processing device 1 displays the first color image 151a (see Figure 9) on which the mask 110 is superimposed, it may display the retention possibility score of the recognition target portion 23 covered by the mask 110 near each mask 110 on the first color image 151a.
[0077] The selection unit 122 may select the recognition target portion 23 to be held by the robot 10 from the recognition target portions 23 of the multiple nozzles 20 in the first tray 30 based on two of the classification score, the mask score, and the visibility score. In this case, the selection unit 122 may, for example, determine the holdability score as a value obtained by adding together the two scores multiplied by the weighting coefficient, or may determine the holdability score as a value obtained by multiplying the two scores. The recognition unit 121 does not need to determine one score that is not used by the selection unit 122 from the classification score, the mask score, and the visibility score.
[0078] Furthermore, the selection unit 122 may select the recognition target portion 23 to be held by the robot 10 from the recognition target portions 23 of the multiple nozzles 20 in the first tray 30 based on one of the classification score, the mask score, and the visibility score. In this case, the selection unit 122 may use, for example, one of the scores to be used as the holdability score. The recognition unit 121 does not need to calculate the two scores not used by the selection unit 122 from the classification score, the mask score, and the visibility score.
[0079] After step s4, in step s5, the orientation identification unit 123 identifies the orientation of the selected recognition target portion 23 selected by the selection unit 122. In step s5, the orientation identification unit 123 generates a point cloud that three-dimensionally represents the surface of the selected recognition target portion 23, for example, based on a mask of the selected recognition target portion 23 and the depth image 152 obtained by the image capture by the effector camera 15 in step s2. The point cloud may also be referred to as a point cloud image. The depth image 152 (also referred to as a first depth image 152) obtained in step s2 is a depth image 152 obtained by the effector camera 15 capturing an image of the inside of the first tray 30 from above the first tray 30, as shown in FIG. 6 .
[0080] Next, the orientation identification unit 123 performs three-dimensional matching between the point cloud of the selected recognition target portion 23 and a point cloud serving as a template (also referred to as a template point cloud) that represents the original surface of the recognition target portion 23, thereby identifying the orientation of the selected recognition target portion 23. For example, the orientation identification unit 123 rotates the template point cloud in a virtual three-dimensional space so that the template point cloud matches the point cloud of the selected recognition target portion 23. Then, the orientation identification unit 123 identifies the orientation of the selected recognition target portion 23 based on the amount of rotation of the template point cloud when the template point cloud matches the point cloud of the selected recognition target portion 23 and the orientation of the template point cloud before rotation. The template point cloud may be generated based on, for example, a depth image 152 obtained by the effector camera 15 capturing an image of one recognition target portion 23.
[0081] FIG. 10 is a schematic diagram showing an example of a point cloud 183 of the selected recognition target portion 23 determined by the orientation identification unit 123, and an example of a template point cloud 180. As can be seen from FIG. 10 , the template point cloud 180 is also data of a portion corresponding to the recognition target portion, which is part of the work object. The lower part of FIG. 10 is a schematic diagram showing an example of the template point cloud 180 that matches the point cloud 183 of the selected recognition target portion 23 and is superimposed on the point cloud 183 of the selected recognition target portion 23. In the example of FIG. 10 , sand hatching is shown in the portion of the point cloud 183 of the selected recognition target portion 23 that overlaps with the template point cloud 180.
[0082] The display unit of the processing device 1 may display the template point cloud 180, or may display the point cloud 183 of the selected recognition target portion 23. Furthermore, the control unit 2 may generate the point cloud 183 of the selected recognition target portion 23 in which the matching template point cloud 180 is superimposed on the point cloud 183 of the selected recognition target portion 23, as shown in the lower part of Fig. 10 , and display it on the display unit of the processing device 1.
[0083] Here, if another object overlaps the selected recognition target portion 23, the mask of the selected recognition target portion 23 generated by the recognition unit 121 is a mask that represents the entire selected recognition target portion 23 except for the area where the other object overlaps. Therefore, if the area where the other object overlaps is large in the selected recognition target portion 23, the point cloud of the selected recognition target portion 23 obtained based on the mask of the selected recognition target portion 23 generated by the recognition unit 121 and the depth image 152 may be significantly different from the template point cloud. As a result, the template point cloud may not match the point cloud of the selected recognition target portion 23. In this example, because the recognition target portion 23 is selected based on the visibility score, there is a high possibility that the area where the other object overlaps in the selected recognition target portion 23 will be small. Therefore, there is a high possibility that the template point cloud will match the point cloud of the selected recognition target portion 23, and the orientation of the selected recognition target portion 23 can be appropriately determined.
[0084] After step s5, in step s6, the holding position specifying unit 124 specifies, based on the posture of the selected recognition target portion 23 specified by the posture specifying unit 123, predetermined locations of the connection portion 22 included in the selected recognition target portion 23 as holding positions of the plurality of fingers 13 of the robot 10 when the selected recognition target portion 23 is held by the plurality of fingers 13. The control unit 2 may cause the display unit of the processing device 1 to display a first color image 151a indicating, by color or the like, predetermined locations of the connection portion 22 of the selected recognition target portion 23 held by the plurality of fingers 13.
[0085] Next, in step s7, the holdability determination unit 125 determines, based on the first depth image 152, whether the multiple fingers 13 can hold the connection portion 22 of the selection recognition target portion 23 at the holding position identified by the holding position identification unit 124. The holdability determination unit 125 determines, for example, based on the first depth image 152, whether there is an obstacle in the moving path of the multiple fingers 13 when the multiple fingers 13 attempt to hold a predetermined part of the connection portion 22 of the selection recognition target portion 23 at the holding position. Then, when the holdability determination unit 125 determines that there is no obstacle in the moving path of the multiple fingers 13, it determines that the multiple fingers 13 can hold the selection recognition target portion 23.
[0086] After step s7, in step s8, the robot control unit 120 controls the arm 11 and the end effector 12 of the robot 10 so that the plurality of fingers 13 hold predetermined parts of the connection part 22 of the selected recognition target part 23 at the holding positions. Fig. 11 is a schematic diagram showing an example of how the plurality of fingers 13 hold predetermined parts of the connection part 22 of the selected recognition target part 23.
[0087] If the holdability determination unit 125 determines that an obstacle exists in the travel path of the multiple finger units 13 and that the multiple finger units 13 cannot hold the selected recognition target portion 23, step s4 is executed again, for example. In this re-execution of step s4, the selection unit 122 selects the recognition target portion 23 with the next highest holdability score. Thereafter, the robot system 50 operates in the same manner.
[0088] When the plurality of fingers 13 hold the selection recognition target portion 23, the robot control unit 120 controls the arm 11 so that the selection recognition target portion 23 held by the plurality of fingers 13 (also referred to as the held selection recognition target portion 23) is positioned above the first tray 30, as shown in FIG. 12 . Then, as shown in FIG. 13 , the robot control unit 120 controls the arm 11 to set the held selection recognition target portion 23 to a specific posture (step s9). Hereinafter, the selection nozzle 20 having the held selection recognition target portion 23 may be referred to as the held selection nozzle 20. In this case, setting the held selection recognition target portion 23 to a specific posture can also be said to set the held selection nozzle 20 to a specific posture.
[0089] Note that the setting of the selected nozzle 20 to a specific posture may be performed, for example, by calculating a motion path of the robot 10 from the posture of the selected nozzle 20 before holding to the specific posture of the selected nozzle 20 based on the posture of the selected nozzle 20 before holding identified by the posture identification unit 123, and then having the robot control unit 120 move the robot 10 along the motion path. In this case, the specific posture of the selected nozzle 20 can be set regardless of the gripping posture of the robot 10. Alternatively, the relative gripping posture of the robot 10 with respect to the selected nozzle 20 when holding the selected nozzle 20 can be kept constant, and the robot 10 can be set to a predetermined posture, thereby setting the selected nozzle 20 to a specific posture. In this case, the specific posture of the selected nozzle 20 can be set solely from the posture of the robot 10, regardless of the holding posture of the selected nozzle 20 by the robot 10.
[0090] FIG. 14 is a schematic diagram showing an example of the holding selection recognition target portion 23 set in a specific posture, the end effector 12, the cameras 15 and 16, and the first tray 30 viewed horizontally.
[0091] When the holding selection recognition target portion 23 is set to a specific posture, the optical axis direction of the effector camera 15 (the left-right direction on the paper in FIG. 14 ) is parallel to a direction perpendicular to the optical axis direction of the fixed camera 16 (the front-back direction on the paper in FIG. 14 ). The front-back direction on the paper in FIG. 14 can also be said to be the direction along the direction from the back side of the paper to the front side of the paper in FIG. 14 .
[0092] Here, in the workspace of the robot 10 (robot workspace), the direction corresponding to the front-to-back direction on the page of Fig. 14 is referred to as the x direction, and the direction corresponding to the left-to-right direction on the page of Fig. 14 is referred to as the y direction. The x direction and y direction can be said to be directions along a horizontal plane, or directions along the opening surface of the first tray 30. In other words, the origin of the coordinate system is somewhere on the top surface of the pedestal 18, and the depth direction of the pedestal 18 is the x direction, and the width direction of the pedestal 18 is the y direction.
[0093] The fixed camera 16 takes images along the x direction regardless of whether the holding selection recognition target portion 23 is set to a specific posture. On the other hand, the effector camera 15, which changes depending on the posture of the end effector 12, takes images along the y direction when the holding selection recognition target portion 23 is set to a specific posture. When the holding selection recognition target portion 23 is set to a specific posture, the multiple fingers 13 extend along the y direction, which can also be said to extend along a horizontal plane. When the holding selection recognition target portion 23 is set to a specific posture, the two fingers 13 of the end effector 12 are lined up side by side along a horizontal plane. Hereinafter, a situation in which the holding selection recognition target portion 23 is set to a specific posture may be referred to as a specific situation.
[0094] In certain situations, the hose 24 of the holding selection nozzle 20 extends generally along a direction perpendicular to the x and y directions (also referred to as the z direction, which is the up-down direction on the paper surface of FIG. 14 ). In certain situations, the hose 24 of the holding selection nozzle 20 extends generally along the vertical direction. In certain situations, the fixed camera 16 can photograph the side of the hose 24 of the holding selection nozzle 20 along the x direction, and the effector camera 15 can photograph the side of the hose 24 of the holding selection nozzle 20 along the y direction. In certain situations, the effector camera 15 and the fixed camera 16 can photograph the side of the hose 24 of the holding selection nozzle 20 from different directions.
[0095] FIG. 15 is a schematic diagram showing an example of a color image 161 (also referred to as a second color image 161b) generated by the fixed camera 16 in a specific situation. FIG. 16 is a schematic diagram showing an example of a color image 151 (also referred to as a second color image 151b) generated by the effector camera 15 in a specific situation. The second color image 161b shows a side view of the hose 24 of the holding selection nozzle 20 in the x direction in the specific situation. The second color image 151b shows a side view of the hose 24 of the holding selection nozzle 20 in the y direction in the specific situation. For reference, in FIG. 15, the y direction of the robot workspace shown in the second color image 161b is indicated by an arrow. Similarly, in FIG. 16, the x direction of the robot workspace shown in the second color image 151b is indicated by an arrow.
[0096] In this way, in step S9, the fixed camera 16 photographs the side of the hose 24 of the holding selection nozzle 20 along the x direction, and the arm 11 is controlled by the robot control unit 120 so that the holding selection recognition target portion 23 is in a position that allows the effector camera 15 to photograph the side of the hose 24 of the holding selection nozzle 20 along the y direction.
[0097] When the holding selection recognition target portion 23 is set to a specific posture, in step s10, the tip position specifying unit 126 specifies the position of the tip of the hose 24 of the holding selection nozzle 20 based on the second color image 161a and the second color image 161b. For example, the tip position specifying unit 126 specifies the position of the tip of the hose 24 of the holding selection nozzle 20 in a camera coordinate system set in the effector camera 15 based on the second color image 161a and the second color image 161b. Hereinafter, the term "camera coordinate system" simply refers to the camera coordinate system set in the effector camera 15. The posture of the camera coordinate system changes depending on the posture of the effector camera 15 and changes depending on the posture of the end effector 12.
[0098] Here, the position of the tip of the hose 24 of the hold selection nozzle 20 when it is assumed that the hose 24 extends straight from the connection part 22 of the hold selection nozzle 20 in a specific situation is referred to as the reference tip position. It can also be said that the reference tip position is the position of the tip of the hose 24 when it is assumed that the hose 24 of the hold selection nozzle 20 extends straight in the direction in which the connection part 22 of the hold selection nozzle 20 extends in a specific situation.
[0099] The tip position identifying unit 126 calculates the amount of deviation in the y direction (also referred to as the y-direction deviation) of the position of the tip of the hose 24 of the selected holding nozzle 20 relative to a reference tip position in the camera coordinate system, based on the second color image 161b obtained by the fixed camera 16 capturing images along the x direction. At this time, the tip position identifying unit 126 binarizes the second color image 161b using a predetermined threshold value, for example, to generate a binary image 261 that shows only the hose 24. FIG. 17 is a schematic diagram showing an example of the binary image 261 obtained by binarizing the second color image 161b shown in FIG. 15 using a predetermined threshold value. The white area in the binary image 261 is a hose area 262 that represents the hose 24. The tip position identifying unit 126 identifies the position of a point 263 representing the tip of the hose 24 in the binary image 261, and calculates the amount of deviation in the y direction based on the identified position.
[0100] Furthermore, the tip position identifying unit 126 determines the amount of deviation in the x direction (also referred to as the x-direction deviation) of the position of the tip of the hose 24 of the selected holding nozzle 20 relative to a reference tip position in the camera coordinate system, based on the second color image 151b obtained by the effector camera 15 capturing images along the y direction. At this time, the tip position identifying unit 126 binarizes the second color image 151b using a predetermined threshold value, for example, to generate a binary image 251 that shows only the hose 24. FIG. 18 is a schematic diagram showing an example of the binary image 251 obtained by binarizing the second color image 151b shown in FIG. 16 using a predetermined threshold value. The white area in the binary image 251 is a hose area 252 that represents the hose 24. The tip position identifying unit 126 identifies the position of a point 253 representing the tip of the hose 24 in the binary image 251, and determines the amount of deviation in the x direction based on the identified position.
[0101] The tip position specifying unit 126 specifies the position of the tip of the hose 24 of the holding selection nozzle 20 in the camera coordinate system based on the calculated x-direction deviation amount and y-direction deviation amount. It can be said that the x-direction deviation amount and y-direction deviation amount represent the degree of bending of the hose 24. It can also be said that the tip position specifying unit 126 specifies the position of the tip of the hose 24 of the holding selection nozzle 20 in the camera coordinate system based on the determined degree of bending of the hose 24.
[0102] When the multiple fingers 13 of the end effector 12 hold the selection nozzle 20, the relative positional relationship between the effector camera 15 fixed to the end effector 12 and the tip of the hose 24 of the holding selection nozzle 20 does not change. Therefore, by specifying the position of the tip of the hose 24 of the holding selection nozzle 20 in the camera coordinate system, it can be said that the relative positional relationship between the effector camera 15 and the tip of the hose 24 of the holding selection nozzle 20 is specified. Furthermore, by specifying the position of the tip of the hose 24 of the holding selection nozzle 20 in the camera coordinate system, it can be said that the relative positional relationship between the end effector 12 and the tip of the hose 24 of the holding selection nozzle 20 is specified, and it can also be said that the relative positional relationship between the fingers 13 and the tip of the hose 24 of the holding selection nozzle 20 is specified.
[0103] The display unit of the processing device 1 may display the second color image 161b shown in Fig. 15 or the second color image 151b shown in Fig. 16. The display unit of the processing device 1 may also display the binary image 261 shown in Fig. 17. In this case, a circle may be displayed in the binary image 261 surrounding a portion 263 representing the tip of the hose 24. The display unit of the processing device 1 may also display the binary image 251 shown in Fig. 18. In this case, a circle may be displayed in the binary image 251 surrounding a portion 253 representing the tip of the hose 24.
[0104] After step s10, in step s11, the robot control unit 120 controls the robot 10 so that the hose 24 is inserted from its tip into the container body 25 in the second tray 35 based on the position in the camera coordinate system of the tip of the hose 24 of the holding selection nozzle 20 identified by the tip position identification unit 126.
[0105] In this example, the robot control unit 120 knows in advance the position of the opening 25a of each container body 25 in the second tray 35. As shown in Fig. 19 , the robot control unit 120 controls the arm 11 so that the tip of the hose 24 of the holding selection nozzle 20 moves to directly above the opening 25a of one container body 25 in the second tray while maintaining the holding selection recognition target portion 23 in a specific posture. Because the robot control unit 120 knows the position in the camera coordinate system of the tip of the hose 24 of the holding selection nozzle 20 and the position of the opening 25a of the container body 25, it can control the arm 11 so that the tip of the hose 24 of the holding selection nozzle 20 moves to directly above the opening 25a of the container body 25. Next, the robot control unit 120 controls the arm 11 so that the end effector 12 moves straight down, and as shown in Figure 20, inserts the hose 24 of the holding selection nozzle 20 from its tip into the opening 25a of the container body 25, and inserts the hose 24 into the container body 25.
[0106] When the hose 24 of the held selection nozzle 20 has entered the container body 25 to some extent, the robot control unit 120 controls the end effector 12 to open the multiple fingers 13 in step s12, thereby releasing the fingers 13 from holding the selection nozzle 20. Thereafter, the robot control unit 120 controls the arm 11 to move the multiple fingers 13 in the y direction away from the selection nozzle 20, as shown in FIG. 21 . Thereafter, step s1 is executed again, and thereafter the robot system 50 operates in the same manner. As a result, the multiple nozzles 20 in the first tray 30 are respectively combined with the multiple container bodies 25 in the second tray 35.
[0107] The timing for opening the multiple fingers 13 may be set in advance based on, for example, the amount of movement of the arm 11 in the z direction. Specifically, the multiple fingers 13 may be set to open after the arm 11 has moved a predetermined amount in the z direction from the point at which the insertion operation of the hose 24 of the holding selection nozzle 20 into the container body 25 begins. In this case, the z-direction coordinate of the insertion operation start point may be set to a constant value, or may be set based on the height of the container body 25 acquired based on an image from the effector camera 15 or the fixed camera 16. The timing for opening the multiple fingers 13 may also be set based on the length of the hose 24. In this case, for example, the robot 10 may be controlled to open the multiple fingers 13 after moving the arm in the z direction by a predetermined percentage of the length of the hose 24 based on the length of the hose 24 acquired in advance or the length of the hose 24 calculated based on an image from the effector camera 15 or the fixed camera 16.
[0108] As described above, the recognition unit 121 (in other words, the trained model 121) recognizes only a portion of the work object, rather than the entire work object, of the robot 10. This allows for the work efficiency to be improved by appropriately setting only a portion of the work object as the recognition target. In the above example, the recognition unit 121 does not recognize the entire nozzle 20 as the work object, but rather the hard portion 23 that constitutes part of the nozzle 20. This allows the robot 10 to reliably perform the holding operation by first holding the nozzle 20, which is easier to hold. This allows for the work efficiency of the robot 10 to be improved. This point will be described in detail below.
[0109] For example, consider a case where, unlike this example, the recognition unit 121 recognizes the entire nozzle 20 and outputs a visibility score (also referred to as an overall visibility score) indicating the proportion of the portion of the nozzle 20 that appears in the color image 151 relative to the overall image of the nozzle 20. In this case, the overall visibility score may be small even if no other object overlaps the hard portion 23 of the nozzle 20 that the robot 10 is holding. For example, if another object (e.g., the head 21 of another nozzle 20) overlaps a hose 24 that the robot 10 is not holding, the overall visibility score may be small. Therefore, even if the hard portion 23 of a certain nozzle 20 is easily held by the robot 10, the certain nozzle 20 may not be selected as the nozzle 20 to be held by the robot 10 in step s4.
[0110] In contrast, in this example, when the hard portion 23 that constitutes part of the nozzle 20 and is held by the robot 10 is set as the recognition target, even if another object overlaps the hose 24 that the robot 10 is not holding, this does not cause the visibility score output by the recognition unit 121 to decrease. As a result, in step s4, the nozzles 20 that are easy for the robot 10 to hold are more likely to be selected as the nozzles 20 to be held by the robot 10. Therefore, the robot 10 can reliably perform the holding task starting with the nozzles 20 that are easy to hold. As a result, the work efficiency of the robot 10 is improved.
[0111] Furthermore, if another object is overlapping the relatively hard head 21, the robot 10 may have difficulty holding the connection portion 22 and lifting the nozzle 20, which may result in the robot 10 failing to lift the nozzle 20. In this example, since not only the connection portion 22 held by the robot 10 but also the head 21 is targeted for recognition, the visibility score decreases when another object is overlapping the head 21. As a result, a nozzle 20 with another object overlapping the head 21 is less likely to be selected as the nozzle 20 to be held by the robot 10 in step s4. This makes it less likely that the robot 10 will fail to lift the nozzle 20, improving the work efficiency of the robot 10.
[0112] Furthermore, even if another object is placed on top of the hose 24 of the nozzle 20 being held, since the hose 24 is flexible, when the nozzle 20 is lifted by holding the connection part 22, the hose 24 will deform itself to release the force, so that it is unlikely that an external force will be applied to the nozzle 20 and the nozzle 20 will not be able to be held.
[0113] Furthermore, in the above example, the robot system 50 is described as being equipped with the camera 15 and the camera 16. However, the robot system 50 may be equipped with only one of the cameras 15 and 16 as long as the robot system 50 can perform the task. In this case, the orientation of the hose 24 may be estimated based on images obtained by two captures by the camera. That is, after the camera captures the first image, the camera may rotate at a right angle along the planar direction, and then the camera may capture the second image.
[0114] In the above example, the hard portion 23 constituting part of the nozzle 20 is the recognition target portion, but the hard portion 23 and the end of the hose 24 on the connection portion 22 side may also be the recognition target portions. In this case, the robot 10 may hold the end of the hose 24 on the connection portion 22 side.
[0115] Furthermore, in the above example, various scores are output from the recognition unit 121 as recognition results, but the various scores do not have to be output from the recognition unit 121. In this case, when outputting a mask, the recognition unit 121 may use the scores as a criterion for selecting a mask to be output.
[0116] In addition, in step s4, the selected nozzle 20 may be selected based on the proximity of the nozzle 20 to the current position of the end effector 12, or may be selected randomly from nozzles 20 with a holdability score above a certain level.
[0117] Furthermore, in the above example, the robot control unit 120 has previously determined the position of the opening 25 a of each container body 25 in the second tray 35, i.e., the position of the opening 25 a of each container body 25 has been set in advance. However, the control unit 2 may also recognize the position of each container body 25. In this case, for example, the effector camera 15 may photograph each container body 25 in the second tray 35, and the control unit 2 may recognize each container body 25 using machine learning or the like based on the image obtained by the effector camera 15, thereby estimating the position of the opening 25 a of each container body 25. In this case, the center of the recognized outer shape of each container body 25 may be set as the position of the opening 25 a of each container body 25, or the opening 25 a of each container body 25 itself may be recognized. Note that a learning method for recognition using machine learning may be used in conjunction with a learning method for recognition of a work object.
[0118] <Example of Learning Method> Next, a description will be given of an example of a method for learning a pre-learning model to generate the trained model 121. Hereinafter, the pre-learning model that is the basis of the trained model 121 will be referred to as the pre-learning model 221. The trained model 121 is generated by learning the pre-learning model 221.
[0119] 22 is a schematic diagram illustrating an example of a learning method for the pre-training model 221. The learning of the pre-training model 221 may be performed by the processing device 1, or may be performed by a device other than the processing device 1 (e.g., a cloud server). In other words, the generation of the trained model 121 may be performed by the processing device 1, or may be performed by a device other than the processing device 1. The device that learns the pre-training model 221 can also be called a learning device.
[0120] The learning device has a pre-learning model 221. The learning device also stores learning data 230 used in learning the pre-learning model 221. When the processing device 1 is a learning device, for example, the control unit 2 has the pre-learning model 221, and the memory unit 3 stores the learning data 230. When the control unit 2 executes the program 3a in the memory unit 3, a learning unit that learns the pre-learning model 221 and generates the trained model 121 is formed as a functional block in the control unit 2.
[0121] The pre-training model 221 is trained based on training data 230. The training data 230 includes a plurality of data sets each consisting of training images 231 and corresponding annotation data 232.
[0122] The learning image 231 may be, for example, a CG (Computer Graphics) image. A CG image may also be, for example, a composite image. Furthermore, the learning image 231 may depict the hard portion 23 of the nozzle 20 (in other words, the recognition target portion 23) and the hose 24 of the nozzle 20 (in other words, the portion of the nozzle 20 other than the recognition target portion 23) separately from each other. In other words, the learning image 231 may depict the hard portion 23 and the hose 24 as separate components.
[0123] 23 is a schematic diagram showing an example of a training image 231 (also referred to as a first training CG image 231a) in which the hard portions 23 and the hoses 24 are shown separately. The hard portions 23 and the hoses 24 shown in the first training CG image 231a are CG images, not actual objects. The first training CG image 231a shows a state in which a plurality of hard portions 23 and a plurality of hoses 24 are randomly stacked. The training data 230 includes a plurality of first training CG images 231a each showing a different number of hard portions 23 and hoses 24. The number of hard portions 23 shown in the first training CG image 231a and the number of hoses 24 shown in the first training CG image 231a may be the same as or different from each other.
[0124] In the first learning CG image 231a, a pattern that does not exist in the real world may be applied to the hard portion 23. In this case, the patterns may be different among the multiple hard portions 23 depicted in the first learning CG image 231a. In addition, in the first learning CG image 231a, a pattern that does not exist in the real world may be applied to the hose 24. In this case, the patterns may be different among the multiple hoses 24 depicted in the first learning CG image 231a.
[0125] The first learning CG image 231a is generated using, for example, 3DCG technology. 3DCG is an abbreviation for 3-Dimensional Computer Graphics. When the first learning CG image 231a is generated, 3D CAD data of the nozzle 20 is divided to generate 3D CAD data of the hard portion 23 and 3D CAD data of the hose 24. CAD is an abbreviation for Computer Aided Design. Then, the hard portion 23 is modeled using 3DCG technology based on the 3D CAD data of the hard portion 23, generating a 3DCG model of the hard portion 23 (also referred to as a hard portion model). Similarly, the hose 24 is modeled using 3DCG technology based on the 3D CAD data of the hose 24, generating a 3DCG model of the hose 24 (also referred to as a hose model). Next, multiple hard portion models and multiple hose models are randomly stacked in a virtual three-dimensional space. For example, by using simulation software, the plurality of hard part models and the plurality of hose models can be piled up randomly by allowing each hard part model and each hose model to fall naturally one by one in a virtual three-dimensional space. Then, a virtual image obtained by photographing the randomly piled hard part models and hose models from above with a virtual camera is set as the first learning CG image 231a. Note that the method for generating the first learning CG image 231a is not limited to this.
[0126] The annotation data 232 includes a mask, a classification score, a mask score, and a visibility score for each hard portion 23 (in other words, a hard portion model) appearing in the corresponding learning image 231. Since the mask included in the annotation data 232 is a correct mask of the hard portion 23 appearing in the learning image 231, the mask score included in the annotation data 232 is 1. The classification score included in the annotation data 232 is also 1. The visibility score included in the annotation data 232 indicates the proportion of the portion of the hard portion 23 that appears in the learning image 231 with respect to the entire image of the hard portion 23 (in other words, the hard portion model). The annotation data 232 includes a correct visibility score for the hard portion 23 appearing in the learning image 231. As described above, when multiple hard part models and multiple hose models are stacked randomly using simulation software, the positions and postures of the multiple hard part models and multiple hose models stacked randomly can be determined using the simulation software, making it possible to easily generate annotation data.
[0127] As described above, the trained model 121 is generated by training the pre-trained model 221 based on the training data 230 including the training image 231 and the annotation data 232 related to the recognition target portion 23. The trained model 121 trained based on the training data 230 including the training image 231 and the annotation data 232 related to the recognition target portion 23 recognizes the recognition target portion 23, thereby enabling the recognition target portion 23 to be properly recognized.
[0128] 23 , when the recognition target portion 23 and an area other than the recognition target portion 23 are shown separately in the learning image 231, it becomes easy to create various situations in which other objects overlap the recognition target portion 23 in the learning image 231. This makes it possible to easily improve the learning accuracy of the pre-learning model 221.
[0129] Furthermore, when the training images 231 are CG images, it is easy to generate a variety of training images 231. This makes it possible to easily improve the training accuracy of the pre-training model 221.
[0130] 24, the learning image 231 may be a CG image in which the hard portion 23 and the hose 24 are depicted in an inseparable state. In other words, the learning image 231 may be a CG image in which the hard portion model and the hose model are depicted in a combined state. Hereinafter, the CG learning image 231 in which the hard portion 23 and the hose 24 are depicted in an inseparable state, as in FIG. 24, will be referred to as a second learning CG image 231b.
[0131] 25, the learning image 231 may be a CG image showing only the hard portion 23. Hereinafter, the CG learning image 231 showing only the hard portion 23 as shown in FIG. 25 will be referred to as a third learning CG image 231c.
[0132] The learning data 230 may include at least two types of images from among a first learning CG image 231a, a second learning CG image 231b, and a third learning CG image 231c.
[0133] The learning data 230 may also include a learning image 231 (also referred to as a first learning actual image 231) that is not a CG image and that depicts the actual nozzle 20. The learning data 230 may also include a learning image 231 (also referred to as a second learning actual image 231) in which the actual nozzle 20 is decomposed into the hard portion 23 and the hose 24, and the actual hard portion 23 and the actual hose 24 are separately depicted. The learning data 230 may also include a learning image 231 (also referred to as a third learning actual image 231) in which only the actual hard portion 23 is depicted. The learning data 230 may also include a learning image 231 that depicts the hard portion 23 and a learning image 231 that depicts the hose 24 separately.
[0134] The learning data 230 may include at least two types of images: a first learning real image 231, a second learning real image 231, and a third learning real image 231. The learning data 230 may also include at least one type of image: a first learning CG image 231a, a second learning CG image 231b, and a third learning CG image 231c, and at least one type of image: a first learning real image 231, a second learning real image 231, and a third learning real image 231.
[0135] In the above example, the hard portion 23 of the nozzle 20 is set as the recognition target portion, but only the connection portion 22 of the nozzle 20 may be set as the recognition target portion. Also, only the head 21 of the nozzle 20 may be set as the recognition target portion. Also, only the hose 24 of the nozzle 20 may be set as the recognition target portion.
[0136] Also, a grayscale image may be used instead of the color image 151. Also, a grayscale image may be used instead of the color image 161.
[0137] Although the processing device and the robot system have been described in detail above, the above description is merely illustrative in all respects and does not limit the scope of this disclosure. Furthermore, the various examples described above can be combined and applied as long as they are not mutually inconsistent. It is understood that countless examples not illustrated can be envisioned without departing from the scope of this disclosure.
[0138] This disclosure includes the following:
[0139] In one embodiment, (1) a learning method generates a trained model that recognizes a recognition target portion, which is part of a work object, based on an inference image that shows multiple work objects that are the targets of work by a robot, by learning a pre-trained model based on training data that includes the training image and annotation data related to the recognition target portion.
[0140] (2) The learning method of (1) above, wherein the learning image shows the recognition target portion and the portion of the work object other than the recognition target portion, separated from each other.
[0141] (3) In the learning method of (1) or (2) above, the learning images are CG (Computer Graphics) images.
[0142] (4) A learning method according to any one of (1) to (3) above, wherein the annotation data includes a visibility score indicating the proportion of the portion of the recognition target part that appears in the learning image relative to the overall image of the recognition target part.
[0143] (5) The learning device executes any one of the learning methods (1) to (4) above.
[0144] (6) The processing device includes a recognition unit that recognizes a recognition target portion, which is part of a work object, based on an image showing the work object that is the target of work by the robot.
[0145] (7) In the processing device of (6) above, the recognition result of the recognition unit includes a visibility score indicating the proportion of the part of the recognition target part that appears in the image to the entire image of the recognition target part.
[0146] (8) The processing device of (7) above, further comprising a selection unit that selects a work object from the plurality of work objects for the robot to perform work on, based on the recognition result including the visibility score.
[0147] In one embodiment, the program (9) is a program for causing a computer device to execute any one of the learning methods (1) to (4) above.
[0148] In one embodiment, the program (10) is a program for causing a computer device to function as any one of the processing devices (6) to (8) above.
[0149] REFERENCE SIGNS LIST 1 Processing device (learning device) 3a Program 10 Robot 20 Nozzle 23 Hard part (recognition target part) 24 Hose 121 Recognition unit (trained model) 122 Selection unit 221 Pre-learning model 230 Learning data 231 Learning image 231a First learning CG image 231b Second learning CG image 231c Third learning CG image 232 Annotation data
Claims
1. A learning method that generates a trained model that recognizes a recognition target portion, which is part of a work object, based on an inference image that shows multiple work objects that are the target of robot work, by learning a pre-trained model based on training data that includes the training image and annotation data related to the recognition target portion.
2. A learning method according to claim 1, wherein the learning image shows the portion to be recognized and the portion of the work object other than the portion to be recognized, separated from each other.
3. A learning method according to claim 1 or 2, wherein the learning images are CG (Computer Graphics) images.
4. A learning method according to any one of claims 1 to 3, wherein the annotation data includes a visibility score indicating the proportion of the part of the recognition target portion that appears in the learning image relative to the entire image of the recognition target portion.
5. A learning device that executes the learning method according to any one of claims 1 to 4.
6. A processing device having a recognition unit that recognizes a recognition target portion that is part of a work object based on an image showing multiple work objects that are the target of work by a robot.
7. A processing device according to claim 6, wherein the recognition result of the recognition unit includes a visibility score indicating the proportion of the part of the recognition target that appears in the image relative to the entire image of the recognition target.
8. A processing device according to claim 7, comprising a selection unit that selects a work object for causing the robot to perform a task from the plurality of work objects based on the recognition result including the visibility score.
9. A program for causing a computer device to execute the learning method according to any one of claims 1 to 4.
10. A program for causing a computer device to function as a processing device according to any one of claims 6 to 8.
Citation Information
Patent Citations
Recognition method, recognition system, robot control method, robot control system, robot system, recognition program, and robot control program
JP2020107142A
Object measurement method, measuring device, program, and computer-readable recording medium
JP2020197983A
Target object recognition device, manipulator, and mobile robot
WO2020049766A1
Data set generation device, method, program, and system
WO2022097353A1
Recognition model generation method and recognition model generation device
WO2023286847A1