Processing device, robot system, and program
The robot system addresses the challenge of handling and aligning rod-shaped objects by using a processing device to estimate and correct their longitudinal direction, enabling precise manipulation and alignment into storage sections, thereby enhancing operational efficiency.
Patent Information
- Application Number
- PCT/JP2025/027367
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-02
- Filing Date
- 2025-08-01
- Publication Date
- 2026-02-05
AI Technical Summary
Existing robotic systems face challenges in efficiently handling and aligning rod-shaped objects with varying lengths and end surface configurations, particularly when they are randomly stacked, due to difficulties in accurately detecting and orienting these objects for precise manipulation and alignment.
A robot system equipped with a processing device that estimates the longitudinal direction of rod-shaped objects using image capture, detects deviations, and controls a robotic arm to align and manipulate these objects into storage sections, utilizing suction mechanisms and cameras for precise positioning and orientation, regardless of object type.
The system enhances the efficiency of handling and aligning rod-shaped objects by ensuring accurate orientation and placement, improving the overall operational efficiency of subsequent processes.
Smart Images

Figure JP2025027367_05022026_PF_FP_ABST
Abstract
Description
Processing device, robot system and program
[0001] The present disclosure relates to robots.
[0002] Patent Document 1 describes a technology related to a robot.
[0003] JP 2010-120414 A
[0004] A processing device, a robot system, and a program are disclosed. In one embodiment, the processing device includes a processing unit. The processing unit estimates the longitudinal direction of a rod-shaped object that is a work target of the robot, and detects a deviation of the estimated longitudinal direction based on a captured image of the object.
[0005] In one embodiment, a robot system includes the processing device described above and a robot controlled by the processing device.
[0006] In one embodiment, the program is a program for causing a computer device to function as a processing unit of the processing device.
[0007] FIG. 1 is a schematic diagram showing an example of a robot system. FIG. 2 is a schematic diagram showing an example of an object. FIG. 3 is a schematic diagram showing an example of an alignment container and an alignment tray. FIG. 4 is a schematic diagram showing an example of the operation of the robot system. FIG. 5 is a schematic diagram showing an example of the operation of the robot system. FIG. 6 is a schematic diagram showing an example of an alignment container. FIG. 7 is a schematic diagram showing an example of an alignment container. FIG. 8 is a schematic diagram showing an example of the operation of the robot system. FIG. 9 is a schematic diagram showing an example of the operation of the robot system. FIG. 10 is a schematic diagram showing an example of the configuration of a processing device. FIG. 11 is a flowchart showing an example of the operation of the robot system. FIG. 12 is a schematic diagram showing an example of an object mask. FIG. 13 is a schematic diagram showing an example of a selected object point cloud. FIG. 14 is a schematic diagram for explaining an example of the operation of the processing device. FIG. 15 is a schematic diagram showing an example of a selected object point cloud and a parallel plane peripheral point cloud. FIG. 16 is a schematic diagram for explaining an example of a parallel plane peripheral point cloud. FIG. 17 is a schematic diagram for explaining an example of the operation of the processing device. FIG. 18 is a schematic diagram for explaining an example of the operation of the processing device. FIG. 19 is a schematic diagram for explaining an example of the operation of the robot system. FIG. 20 is a schematic diagram showing an example of a color image. Fig. 21 is a schematic diagram showing an example of a color image, and Fig. 22 is a schematic diagram for explaining an example of the operation of the processing device.
[0008] Fig. 1 is a schematic diagram showing an example of a robot system 80. As shown in Fig. 1, the robot system 80 includes, for example, a robot 10 and a processing device 1. The processing device 1 is capable of controlling, for example, the robot 10. The processing device 1 in this example functions as a control device that controls the robot 10.
[0009] The robot 10 is capable of, for example, moving an object from a first location to a second location different from the first location under the control of the processing device 1. The robot 10 is, for example, an arm-type robot and includes an arm 11 and an end effector 12 connected to the arm 11. The processing device 1 is capable of controlling the arm 11 and the end effector 12.
[0010] The end effector 12 is capable of holding an object 20 (also referred to as a work object 20) that is the work target of the robot 10 under the control of the processing device 1. In this case, the end effector 12 can also be called a holding mechanism.
[0011] The end effector 12 is capable of, for example, suction-holding the target object 20. In this case, the end effector 12 can also be called a suction mechanism. The end effector 12 includes, for example, a long and thin suction nozzle 12a and a suction part 12b attached to the tip of the suction nozzle 12a. The suction part 12b is capable of suction-holding the target object 20. The suction part 12b is made of an elastic material such as synthetic rubber and is elastically deformable. The suction part 12b is also called a suction pad.
[0012] The suction portion 12b has a suction opening 12bb at its tip. The outer shape of the suction opening 12bb is, for example, circular. When the suction portion 12b suctions the target object 20, the edge of the suction opening 12bb abuts against the target object 20, and the suction opening 12bb is blocked by the target object 20. Then, by reducing the pressure inside the suction portion 12b, the edge of the suction opening 12bb comes into close contact with the target object 20. As a result, the target object 20 is sucked onto the suction portion 12b. Note that the end effector 12 may be a gripping mechanism that grips the target object 20 with multiple fingers.
[0013] The arm 11 includes, for example, a plurality of joints. The posture of the arm 11 changes as the amount of rotation of at least one of the plurality of joints changes. The processing device 1 can change the posture of the arm 11. The change in posture of the arm 11 changes the position and posture of the end effector 12. The change in posture of the arm 11 also changes the position and posture of the object 20 held by the end effector 12.
[0014] In this example, a plurality of objects 20 are randomly stacked in a tray 30. The end effector 12, for example, sucks and holds the objects 20 in the tray 30 one by one with the suction portion 12b. Then, the robot 10 changes the posture of the arm 11 to move the held object 20 from the tray 30 to another location.
[0015] The object 20 is, for example, a rod-shaped member. The object 20 is, for example, an insulator used to insulate electric wires. The object 20 is, for example, cylindrical. Furthermore, the object 20 is, for example, hollow cylindrical. Openings are provided on both end surfaces in the longitudinal direction (also referred to as the length direction) of the object 20. For example, the end surfaces of the object 20 are flat. For example, the side surfaces of the object 20 are curved. A plurality of objects 20 are, for example, laid horizontally and stacked in bulk in a tray 30. The end effector 12 holds, for example, the side surfaces 25 (also referred to as peripheral side surfaces) of the objects 20 in the tray 30. Note that the shape of the object 20 does not have to be cylindrical, and may be a polygonal prism, a cone, a polygonal pyramid, or the like.
[0016] FIG. 2 is a schematic diagram showing an example of both end faces of an object 20. The left side of FIG. 2 shows one longitudinal end face 21 (also referred to as a first end face 21) of the object 20. The right side of FIG. 2 shows an example of the other longitudinal end face 22 (also referred to as a second end face 22) of the object 20. The outer shapes of the first end face 21 and the second end face 22 are, for example, circular and are the same. Note that at least one of the shapes and sizes of the first end face 21 and the second end face 22 may be different from each other. For example, the shapes of the first end face 21 and the second end face 22 may be similar.
[0017] The second end face 22 is provided with an opening 22a (also referred to as the second opening 22a) that is, for example, generally minus-shaped. In contrast, the first end face 21 is provided with an opening 21a (also referred to as the first opening 21a) that is, for example, generally plus-shaped. The first opening 21a and the second opening 22a are connected to each other. The appearance of the first end face 21 and the appearance of the second end face 22 are different from each other. As shown on the left side of FIG. 2 , when the object 20 is viewed from the first end face 21 side, not only the first opening 21a but also the second opening 22a provided in the second end face 22 is visible.
[0018] In the example of FIG. 2 , the appearances of the first end face 21 and the second end face 22 differ from each other due to the difference in opening shape. However, the appearances of the first end face 21 and the second end face 22 may be different from each other by other means. For example, the appearances of the first end face 21 and the second end face 22 may be different from each other by applying different colors to the first end face 21 and the second end face 22. The appearances of the first end face 21 and the second end face 22 may be different from each other by applying different patterns to the first end face 21 and the second end face 22. The appearances of the first end face 21 and the second end face 22 may be different from each other by applying different outer shapes to the first end face 21 and the second end face 22.
[0019] In this example, there are multiple types of objects 20 that differ only in length. The multiple types of objects 20 include, for example, a first object 20, a second object 20 that is longer than the first object 20, and a third object 20 that is longer than the second object 20.
[0020] Only objects 20 of the same type (i.e., objects 20 of the same length) are bulk loaded in one tray 30. In this example, a tray 30 (also referred to as a first tray 30) on which a plurality of first objects 20 are bulk loaded, a tray 30 (also referred to as a second tray 30) on which a plurality of second objects 20 are bulk loaded, and a tray 30 (also referred to as a third tray 30) on which a plurality of third objects 20 are bulk loaded are prepared.
[0021] When a first tray 30 is placed within the working range of the robot 10, the robot 10 holds and moves the first object 20 from the first tray 30. When a second tray 30 is placed within the working range of the robot 10, the robot 10 holds and moves the second object 20 from the second tray 30. When a third tray 30 is placed within the working range of the robot 10, the robot 10 holds and moves the third object 20 from the third tray 30. The robot 10 performs the same operation under the control of the processing device 1 regardless of the type of object 20 it holds. In other words, the processing device 1 causes the robot 10 to operate in the same way regardless of the type of object 20 it holds.
[0022] The object 20 is not limited to the above examples. For example, the object 20 may be an industrial product such as an electronic device or a screw, a food product such as a vegetable or bread, or a household item such as a diaper, a toothbrush, or a cup. Furthermore, the object 20 may be any object that can be adsorbed and whose tilt can be detected.
[0023] The robot system 80 includes an alignment container 40 and an alignment tray 50. The robot 10 holds the objects 20 in the tray 30, moves the held objects 20 to the alignment container 40, and places them in the alignment container 40. The alignment container 40 can, for example, align and store N objects 20 (N is an integer greater than or equal to 2) so that the N objects 20 are lined up in a row. The alignment container 40 has N storage sections 41 that can store the objects 20. The N storage sections 41 are, for example, lined up along a horizontal plane. The robot 10 moves the objects 20 in the tray 30 one by one to the alignment container 40 and places them in the storage sections 41. The objects 20 are inserted and placed in the storage sections 41 from the upper side of the storage section 41. The robot 10 may, for example, insert and place the objects 20 into the N storage sections 41 in order from the end. In this example, N=9, and the alignment container 40 has nine storage sections 41. However, the value of N is not limited to this.
[0024] The alignment container 40 is movable. Furthermore, the bottom of each storage section 41 of the alignment container 40 can be opened and closed. When the bottom of the storage section 41 containing the objects 20 opens, the objects 20 fall out of the storage section 41. The movement of the alignment container 40 and the opening and closing of the bottom of the storage section 41 are performed by, for example, a drive mechanism including an actuator. The processing device 1 can control the movement of the alignment container 40 and the opening and closing of the bottom of the storage section 41 by, for example, controlling the drive mechanism. Note that the movement of the alignment container 40 and the opening and closing of the bottom of the storage section 41 may be controlled by a device separate from the processing device 1.
[0025] The objects 20 are placed in the arranging container 40, and then placed from the arranging container 40 onto the arranging tray 50. FIG. 3 is a schematic diagram showing an example of the arranging container 40 and the arranging tray 50.
[0026] The sorting tray 50 can, for example, align and store the plurality of objects 20 so that the objects 20 are lined up in a matrix. The sorting tray 50 includes, for example, a plate-shaped insert member 51 into which the objects 20 are inserted, and a plate-shaped support member 55 that supports the objects 20 inserted into the insert member 51 from below. The insert member 51 is provided with a plurality of through holes 52. One object 20 is inserted into one through hole 52.
[0027] The multiple through holes 52 are arranged, for example, in a matrix along a horizontal plane. Here, the direction along which the multiple storage sections 41 of the alignment container 40 are arranged is called the column direction, and the direction perpendicular to the column direction and along the horizontal direction is called the row direction. In the example of FIG. 3 , the multiple through holes 52 are arranged, for example, N in number in the column direction and M in number in the row direction. Hereinafter, the N through holes 52 arranged in the column direction will be collectively referred to as a through hole row. The multiple through holes 52 are configured as M through hole rows. In the example of FIG. 3 , M=8, but the value of M is not limited to this.
[0028] 4, when the N objects 20 are aligned in the alignment container 40, the alignment container 40 moves to above a row of through-holes in the alignment tray 50 into which no objects 20 have yet been inserted. The bottoms of the N storage sections 41 of the alignment container 40 then open, and the N objects 20 fall out of the N storage sections 41. The bottoms of the storage sections 41 into which the objects 20 have fallen then close.
[0029] When the N objects 20 drop from the sorting container 40, the N objects 20 are dropped and inserted into the N through-holes 52 that make up the row of through-holes located below the N storage sections 41, as shown in Fig. 5. As a result, the N objects 20 are collectively arranged in the sorting tray 50.
[0030] The sorting container 40 then moves toward the tray 30 and returns to its original position. When the robot 10 then places N objects 20 in the empty sorting container 40, the robot system 80 operates in the same manner, and the N objects 20 in the sorting container 40 are inserted into the N through-holes 52 that make up the through-hole row of the sorting tray 50, respectively. Thereafter, the robot system 80 operates in the same manner. The robot system 80 may also photograph the sorting container 40 and determine whether a predetermined number of objects 20 have been placed in the sorting container 40 based on the photographed image. In this case, for example, the processing device 1 may count the number of objects 20 in the sorting container 40 from the photographed image. The processing device 1 may also determine whether a predetermined number of objects 20 have been placed in the sorting container 40 using a trained model that has been machine-learned to learn the state of the sorting container 40 in which the predetermined number of objects 20 have been placed. Various sensors may be arranged in the sorting container 40, and the processing device 1 may determine, based on the detection results of the sensors, that a predetermined number of objects 20 have been arranged in the sorting container 40. The processing device 1 may also determine that a predetermined number of objects 20 have been arranged in the sorting container 40 when the robot 10 holds the predetermined number of objects 20 or when the robot 10 arranges the predetermined number of objects 20 in the sorting container 40.
[0031] The robot 10 may drop and insert the N objects 20 contained in the sorting container 40 into the N rows of through holes of the sorting tray 50 in order from the end in the row direction. For example, the robot 10 may drop and insert the N objects 20 contained in the sorting container 40 into the N rows of through holes of the sorting tray 50 in order from the farthest from the robot 10.
[0032] 6 and 7 are schematic diagrams showing an example of an alignment container 40. The inner surface of the storage section 41 of the alignment container 40 has, for example, a shape roughly like a funnel split in half except for the tip half of the tapered section. The object 20 is inserted into the storage section 41 from the upper opening of the storage section 41 and placed in the storage section 41. The depth direction of the storage section 41 is, for example, along the vertical direction.
[0033] The housing portion 41 has a first partial housing portion 42, a second partial housing portion 43, and a third partial housing portion 44. The first partial housing portion 42, the second partial housing portion 43, and the third partial housing portion 44 are arranged in this order from the top of the housing portion 41. The top and bottom ends of the first partial housing portion 42 are open, and the top and bottom ends of the second partial housing portion 43 are open. The top end of the third partial housing portion 44 is open, and the bottom end of the third partial housing portion 44 is openable. The bottom end of the third partial housing portion 44, i.e., the bottom of the third partial housing portion 44, constitutes the openable bottom 45 of the housing portion 41 (see FIG. 7 ). The internal space of the second partial housing portion 43 is in communication with the internal space of the first partial housing portion 42 and the internal space of the third partial housing portion 44.
[0034] The inner side surface of the first partial accommodating portion 42 has a shape roughly obtained by halving the side surface of a truncated cone. The inner side surface of the second partial accommodating portion 43 has a shape roughly obtained by halving the side surface of a cylinder. The inner side surface of the third partial accommodating portion 44 has the same shape as the side surface of a cylinder. It can also be said that the inner side surface of the partial accommodating portion consisting of the second partial accommodating portion 43 and the third partial accommodating portion 44 has a shape roughly obtained by halving the upper half of the side surface of a cylinder.
[0035] The object 20 is inserted into the storage section 41 from the first partial storage section 42 side and placed within the storage section 41. The object 20 is inserted into the storage section 41 from the upper end opening of the first partial storage section 42 (i.e., the opening surrounded by the upper edge of the inner side surface of the first partial storage section 42) and placed in the first partial storage section 42, the second partial storage section 43, and the third partial storage section 44. When the object 20 is stored in the storage section 41, the first partial storage section 42 and the second partial storage section 43 cover approximately half of the circumference of the side surface 25 of the cylindrical object 20, and the third partial storage section 44 surrounds the entire circumference of the side surface 25 of the object 20. The posture of the object 20 within the storage section 41 is stabilized by the object 20 being surrounded by the third partial storage section 44 in the circumferential direction.
[0036] The first partial accommodating portion 42 has a tapered inner side surface whose diameter gradually decreases from the insertion side of the object 20. The diameter of the inner side surface of the first partial accommodating portion 42 gradually decreases from the insertion side of the object 20 toward the second partial accommodating portion 43. In other words, the diameter of the inner side surface of the first partial accommodating portion 42 gradually decreases from the upper end opening of the accommodating portion 41 toward the bottom 45.
[0037] After suction-holding the side surface 25 of the object 20 in the tray 30 with the suction unit 12b, the robot 10 moves the object 20 above the empty storage section 41 as shown in FIG. 8 . At this time, the robot 10 positions the object 20 directly above the empty storage section 41 so that the longitudinal direction of the object 20 is generally along the vertical direction. As a result, the longitudinal direction of the object 20 and the depth direction of the empty storage section 41 are generally aligned on the same straight line. The robot 10 then moves the object 20 held by the suction unit 12b directly downward, and inserts the object 20 into the empty storage section 41 through the upper opening thereof as shown in FIG. 9 . At this time, the robot 10 moves the held object 20 directly downward so that the object 20 is inserted into the third partial storage section 44. Thereafter, the robot 10 causes the suction unit 12b to release the suction-holding of the object 20. This completes the storage of the object 20 in the storage section 41.
[0038] When the storage of the object 20 in the storage section 41 is completed, the robot 10 picks up a new object 20 from the tray 30 and similarly inserts and places the held object 20 into an empty storage section 41 of the sorting container 40.
[0039] Under the control of the processing device 1, the robot 10 aligns the orientation of the N objects 20 and places the N objects 20 in the sorting container 40. This improves the efficiency of subsequent processes. As shown in FIG. 4 , the robot 10 places the N objects 20 in the sorting container 40 so that the first end surface 21 of each object 20 is located on the upper side. The robot 10 inserts the objects 20 into the receptacle 41 of the sorting container 40 from the second end surface 22 side. The robot 10 may also place the N objects 20 in the sorting container 40 so that the second end surface 22 of each object 20 is located on the upper side.
[0040] In this example, because the inner side surfaces of the first partial storage section 42 and the second partial storage section 43 are partially open, when the end effector 12 holds the side surface 25 of the object 20 and inserts and places the object 20 into the storage section 41, the end effector 12 is less likely to interfere with the storage section 41, as shown in Fig. 9. This makes it easier for the robot 10 to store the object 20 in the storage section 41.
[0041] Furthermore, in this example, the storage section 41 has a tapered inner side surface whose diameter gradually decreases from the insertion side of the object 20, which makes it easier for the robot 10 to insert the object 20 into the storage section 41. Therefore, the robot 10 can more easily store the object 20 in the storage section 41.
[0042] At least one of the N storage sections 41 does not have to have a tapered inner side surface whose diameter gradually decreases from the insertion side of the target object 20. In this case, the inner surface of at least one of the N storage sections 41 may have the same shape as the side surface of a cylinder with a constant diameter, for example.
[0043] Furthermore, the entire area of the inner side surfaces of the first partial accommodating portion 42 and the second partial accommodating portion 43 of at least one of the N accommodating portions 41 may be closed.
[0044] The robot 10 may also move the target object 20 to a location other than the sorting container 40. For example, the robot 10 may move the target object 20 to a belt conveyor or to a tray other than the tray 30.
[0045] The robot system 80 includes a camera 15 , a camera 16 , and a background member 18 in addition to the robot 10 described above.
[0046] The camera 15 is fixed to, for example, the end effector 12. Therefore, the imaging range of the camera 15 changes depending on the position and posture of the end effector 12. It can also be said that the imaging range of the camera 15 changes depending on the posture of the arm 11. The camera 15 fixed to the end effector 12 can capture, for example, the suction portion 12b side of the end effector 12.
[0047] A camera coordinate system is set for the camera 15. The camera coordinate system is, for example, an xyz Cartesian coordinate system set for the camera 15. The camera coordinate system is a coordinate system that expresses the real space viewed from the camera 15 as a coordinate space. The z direction of the camera coordinate system is set, for example, in the optical axis direction of the camera 15. The optical axis direction of the camera 15 is set, for example, on the same line as the longitudinal direction of the suction nozzle 12a. Therefore, the z direction of the camera coordinate system is set on the same line as the longitudinal direction of the suction nozzle 12a. Hereinafter, simply referring to the camera coordinate system means the camera coordinate system set for the camera 15.
[0048] When the end effector 12 is positioned above the tray 30 so that the longitudinal direction of the suction nozzle 12a is aligned perpendicular to the opening at the top end of the tray 30 and the suction portion 12b faces the tray 30, the camera 15 can photograph the state inside the tray 30. In other words, the camera 15 can photograph multiple objects 20 piled up randomly in the tray 30.
[0049] The camera 15 is, for example, a three-dimensional camera. The camera image (also referred to as a first camera image) generated by the camera 15 includes, for example, a color image and a depth image. The camera 15 outputs the generated first camera image to the processing device 1. The depth image generated by the camera 15 may be generated using a stereo system, a projector system, a combination of the stereo system and the projector system, or another system.
[0050] The depth image (also referred to as the first depth image) included in the first camera image is, for example, a grayscale image. Each pixel value in the first depth image indicates, for example, the distance from the camera 15 to the measurement point in the z direction of the camera coordinate system. The color image (also referred to as the first color image) included in the first camera image may be, for example, an RGB image. The first color image can also be said to be a captured image that shows the state of the camera 15's capture range. Hereinafter, the camera 15 may be referred to as the effector camera 15. Note that a grayscale image may be used instead of the first color image.
[0051] The camera 16 is fixed, for example, at a predetermined location. The imaging range of the camera 16 is fixed. The camera 16 can, for example, capture images of the area above the tray 30 from the side. The optical axis direction of the camera 16 is set, for example, parallel to the horizontal plane. The camera 16 can capture images along the horizontal direction. It can also be said that the optical axis direction of the camera 16 is set parallel to the opening plane of the tray 30.
[0052] The camera 16 is, for example, a three-dimensional camera. The camera image (also referred to as a second camera image) generated by the camera 16 includes, for example, a color image and a depth image. The camera 16 outputs the generated second camera image to the processing device 1. The depth image generated by the camera 16 may be generated using a stereo system, a projector system, a combination of the stereo system and the projector system, or another system.
[0053] The depth image (also referred to as the second depth image) included in the second camera image is, for example, a grayscale image. The color image (also referred to as the second color image) included in the second camera image may be, for example, an RGB image. The second color image can also be considered a captured image that shows the state of the camera 16's capture range. Note that the camera 16 may generate only the second color image of the second color image and the second depth image. In this case, the camera 16 can also be considered, for example, a color camera. Hereinafter, the camera 16 may be referred to as the fixed camera 16. A grayscale image may be used instead of the second color image.
[0054] The background member 18 is a member that serves as the background when the fixed camera 16 takes a photograph. It can also be said that the background member 18 constitutes the photographing background of the fixed camera 16. The shape of the background member 18 may be a plate shape or may be another shape.
[0055] As will be described later, the fixed camera 16 captures an image of the object 20 held by the end effector 12. The surface color of the background member 18 is, for example, a single color that is not included in the surface color of the object 20. It can also be said that the surface color of the background member 18 is a color that can be distinguished from the surface color of the object 20. For example, if the surface color of the object 20 is white or milky white, the surface color of the background member 18 may be black. In this case, the background of the second color image generated by the fixed camera 16 will be black. The surface color of the background member 18 may be a color that does not easily reflect illumination light emitted by lighting fixtures in the imaging environment of the fixed camera 16 (in other words, a color that easily absorbs illumination light).
[0056] <Configuration Example of Processing Device> The processing device 1 performs processing based on the first camera image and the second camera image, and is capable of controlling the robot 10 based on the processing results. Figure 10 is a schematic diagram showing an example of the configuration of the processing device 1.
[0057] 10 , the processing device 1 includes, for example, a processing unit 2, a storage unit 3, and interfaces 4, 5, and 6. The processing device 1 can also be referred to as, for example, a computer device. The processing device 1 can also be referred to as, for example, a processing circuit. The processing device 1 can also be referred to as, for example, a processing system.
[0058] The interface 5 is capable of communicating with the effector camera 15. The processing unit 2 can acquire the first camera image generated by the effector camera 15 through the interface 5. The interface 5 can also be called, for example, an interface circuit, a communication unit, or a communication circuit. The interface 5 may communicate with the effector camera 15 via wired or wireless communication.
[0059] The interface 6 is capable of communicating with the fixed camera 16. The processing unit 2 can acquire the second camera image generated by the fixed camera 16 through the interface 6. The interface 6 can also be referred to as, for example, an interface circuit, a communication unit, or a communication circuit. The interface 6 may communicate with the fixed camera 16 via wired communication or wireless communication.
[0060] The interface 4 is capable of communicating with the robot 10. The processing unit 2 is capable of controlling the robot 10 through the interface 4. The interface 4 may also be referred to as, for example, an interface circuit, a communication unit, or a communication circuit. The interface 4 may communicate with the robot 10 via wired communication or wireless communication.
[0061] The processing unit 2 can generally manage the operation of the processing unit 1 by controlling the other components of the processing unit 1. The processing unit 2 can also be referred to as, for example, a control unit. The processing unit 2 can also be referred to as, for example, a processing circuit or a control circuit. The processing unit 2 includes at least one processor to provide control and processing capabilities for performing various functions, as described in more detail below.
[0062] According to various embodiments, the at least one processor may be implemented as a single integrated circuit (IC) or as multiple communicatively connected integrated circuits ICs and / or discrete circuits. The at least one processor may be implemented according to various known techniques.
[0063] In one embodiment, a processor includes one or more circuits or units configured to perform one or more data computational procedures or processes, for example, by executing instructions stored in associated memory. In other embodiments, a processor may be firmware (e.g., discrete logic components) configured to perform one or more data computational procedures or processes.
[0064] According to various embodiments, the processor may include one or more processors, controllers, microprocessors, microcontrollers, application specific integrated circuits (ASICs), digital signal processors, programmable logic devices, field programmable gate arrays, or any combination of these devices or configurations, or other known devices and configurations, to perform the functions described below.
[0065] The processing unit 2 may include, for example, a CPU (Central Processing Unit) as a processor. The storage unit 3 may include a non-transitory recording medium readable by the CPU of the processing unit 2, such as a ROM (Read Only Memory) and a RAM (Random Access Memory). The storage unit 3 stores, for example, a program 3a for controlling the processing device 1. Various functions of the processing unit 2 are realized, for example, by the CPU of the processing unit 2 executing the program 3a in the storage unit 3.
[0066] The configuration of the processing device 1 is not limited to the above example. For example, the processing unit 2 may include multiple CPUs. The processing unit 2 may also include at least one DSP (Digital Signal Processor). All or some of the functions of the processing unit 2 may be realized by a hardware circuit that does not require software to realize the function. The storage unit 3 may also include a computer-readable non-transitory recording medium other than ROM and RAM. The storage unit 3 may also include, for example, a small hard disk drive or SSD (Solid State Drive).
[0067] The processing device 1 may also include a display unit controlled by the processing unit 2. The display unit may be, for example, a liquid crystal display, an organic electroluminescence (EL) display, or a plasma display. The display unit of the processing device 1 may, for example, display at least one of a first color image and a first depth image generated by the effector camera 15, or at least one of a second color image and a second depth image generated by the fixed camera 16. The display unit of the processing device 1 may also display various recognition results of the object 20.
[0068] The processing device 1 may also include an input unit that accepts input from a user. The input unit may include, for example, a mouse and a keyboard. The input unit may also include a touch sensor that accepts touch operations by the user. The input unit may also include a microphone that accepts voice input by the user.
[0069] The processing device 1 may also be configured with a plurality of computer devices. The processing device 1 may also be a cloud server. In this case, the interface 4 of the processing device 1 may communicate with the robot 10 through a network including the Internet.
[0070] <Example of Operation of Robot System> Fig. 11 is a flowchart showing an example of operation of the robot system 80. As shown in Fig. 11, in step s1, the processing unit 2 controls the arm 11 to move the end effector 12 so that the effector camera 15 can photograph the state inside the tray 30 from above.
[0071] Next, in step s2, the processing unit 2 causes the effector camera 15 to take an image. The effector camera 15 takes an image of the state inside the tray 30 from directly above the tray 30. The effector camera 15 generates a first color image showing the state inside the tray 30.
[0072] Next, in step s3, the processing unit 2 recognizes each object 20 in the tray 30 based on the first color image obtained in step s2. It can also be said that the processing unit 2 detects or searches whether or not the object 20 is captured in the first color image.
[0073] The processing unit 2 may recognize the object 20 using, for example, a trained model that has been trained. The training model used by the processing unit 2 may be configured, for example, by a neural network that realizes instance segmentation. The training model may be configured, for example, by a Mask Scoring R-CNN. R-CNN is an abbreviation for Region Based Convolutional Neural Networks. The training model may also be referred to as, for example, a machine learning model. Note that in this example, the tasks of the robot 10 are performed individually for the first object 20, the second object 20, and the third object 20, but during machine learning, machine learning may be performed for the first object 20, the second object 20, and the third object 20. Furthermore, during machine learning, machine learning may be performed using images in which the first object 20, the second object 20, and the third object 20 are simultaneously captured, or may be performed using images in which the first object 20, the second object 20, and the third object 20 are individually captured. Furthermore, the machine learning process may be performed individually or simultaneously on the first object 20, the second object 20, and the third object 20. The robot 10 performs work individually on the first object 20, the second object 20, and the third object 20, but by performing machine learning simultaneously on the first object 20, the second object 20, and the third object 20 of the same type, the work can be made more efficient.
[0074] A first color image showing the state inside the tray 30 is input to the trained model. The trained model recognizes each object 20 inside the tray 30 based on the input first color image. The first color image can be considered an image for inference. The trained model then outputs a recognition result for each object 20 inside the tray 30.
[0075] The recognition result output by the trained model includes, for example, the mask of each object 20 in the tray 30. Specifically, the recognition result output by the trained model includes position information and shape information of the mask of each object 20 in the tray 30. The position and shape of the mask of the object 20 represent the position and shape of the object 20. It can also be said that the trained model recognizes the position and shape of the object 20.
[0076] Furthermore, the recognition result output by the trained model includes, for example, a classification score for each object 20 in the tray 30, which indicates the accuracy of classification of the object 20. The classification score can be said to indicate the degree of confidence in the classification of the object 20, the likelihood of the classification of the object 20, or the likelihood of the object 20. The classification score is, for example, a numerical value greater than or equal to 0 and less than or equal to 1. The larger the classification score, the higher the accuracy of classification of the object 20. The classification score can also be said to indicate the probability that the mask generated by the trained model is a mask of the object 20. The score can also be said to be an evaluation value.
[0077] Furthermore, the recognition result output by the trained model includes, for example, a mask score indicating the accuracy of generation of each mask generated by the trained model. The mask score is, for example, a numerical value greater than or equal to 0 and less than or equal to 1. The larger the mask score, the higher the accuracy of mask generation. The mask score can be said to indicate the degree of confidence in the position information and shape information of the mask generated by the trained model, or the likelihood of the position information and shape information of the mask, or the certainty of the position information and shape information of the mask.
[0078] Furthermore, the recognition result output by the trained model includes, for each object 20 in the tray 30, a visibility score indicating the proportion of the portion of the object 20 that appears in the first color image to the overall image of the object 20. The visibility score is, for example, a numerical value greater than or equal to 0 and less than or equal to 1. The larger the visibility score, the greater the proportion of the portion of the object 20 that appears in the first color image to the overall image of the object 20.
[0079] Here, the object 20 to be explained in the tray 30, in other words, the object 20 of interest in the tray 30, is referred to as the object of interest 20.
[0080] For example, consider a case where another object 20 overlaps the target object 20, and a part of the target object 20 is not captured in the first color image. In this case, the visibility score of the target object 20 will be small. Also consider a case where a part of the target object 20 is located outside the capture range of the effector camera 15, and a part of the target object 20 is not captured in the first color image. In this case, the visibility score of the target object 20 will also be small.
[0081] On the other hand, if the entire target object 20 is located within the shooting range of the effector camera 15 and no other target objects 20 overlap the target object 20, the visibility score of the target object 20 will be the maximum value of 1.
[0082] The area of the region in the first color image in which the target object 20 actually appears is referred to as the visible area of the target object 20. Furthermore, the area of the region in the first color image in which the target object 20 appears, assuming that the entire target object 20 appears in the first color image, is referred to as the total area of the target object 20. For example, a value obtained by dividing the visible area of the target object 20 by the total area of the target object 20 may be used as the visibility score of the target object 20. The visible area of the target object 20 can also be said to be the area of a mask of the target object 20.
[0083] When a part of the target object 20 is located outside the shooting range of the effector camera 15, the total area is the area of the region in which the target object 20 appears in the first color image if the shooting range of the effector camera 15 were to expand so that the part were included in the shooting range. Also, when the entire target object 20 is located within the shooting range of the effector camera 15 but another object 20 overlaps the target object 20, the total area is the area of the region in which the target object 20 appears in the first color image if the other object 20 does not overlap the target object 20. When the entire target object 20 is located within the shooting range of the effector camera 15 and another object 20 does not overlap the target object 20, the visible area and the total area of the target object 20 are the same.
[0084] The visibility score of the target object 20 can also be said to indicate the proportion of the target object 20 that is actually visible from the effector camera 15 to the entire target object 20 that is visible from the effector camera 15, assuming that the entire target object 20 is visible from the effector camera 15.
[0085] If the camera 15 fixed to the end effector 12 is considered to be the eyes of the robot 10, the first color image obtained by the camera 15 can be considered to represent what the robot 10 sees with its eyes. Therefore, the visibility score of the target object 20 can be said to represent the visibility (in other words, the degree of visibility) of the target object 20 when viewed from the robot 10. Therefore, the larger the visibility score of the target object 20, the easier it is for the robot 10 to see the target object 20, and as a result, the easier it is for the robot 10 to hold the target object 20. The visibility score of the target object 20 can also be said to represent the degree of visibility when the robot 10 looks at the target object 20.
[0086] As described above, the trained model possessed by the processing unit 2 outputs a mask, a classification score, a mask score, and a visibility score for each of the multiple objects 20 in the tray 30. Note that the processing unit 2 may recognize the objects 20 using a method different from the above example.
[0087] FIG. 12 is a schematic diagram showing an example of a binary image 100 including a mask 110 of the target object 20. In FIG. 12, for convenience of explanation, the mask 110 is shown with diagonal hatching. In the binary image 100, for example, each pixel value constituting the mask 110 is 0, and each pixel value outside the mask 110 is 1. The binary image 100 is an image of the same size as the first color image. The position of the mask 110 in the binary image 100 is the same as the position of the area in which the target object 20 appears in the first color image. The outer shape of the mask 110 included in the binary image 100 is the same as the outer shape of the area in which the target object 20 appears in the first color image.
[0088] After step s3, in step s4, the processing unit 2 selects an object 20 to be held by the robot 10 from the plurality of objects 20 in the tray 30 based on the recognition result obtained in step s2.
[0089] The processing unit 2 selects an object 20 to be held by the robot 10 from the plurality of objects 20 in the tray 30, for example, based on the classification score, mask score, and visibility score obtained in step s2. For example, the processing unit 2 calculates a holdability score for each object 20, indicating the possibility that the robot 10 can hold the object 20, based on the classification score, mask score, and visibility score of the object 20. Then, the processing unit 2 selects the object 20 with the highest holdability score from the plurality of objects 20 in the tray 30 as the object 20 to be held by the robot 10.
[0090] When calculating the holdability score of the object 20, the processing unit 2, for example, multiplies the classification score, mask score, and visibility score individually by a weighting coefficient. Then, the processing unit 2 determines the value obtained by adding together the classification score, mask score, and visibility score multiplied by the weighting coefficient as the holdability score. Alternatively, the processing unit 2 may determine the value obtained by multiplying the classification score, mask score, and visibility score as the holdability score.
[0091] The processing unit 2 may select an object 20 to be held by the robot 10 from the plurality of objects 20 in the tray 30 using a method different from the above example.
[0092] After step s4, in step s5, the processing unit 2 estimates the orientation of the object 20 selected in step s4 (also referred to as the selected object 20). An example of a method for estimating the orientation of the selected object 20 will be described below.
[0093] The processing unit 2 first generates a point cloud that three-dimensionally represents the surface of the selection object 20, for example, based on the mask of the selection object 20 and the first depth image obtained by photographing with the effector camera 15 in step s2. The first depth image obtained in step s2 is a depth image obtained by the effector camera 15 photographing the state inside the tray 30 from above the tray 30.
[0094] A point cloud (also referred to as a selection object point cloud) that three-dimensionally represents the surface of the selection object 20 is expressed, for example, by the position coordinates of each point that constitutes the point cloud in a camera coordinate system. Hereinafter, each of the multiple points that constitute the selection object point cloud may be referred to as an object point. The point cloud may also be referred to as a point cloud image.
[0095] Fig. 13 is a schematic diagram showing an example of a selection object point cloud 150. In Fig. 13, for convenience of illustration, the outer shape of the selection object point cloud 150 is shown two-dimensionally with solid lines, but in reality, the selection object point cloud 150 is expressed three-dimensionally.
[0096] The selection object point cloud 150 represents a visible area of the side surface 25 of the selection object 20 when the selection object 20 is viewed from the effector camera 15. The area of the side surface 25 of the selection object 20 that is visible to the camera 15 (in other words, the area visible from the camera 15) can be said to be an area that appears in the first color image, or an area that appears in the first depth image. The selection object point cloud 150 can be said to represent an area of the side surface of the selection object 20 that appears in the first color image, or an area of the surface of the selection object 20 that appears in the first depth image. Hereinafter, the area of the side surface of the selection object 20 that is visible to the camera 15 will be referred to as the visible area.
[0097] Next, the processing unit 2 estimates parallel planes parallel to tangent planes tangent to the side surfaces of the selection object 20 based on the selection object point cloud 150. Herein, in the present disclosure, parallel planes parallel to tangent planes tangent to the surface of the selection object 20 also include planes located on the same plane as the tangent plane. Hereinafter, the term "parallel plane" simply refers to a parallel plane parallel to a tangent plane tangent to the surface of the selection object 20. The processing unit 2 estimating parallel planes can also be seen as the processing unit 2 estimating the curved surface (in other words, the side surface) of the object 20 as a pseudo-plane.
[0098] The processing unit 2 estimates parallel planes using, for example, RANSAC (random sample consensus), which is a plane estimation algorithm. The processing unit 2 randomly selects three object points from the multiple object points that make up the selected object point cloud. The processing unit 2 then sets one plane that passes through the three selected object points as one candidate plane. The processing unit 2 sets multiple candidate planes in the camera coordinate system by repeating this process.
[0099] When a candidate surface is set, the processing unit 2 calculates the number of object points located around the candidate surface (also called peripheral object points) by counting the number of object points whose distance from the set candidate surface is equal to or less than a threshold value. The distance from the candidate surface to an object point is the length of a perpendicular line drawn from the object point to the candidate surface. Hereinafter, the threshold value used to compare the distance of an object point from the candidate surface is referred to as the estimation threshold value.
[0100] The processing unit 2 calculates the number of peripheral object points for each of the plurality of candidate surfaces. The processing unit 2 identifies the candidate surface with the largest number of peripheral object points from among the plurality of candidate surfaces. The processing unit 2 then determines the identified candidate surface with the largest number of peripheral object points as a parallel plane.
[0101] Fig. 14 is a schematic diagram showing an example of a parallel plane 200 estimated by the processing unit 2. Fig. 14 schematically shows the selection object 20 as viewed from one end face side. Note that Fig. 14 does not show the openings formed in the end face of the object 20.
[0102] The processing unit 2 estimates a parallel plane 200 that is parallel to a tangent plane 210 that is tangent to a point 220 that is closest to the camera 15 when the side surface 25 (also referred to as the circumferential side surface) of the cylindrical selection object 20 is viewed along the circumferential direction. For example, the processing unit 2 estimates a parallel plane 200 that is parallel to the tangent plane 210 that is tangent to a point 220 that is closest to the origin 15a of the camera coordinate system when the side surface 25 of the selection object 20 is viewed along the circumferential direction. The tangent plane 210 is a plane that is tangent to a straight line that is located on the side surface 25 of the selection object 20 and extends along the longitudinal direction of the selection object 20. By appropriately setting the estimation threshold value, the processing unit 2 can estimate the parallel plane 200 that is parallel to the tangent plane 210. The parallel plane 200 estimated by the processing unit 2 may be a plane passing through the object to be selected 20 (in other words, a plane cutting the object to be selected 20), a plane located on the same plane as the tangent plane 210, or a plane slightly separated from the object to be selected 20, as in the example of Figure 14.
[0103] Next, the processing unit 2 estimates the orientation of the selection object 20 based on the estimated parallel plane 200, which is the estimated parallel plane 200, and the selection object point cloud 150. This allows the processing unit 2 to appropriately estimate the orientation of the selection object 20.
[0104] The processing unit 2 determines, in the camera coordinate system, the vertical direction perpendicular to the estimated parallel plane 200 as the normal direction perpendicular to the tangent plane 210. The vertical direction perpendicular to the estimated parallel plane 200 can be said to be the estimation result of the normal direction perpendicular to the tangent plane 210. The normal direction is a direction perpendicular to the longitudinal direction of the object 20, and is therefore one piece of information representing the orientation of the selection object 20. The vertical direction perpendicular to the estimated parallel plane 200 is a direction perpendicular to the longitudinal direction of the selection object 20, and is therefore one piece of information representing the orientation of the selection object 20. Hereinafter, the normal direction may be referred to as the object z direction.
[0105] Furthermore, the processing unit 2 identifies a partial point group (also referred to as a parallel plane peripheral point group) located in the periphery of the estimated parallel plane 200 from the selected object point group 150. For example, the processing unit 2 sets, as the parallel plane peripheral point group, a plurality of object points located in the periphery of the candidate surface adopted as the parallel plane 200 from among the plurality of object points constituting the selected object point group. The plurality of object points located in the periphery of the candidate surface are a plurality of object points from the selected object point group 150 whose distance from the candidate surface is equal to or less than an estimation threshold. The parallel plane peripheral point group is expressed three-dimensionally in the camera coordinate system. The estimation threshold may be determined based on the deformability of the suction unit 12b.
[0106] 15 is a schematic diagram showing an example of a parallel plane peripheral point cloud 250 included in the selection object point cloud 150. In FIG. 15, for convenience of illustration, the parallel plane peripheral point cloud 250 is shown two-dimensionally with diagonal lines. As described above, the selection object point cloud 150 represents the viewable area of the side surface 25 of the selection object 20 (i.e., the area of the side surface of the selection object 20 that is viewable by the camera 15). The parallel plane peripheral point cloud 250 represents at least a part of the viewable area of the side surface of the selection object 20.
[0107] Here, two planes that are parallel to the estimated parallel plane 200 and spaced apart from the estimated parallel plane 200 by an estimated threshold value are defined as a first threshold plane 205 and a second threshold plane 206. Fig. 16 is a schematic diagram showing an example of the estimated parallel plane 200, the first threshold plane 205, and the second threshold plane 206. The parallel plane surrounding point cloud 250 is made up of a plurality of object points that represent a region 26 (shown by a thick line in Fig. 16 ) of the side surface 25 of the selected object 20 that is located between the first threshold plane 205 and the second threshold plane 206.
[0108] Next, the processing unit 2 projects the obtained parallel plane peripheral point group 250 onto a projection plane 260. The projection plane 260 is a plane, for example, the xy plane of the image coordinate system of the image generated by the camera 15. The processing unit 2 then acquires a minimum circumscribing rectangle 280 that circumscribes the parallel plane peripheral point group 250 (also referred to as a projection point group 255) projected onto the projection plane 260. Fig. 17 is a schematic diagram showing an example of the projection point group 255 and the minimum circumscribing rectangle 280. One object point in the camera coordinate system is projected onto one point in the image coordinate system that corresponds to the one object point.
[0109] Next, the processing unit 2 identifies the long side of the minimum circumscribing rectangle 280. Then, the processing unit 2 estimates the longitudinal direction of the object to be selected 20 based on the long side of the minimum circumscribing rectangle 280. This allows the processing unit 2 to appropriately estimate the longitudinal direction of the object to be selected 20. For example, the processing unit 2 identifies two endpoints at both ends of the long side of the minimum circumscribing rectangle 280 in the image coordinate system. Next, the processing unit 2 projects the two endpoints in the identified image coordinate system onto the camera coordinate system. Then, the processing unit 2 determines the direction connecting the two endpoints projected onto the camera coordinate system as the longitudinal direction of the object to be selected 20.
[0110] Hereinafter, the longitudinal direction of the selection object 20 may be referred to as the object x-direction. Furthermore, when simply referring to the longitudinal direction in the description of the object 20, it means the longitudinal direction of the object 20. Furthermore, when simply referring to the visible area of the selection object 20, it means the visible area of the side of the selection object 20.
[0111] In step s5, the processing unit 2 calculates the distance between the two end points projected onto the camera coordinate system. Then, the processing unit 2 sets the calculated distance as the longitudinal length of the visible region of the selection object 20. It can be said that the processing unit 2 estimates the longitudinal length of the visible region of the selection object 20 based on the long side of the minimum circumscribing rectangle 280.
[0112] Hereinafter, the term "length of the visible region of the selection object 20" simply refers to the longitudinal length of the visible region on the side of the selection object 20. The length of the visible region of the selection object 20 is used in step s7 described below.
[0113] The processing unit 2 estimates an object y direction perpendicular to the object z direction and the object x direction. For example, the processing unit 2 determines the direction orthogonal to the object z direction (i.e., normal direction) estimated for the selected object 20 and the estimated object x direction (i.e., longitudinal direction of the object 20) as the object y direction.
[0114] The object x direction, object y direction, and object z direction represent the three-dimensional orientation of the selection object 20. The processing unit 2 estimates the object x direction, object y direction, and object z direction of the selection object 20, thereby estimating the three-dimensional orientation of the selection object 20. In this example, the processing unit 2 estimates the object x direction, object y direction, and object z direction based on the estimated parallel plane 200 and the object point cloud 150, and therefore can appropriately estimate the object x direction, object y direction, and object z direction.
[0115] When the orientation of the selection object 20 is estimated in step s5, the processing unit 2 determines in step s6 a holding position on the side surface 25 of the selection object 20 where the robot 10 will hold the object. The processing unit 2, for example, identifies the center point of the minimum circumscribing rectangle 280 obtained in step s5. Next, the processing unit 2 projects the identified center point onto the camera coordinate system. Then, the processing unit 2 determines the center point projected onto the camera coordinate system (also referred to as the projection center point) as the holding position.
[0116] The projection center point, i.e., the holding position, corresponds to the center point of the viewing area of the selection object 20. The holding position can be said to be an estimation result of the center point of the viewing area of the selection object 20. The holding position can also be said to be an estimation result of the center point of the area of the side of the selection object 20 that appears in the first color image generated by the camera 15, or can be said to be an estimation result of the center point of the area that appears in the first depth image generated by the camera 15. When the side surface 25 of the selection object 20 is viewed in the circumferential direction, the holding position coincides with or is close to the point 220 (see FIG. 14 ) that is closest to the camera 15. The holding position is located on the tangent plane 210 or close to the tangent plane 210.
[0117] For example, consider a case where no other objects 20 overlap the selection object 20 and the side of the selection object 20 is visible from end to end along the longitudinal direction by the camera 15. In this case, the holding position coincides with or is close to the center in the longitudinal direction of the side of the selection object 20. In this case, when the robot 10 holds the selection object 20 at the holding position, the robot 10 holds the center of the selection object 20 in the longitudinal direction.
[0118] On the other hand, if another object 20 overlaps the selection object 20 and the side of the selection object 20 cannot be seen from end to end along the longitudinal direction by the camera 15, the holding position may be shifted from the center in the longitudinal direction of the side of the selection object 20. In this case, when the robot 10 holds the selection object 20 at the holding position, the robot 10 may hold a portion of the selection object 20 that is away from the center in the longitudinal direction.
[0119] In this way, in this example, the processing unit 2 determines the holding position based on the estimated parallel plane 200 and the object point cloud 150, and therefore can determine an appropriate holding position.
[0120] Once the holding position is determined in step s6, the processing unit 2 determines in step s7 whether to reselect the object 20 to be held by the robot 10. If the processing unit 2 determines to reselect the object 20, it executes step s4 again to reselect the object 20. At this time, the processing unit 2 selects, for example, the object 20 with the next highest holdability score. Thereafter, step s5 is executed again, and thereafter the robot system 80 operates in the same manner.
[0121] On the other hand, if in step s7 the processing unit 2 does not determine to reselect the object 20 to be held by the robot 10, in step s8 the robot 10 holds the currently selected object 20 with the end effector 12 under the control of the processing unit 1.
[0122] In step s7, the processing unit 2 performs a first determination process, a second determination process, and a third determination process to determine whether to reselect the object 20.
[0123] <First Determination Process> In the first determination process, the processing unit 2 determines whether to reselect the object 20 based on the length of the visible region of the selection object 20 estimated in step s5 (also referred to as the estimated length of the visible region of the selection object 20). For example, the processing unit 2 determines to reselect the object 20 when the estimated length of the visible region of the selection object 20 is equal to or greater than an upper limit. Furthermore, the processing unit 2 determines to reselect the object 20 when the estimated length of the visible region of the selection object 20 is equal to or less than a lower limit.
[0124] Here, in this example, if the robot 10 holds the selection object 20 away from the longitudinal center, the robot 10 may have difficulty storing the selection object 20 in the storage section 41 of the sorting container 40.
[0125] For example, consider a case where the selected object 20 is the third object 20, which is the longest, and the robot 10 holds a portion of the third object 20 that is away from the center and close to the first end face 21. In this case, if the robot 10 inserts the held third object 20 into the storage section 41 of the sorting container 40 from the second end face 22 side, as shown in FIGS. 8 and 9 above, the bottom end of the third object 20 (i.e., the end on the second end face 22 side) may interfere with the bottom of the storage section 41. As a result, the robot 10 may have difficulty storing the third object 20 in the storage section 41.
[0126] As another example, consider a case where the selected object 20 is the shortest first object 20, and the robot 10 holds a portion of the first object 20 near the second end surface 22, away from the center. In this case, if the robot 10 inserts the held first object 20 into the storage section 41 of the sorting container 40 from the second end surface 22 side, as shown in FIGS. 8 and 9 above, there is a possibility that the robot 10 will release its hold on the first object 20 before the first object 20 reaches the cylindrical third portion storage section 44 located at the bottom of the storage section 41. As a result, there is a possibility that the first object 20 will not be properly stored in the storage section 41.
[0127] In this way, in this example, the robot 10 holds the longitudinal center of the selection object 20, thereby performing the same operation regardless of the length of the selection object 20, making it easier to store the selection object 20 in the storage section 41.
[0128] On the other hand, as described above, if the side surface 25 of the selection object 20 is not visible from end to end along the longitudinal direction by the camera 15, when the robot 10 holds the selection object 20 in the holding position, the robot 10 may hold a portion of the selection object 20 that is away from the center in the longitudinal direction. When the side surface 25 of the selection object 20 is not visible from end to end along the longitudinal direction by the camera 15, the estimated length of the visible region of the selection object 20 becomes small. Therefore, when the estimated length of the visible region of the selection object 20 is equal to or less than the lower limit, when the robot 10 holds the selection object 20 in the holding position, the robot 10 may hold a portion of the selection object 20 that is away from the center in the longitudinal direction, and as a result, it may be difficult to store the selection object 20 in the storage unit 41.
[0129] Therefore, in this example, as described above, when the estimated length of the visible area of the selection object 20 is equal to or less than the lower limit, the processing unit 2 determines to reselect the object 20. As a result, the currently selected selection object 20 is no longer held by the robot 10, and a new object 20 in the tray 30 is selected. This reduces the possibility that the object 20 will be difficult to accommodate in the sorting container 40.
[0130] Furthermore, if the estimated length of the visual recognition area of the selection object 20 is too large and the estimation accuracy of the length of the visual recognition area of the selection object 20 is poor, when the robot 10 holds the selection object 20 at the holding position, the robot 10 may hold a portion of the selection object 20 that is far from the center in the longitudinal direction. Therefore, in this example, as described above, the processing unit 2 determines to reselect the object 20 when the estimated length of the visual recognition area of the selection object 20 is equal to or greater than the upper limit value. As a result, the currently selected selection object 20 is no longer held by the robot 10, and a new object 20 in the tray 30 is selected. This reduces the possibility that the object 20 will be difficult to accommodate in the sorting container 40.
[0131] The processing unit 2 sets the upper limit and lower limit used in step s7 based on, for example, the actual length in the longitudinal direction of the object 20 that is the current work target of the robot 10. The actual length in the longitudinal direction of the object 20 (also referred to as the actual length of the object 20) is stored in the memory unit 3. The actual length of the object 20 may be input to the processing device 1 by a user.
[0132] For example, let L1 be the actual length of the first object 20, L2 be the actual length of the second object 20, and L3 be the actual length of the third object 20. When the current work object is the first object 20, i.e., when the first tray 30 is located within the work range of the robot 10, the processing unit 2 sets the upper and lower limits based on the actual length L1. When the current work object is the second object 20, the processing unit 2 sets the upper and lower limits based on the actual length L2. When the current work object is the third object 20, the processing unit 2 sets the upper and lower limits based on the actual length L3. For example, the processing unit 2 may set the upper limit to α times the actual length L of the current work object 20 and the lower limit to β times the actual length L (β<α). α may be, for example, 1.0 or greater and 1.2 or less. β may be, for example, 0.8 or greater and less than 1.0.
[0133] The values of α and β are not limited to the above example. Furthermore, instead of the processing unit 2 determining the upper limit value and the lower limit value, the upper limit value and the lower limit value for each type of object 20 may be stored in advance in the storage unit 3. Furthermore, the user may input the upper limit value and the lower limit value for each type of object 20 to the processing device 1.
[0134] <Second judgment process> In the second judgment process, the processing unit 2 determines whether to reselect the object 20 based on a comparison between the normal direction (also called the estimated normal direction) of the selection object 20 estimated in step s5 and the vertical direction (in other words, the direction of gravity).
[0135] Here, if the estimated normal direction is significantly tilted relative to the vertical direction, when the end effector 12 approaches the object to be selected 20 with the suction portion 12b in order to hold the object to be selected 20, it may interfere with the tray 30 and be unable to hold the object to be selected 20.
[0136] Therefore, in the second determination process, the processing unit 2 calculates the inclination angle of the estimated normal direction with respect to the vertical direction. If the calculated inclination angle is equal to or greater than the first angle, the processing unit 2 determines that the target object 20 should be reselected. This reduces the possibility that the end effector 12 will interfere with the tray 30. The first angle is set, for example, to between 45 degrees and 60 degrees. However, the first angle is not limited to this.
[0137] <Third judgment process> In the third judgment process, the processing unit 2 judges whether to reselect the object 20 based on the degree of overlap (also simply referred to as the degree of overlap of the suction portion) between the suction opening 12bb and the selection object 20 when the suction portion 12b holds the selection object 20 in the holding position.
[0138] For example, as shown in FIG. 18 , the processing unit 2 places a graphic 350 on the mask 110 of the selection object 20 shown in the binary image 100 described above, which corresponds to the suction opening 12bb when the suction unit 12b holds the selection object 20 at the holding position. Next, the processing unit 2 calculates the overlap ratio of the graphic 350 with the mask 110. For example, the processing unit 2 divides the area of the region of the graphic 350 that overlaps with the mask 110 by the total area of the graphic 350 to obtain the overlap ratio. The overlap ratio represents the degree of overlap of the suction units. If the calculated overlap ratio is less than a threshold value, the processing unit 2 determines to reselect the object 20. This reduces the likelihood that the robot 10 will attempt to suction the selection object 20 when another selection object 20 overlaps the selection object 20 and the suction unit 12b is unlikely to be able to properly suction the selection object 20. The threshold value may be, for example, 0.99 or another value.
[0139] If the processing unit 2 does not determine to reselect the object 20 in each of the first judgment process, the second judgment process, and the third judgment process, step s8 is executed and the selected object 20 is held by the end effector 12.
[0140] In step s8, the processing unit 2 controls the robot 10 so that the suction unit 12b approaches the holding position of the selection object 20 along a predetermined direction, thereby causing the suction unit 12b to suction and hold the selection object 20 at the holding position. If the inclination angle of the estimated normal direction with respect to the vertical direction (also referred to as the inclination angle of the estimated normal direction) is less than a second angle, the processing unit 2 controls the robot 10 so that the suction unit 12b approaches along the estimated normal direction. The second angle is a value smaller than the first angle used in the second determination processing of step s7. On the other hand, if the inclination angle of the estimated normal direction is equal to or greater than the second angle, the processing unit 2 controls the robot 10 so that the suction unit 12b approaches the holding position along a direction (also referred to as the relaxation direction) inclined from the vertical direction toward the estimated normal direction by the second angle. In this case, the suction unit 12b approaches from a direction inclined toward the vertical by a difference angle obtained by subtracting the first angle from the second angle with respect to the estimated normal direction.
[0141] For example, consider a case where the second angle is 35 degrees and the first angle is 45 degrees. If the inclination angle of the estimated normal direction is 45 degrees or more, the processing unit 2 determines in step s7 to reselect the object 20, and therefore the suction unit 12b does not approach the currently selected selected object 20.
[0142] When the inclination angle of the estimated normal direction is less than 35 degrees, the suction unit 12b approaches along the estimated normal direction. When the inclination angle of the estimated normal direction is 35 degrees or more and less than 45 degrees, the suction unit 12b approaches the holding position along a direction inclined 35 degrees toward the estimated normal direction from the vertical. In this case, the suction unit 12b approaches from a direction inclined 10 degrees toward the vertical from the estimated normal direction. In this way, even when the suction unit 12b approaches from a direction slightly inclined toward the vertical from the estimated normal direction, the elastic deformation of the suction unit 12b makes it possible to appropriately suction the selection target object 20.
[0143] The second angle is set, for example, to a limit value at which the end effector 12 does not interfere with the tray 30 when the suction portion 12b approaches along the relaxation direction. The difference angle between the third angle and the second angle (10 degrees in the above example) is set based on the shape and elastic characteristics of the suction portion 12b. The third angle can also be said to be an angle obtained by adding a margin angle (e.g., 10 degrees) due to elastic deformation of the suction portion 12b to an approach angle (e.g., 35 degrees) at which the end effector 12 interferes with the tray 30.
[0144] After step s8, in step s9, the processing unit 2 detects the deviation in the longitudinal direction (also referred to as the estimated longitudinal direction) of the object to be selected 20 estimated in step s5, and determines whether the detected deviation is large.
[0145] 19 , the processing unit 2 first causes the robot 10 to move the selection object 20 to a position in front of the camera 16 so that the selection object 20 is photographed by the camera 16. At this time, the processing unit 2 controls the robot 10, for example, so that the estimated longitudinal direction of the selection object 20 is positioned on the same line as the optical axis direction of the camera 16, so that the longitudinal direction of the suction nozzle 12a is aligned vertically, and so that the suction opening 12bb of the suction unit 12b faces the tray 30. Next, the processing unit 2 causes the camera 16 to start photographing.
[0146] FIG. 20 is a schematic diagram illustrating an example of the second color image 160 generated by the camera 16 in step s9. The second color image 160 captures either the first end face 21 or the second end face 22 of the selection object 20. In the example of FIG. 20, the second color image 160 captures the second end face 22. In step s9, the processing unit 2 does not know whether the first end face 21 or the second end face 22 of the selection object 20 faces toward the camera 16. Because the camera 16 captures a monochrome background member 18, the background of the second color image 160 is monochrome (e.g., black). The processing unit 2 detects the estimated longitudinal misalignment based on the second color image 160. This allows the processing unit 2 to appropriately detect the estimated longitudinal misalignment.
[0147] If there is no deviation in the estimated longitudinal direction of the selection object 20, i.e., if the estimated longitudinal direction coincides with the actual longitudinal direction of the selection object 20, only one end face of the selection object 20 will be captured in the second color image 160, as shown in the example of FIG. 20 . On the other hand, if the estimated longitudinal direction of the selection object 20 is deviated, and the estimated longitudinal direction of the selection object 20 when held by the suction unit 12b is rotated around the longitudinal direction of the suction nozzle 12a (in other words, the z direction or vertical direction of the camera coordinate system) relative to the actual longitudinal direction, not only one end face but also the side surface 25 of the selection object 20 will be captured, as shown in FIG. 21 . Hereinafter, the longitudinal direction of the suction nozzle 12a may be referred to as the reference direction. The selection object 20 held by the suction unit 12b will be referred to as the held selection object 20.
[0148] The processing unit 2 calculates the width d (also referred to as the image width d) of the retained selection object 20 appearing in the second color image 160. Here, the image width d is the dimension of the retained selection object 20 appearing in the second color image 160 in the y direction of the image coordinate system of the second color image 160 (the horizontal direction in FIGS. 20 and 21 ). For example, the processing unit 2 can calculate the image width d based on a binary image obtained by binarizing the second color image 160. Because the background of the second color image 160 is a single color, the processing unit 2 can appropriately calculate the image width d.
[0149] Based on the calculated image width d, the processing unit 2 calculates a rotation angle θ of the estimated longitudinal direction of the held selection object 20 relative to the actual longitudinal direction (also referred to as the real longitudinal direction) of the held selection object 20. The rotation angle θ is a rotation angle of the estimated longitudinal direction around a reference direction relative to the real longitudinal direction when the selection object 20 is held by the suction unit 12b.
[0150] Fig. 22 is a schematic diagram for explaining an example of a method for determining the rotation angle θ. Fig. 22 shows the holding selection object 20 as viewed from a reference direction 420. In Fig. 22, the actual longitudinal direction 400 of the holding selection object 20 is indicated by a thin solid line, and the estimated longitudinal direction 410 of the holding selection object 20 is indicated by a thin dashed line.
[0151] The actual horizontal width t (also referred to as real width t) of the holding selection object 20 when viewed along the optical axis direction of the camera 16 is expressed by the following equation (1).
[0152]
[0153] In formula (1), l represents the actual length of the selected object 20. In formula (1), h represents the actual diameter of the end face of the selected object 20.
[0154] When the rotation angle θ is sufficiently small, cos θ=1 and sin θ=θ can be approximated, so equation (1) can be approximated to equation (2) below.
[0155]
[0156] The processing unit 2 calculates the actual width t based on the image width d. Then, the processing unit 2 calculates the rotation angle θ by substituting the calculated actual width t, the length l of the object 20 to be held and selected, and the diameter h of the end face of the object to be held and selected into equation (2). The rotation angle θ represents the estimated amount of deviation in the longitudinal direction. In this example, the estimated deviation in the longitudinal direction is detected by calculating the rotation angle θ.
[0157] When the rotation angle θ is equal to or greater than a threshold angle, the processing unit 2 determines that the deviation in the estimated longitudinal direction is large. On the other hand, when the rotation angle θ is less than the threshold angle, the processing unit 2 determines that the deviation in the estimated longitudinal direction is not large. The threshold angle is set to, for example, a few degrees or more and less than 20 degrees.
[0158] If the processing unit 2 determines in step s9 that the estimated longitudinal deviation is large, step s10 is executed. In step s10, the processing unit 2 controls the robot 10 to return the holding selection object 20 to the tray 30. Then, step s1 is executed, and thereafter, the robot system 80 operates in the same manner.
[0159] When inserting and arranging the objects to be held and selected 20 into the storage section 41 of the sorting container 40, the robot 10 inserts the objects to be held and selected 20 into the storage section 41 in a state in which the estimated longitudinal direction of the objects to be held and selected 20 is set to be parallel to the vertical direction. Therefore, if there is a large deviation in the estimated longitudinal direction, the actual longitudinal direction of the objects to be held and selected 20 will be significantly tilted with respect to the depth direction of the storage section 41, and the robot 10 may not be able to properly insert and arrange the objects to be held and selected 20 into the storage section 41.
[0160] Therefore, in this example, as described above, when it is determined that the deviation in the estimated longitudinal direction is large, the currently held object 20 is returned to the tray 30. This reduces the possibility that the held and selected object 20 will not be properly inserted and placed in the storage section 41.
[0161] If it is determined in step s9 that the estimated longitudinal deviation is not large, step s11 is executed. In step s11, the processing unit 2 performs end face determination to determine whether the end face of the holding selection object 20 facing the camera 16 is the first end face 21 or the second end face 22. The processing unit 2 performs the end face determination based on the second color image 160 (see FIGS. 20 and 21 ) obtained in step s9, which shows one end face of the holding selection object 20. Note that step s11 may be executed before step s9.
[0162] The processing unit 2 may perform the end face determination using, for example, a trained model that has been trained. The second color image 160 obtained in step s9 is input to the trained model. The trained model determines whether the end face of the held selection object 20 facing the camera 16 is the first end face 21 or the second end face 22 based on the second color image 160, which captures one end face of the held selection object 20. In other words, the trained model determines whether the end face of the held selection object 20 captured in the second color image 160 is the first end face 21 or the second end face 22 based on the second color image 160. This allows the processing unit 2 to determine whether each end face of the selection object 20 held by the robot 10 is the first end face 21 or the second end face 22. Note that the end face determination may be performed in a state where each end face is not directly facing the camera but is tilted relative to the camera.
[0163] When the end face determination is performed, in step s12, the processing unit 2 controls the robot 10 so that the selection objects 20 are inserted into the storage section 41 of the sorting container 40 from the second end face 22 side. As a result, the selection objects 20 are arranged in the sorting container 40 so that the first end face 21 of the selection objects 20 is positioned on the upper side, as shown in Figures 6 and 7 .
[0164] Thereafter, step s1 is executed again, and thereafter, the robot system 80 operates in the same manner.
[0165] If the determination in step s9 is YES, the processing unit 2 may correct the estimated longitudinal direction by the calculated rotation angle θ. In this case, step s12 may be executed. In step s12, the robot 10 inserts the object 20 to be held and selected into the storage unit 41 with the corrected estimated longitudinal direction set to be parallel to the vertical direction, and stores the object 20 in the storage unit 41.
[0166] In the above example, the processing device 1 controls the robot 10, but the robot 10 may be controlled by a device other than the processing device 1. In this case, the processing device 1 outputs information necessary for controlling the robot 10, such as the estimated posture and holding position of the selection object 20, to the other device that controls the robot 10.
[0167] Although the processing device and the robot system have been described in detail above, the above description is merely illustrative in all respects and does not limit the scope of this disclosure. Furthermore, the various examples described above can be combined and applied as long as they are not mutually inconsistent. It is understood that countless examples not illustrated can be envisioned without departing from the scope of this disclosure.
[0168] This disclosure includes the following:
[0169] In one embodiment, (1) the processing device includes a processing unit that estimates the longitudinal direction of a rod-shaped object that is the work target of the robot, and detects a deviation of the estimated longitudinal direction based on a captured image that shows the object.
[0170] (2) In the processing device of (1) above, the background of the photographed image is monochromatic.
[0171] (3) In a processing device according to (1) or (2) above, the processing unit estimates a parallel plane parallel to a tangent plane that contacts the side of the object based on a point cloud representing the surface of the object, and estimates the longitudinal direction based on the estimated parallel plane, which is the estimated parallel plane, and the point cloud.
[0172] (4) In the processing device of (3) above, the processing unit projects a partial point group of the point group located around the estimated parallel plane onto a projection plane, obtains a minimum circumscribing rectangle that circumscribes the partial point group projected on the projection plane, and estimates the longitudinal direction based on the long side of the obtained minimum circumscribing rectangle.
[0173] (5) In the processing device of (3) or (4) above, the processing unit estimates the posture of the object, including the estimation of the longitudinal direction, based on the estimated parallel plane and the point cloud.
[0174] (6) In the processing device according to any one of (3) to (5) above, the processing unit determines a holding position at which the robot holds the object based on the estimated parallel plane and the point cloud.
[0175] (7) In the processing device of (5) above, the processing unit controls the robot and causes the robot to hold the object based on the estimated posture.
[0176] (8) In the processing device of (7) above, the processing unit causes the robot to align the orientations of a plurality of rod-shaped objects that are work targets of the robot and place the plurality of objects in a container.
[0177] (9) In the processing device of (8) above, the container has a plurality of storage sections into which the plurality of objects are respectively inserted and placed, and at least one of the plurality of storage sections has a tapered inner side surface whose diameter gradually decreases from the insertion side of the object.
[0178] (10) A robot system includes any one of the processing devices described in (7) to (9) above and a robot controlled by the processing device.
[0179] (11) The program causes a computer device to function as the processing unit of any one of the processing devices (1) or (9) above.
[0180] REFERENCE SIGNS LIST 1 Processing device 2 Processing unit 3a Program 10 Robot 20 Object 40 Alignment container 41 Storage unit
Claims
A processing device comprising: a processing unit that estimates the longitudinal direction of a rod-shaped object that is a work target of a robot, and detects a deviation of the estimated longitudinal direction based on a captured image showing the object.
2. The processing device according to claim 1, A processing device, wherein the background of the captured image is monochromatic.
3. The processing apparatus according to claim 1 or 2, The processing unit estimating a parallel plane parallel to a tangent plane that contacts a side surface of the object based on the point cloud representing the surface of the object; a processing device that estimates the longitudinal direction based on an estimated parallel plane that is the estimated parallel plane and the point cloud.
4. The processing device according to claim 3, The processing unit projecting a partial point group located in the periphery of the estimated parallel plane from the point group onto a projection surface; obtaining a minimum circumscribing rectangle that circumscribes the partial point group projected onto the projection plane; A processing device that estimates the longitudinal direction based on the acquired long side of the minimum circumscribing rectangle. The processing apparatus according to claim 3 or claim 4, The processing unit estimates the posture of the object, including the longitudinal direction, based on the estimated parallel plane and the point cloud.
6. The processing apparatus according to claim 3, wherein: The processing unit determines a holding position at which the robot holds the object based on the estimated parallel plane and the point cloud.
6. The processing device according to claim 5, The processing unit Controlling the robot A processing device that causes the robot to hold the object based on the estimated posture.
8. The processing device according to claim 7, The processing unit is a processing device that causes the robot to align the orientations of multiple rod-shaped objects that are work targets of the robot and place the multiple objects in a container.
9. The processing device according to claim 8, the container has a plurality of storage sections into which the plurality of objects are respectively inserted and disposed, At least one of the plurality of storage sections has a tapered inner side surface whose diameter gradually decreases from the insertion side of the object. A processing device according to any one of claims 7 to 9; a robot controlled by the processing device; A robot system comprising: A program for causing a computer device to function as the processing unit of the processing device according to any one of claims 1 to 9.
Citation Information
Patent Citations
Work holding position measuring method using visual sensor and work holder
JP1990250791A
Automatic assembler provided with visual sense
JP1991239487A
Picking feeding device for cigarette
JP1999129179A
Picking device for workpieces loaded in bulk and method for controlling the same
JP2010120141A
Robot apparatus, robot system, and method for manufacturing workpiece
JP2013078825A