Training method, training device, control device, robot system, and program
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- KYOCERA CORP
- Filing Date
- 2026-01-20
- Publication Date
- 2026-07-30
Smart Images

Figure JP2026001661_30072026_PF_FP_ABST
Abstract
Description
Learning Method, Learning Device, Control Device, Robot System, and Program
[0001] The present disclosure relates to a technique for recognizing an object.
[0002] Patent Document 1 describes a technique for recognizing an object.
[0003] Japanese Patent Application Laid-Open No. 2010-120414
[0004] A learning method, a learning device, a control device, a robot system, and a program are disclosed. In one embodiment, the learning method generates a learned model for recognizing an elongated recognition target part shown in an inference image based on an inference image by learning the learning model based on learning data including a learning image and annotation data. The annotation data includes a bounding box set for a recognition target part image in which the recognition target part appears in the learning image. The bounding box includes the recognition target part image and is larger than the recognition target part image.
[0005] Also, in one embodiment, the learning device executes the above learning method.
[0006] Also, in one embodiment, the control device includes a control unit that recognizes a side surface based on an image in which the side surface of a plate-shaped member appears and causes the robot to hold the plate-shaped member based on the recognition result of the side surface.
[0007] Also, in one embodiment, the robot system includes the above control device and a robot controlled by the above control device.
[0008] Also, in one embodiment, the program is a program for causing a computer device to execute the above learning method.
[0009] Also, in one embodiment, the program is a program for causing a computer device to function as the above control device.
[0010] Figure 1 is a schematic diagram showing an example of a robot system. Figure 2 is a schematic diagram showing an example of multiple circuit boards housed in a rack. Figure 3 is a schematic diagram showing an example of robot operation. Figure 4 is a schematic diagram showing an example of robot operation. Figure 5 is a schematic diagram showing an example of robot operation. Figure 6 is a schematic diagram showing an example of robot operation. Figure 7 is a schematic diagram showing an example of the configuration of a control device. Figure 8 is a schematic diagram showing an example of a camera image (inference image). Figure 9 is a schematic diagram showing an example of a trained model provided by the control unit. Figure 10 is a schematic diagram to explain an example of a training method. Figure 11 is a schematic diagram showing an example of a training image. Figure 12 is a schematic diagram showing an example of a mask and bounding box included in annotation data. Figure 13 is a schematic diagram showing an example of a mask included in the recognition result. Figure 14 is a schematic diagram showing an example of a mask included in the recognition result. Figure 15 is a schematic diagram showing an example of the shape of an anchor box. Figure 16 is a schematic diagram showing an example of the relationship between an anchor box and a BBOX equivalent part. Figure 17 is a schematic diagram showing an example of the relationship between an anchor box and a BBOX equivalent part. Figure 18 is a schematic diagram showing an example of a training image. Figure 19 is a schematic diagram showing an example of a mask and bounding box included in annotation data. Figure 20 is a schematic diagram showing an example of the relationship between an anchor box and a BBOX equivalent portion. Figure 21 is a schematic diagram showing an example of an inference image.
[0011] <Outline of an Example of a Robot System> Figure 1 is a schematic diagram showing an example of a robot system 100. As shown in Figure 1, the robot system 100 includes, for example, a robot 10, a camera 15, and a control device 1 that controls the robot 10 and the camera 15.
[0012] The robot 10, under the control of the control device 1, can, for example, hold an object 50 located at a certain location, move the held object 50 to another location, and place it there. The robot 10 may also, under the control of the control device 1, load the held object 50 into another device in the next process. The object 50 can also be called, for example, an object to be held, an object to be transferred, or an object to be worked on. However, the operations performed by the robot 10 are not limited to these.
[0013] The robot 10 is, for example, an arm-type robot. The robot 10 comprises, for example, an arm 11 and an end effector 12 connected to the arm 11. The end effector 12 is, for example, a holding mechanism capable of holding an object. The end effector 12 may be, for example, a hand having a plurality of finger portions 13 capable of grasping an object. In the example of Figure 1, the end effector 12 has two finger portions 13, but the number of finger portions 13 may be three or more. The end effector 12 can hold an object 50 by gripping it with the plurality of finger portions 13. The end effector 12 may also be a suction mechanism having a suction portion for adsorbing the object 50.
[0014] The arm 11 has, for example, multiple joints. The posture of the arm 11 changes as the amount of rotation of at least one of the multiple joints changes. As a result of the change in the posture of the arm 11, the position and posture of the end effector 12 change. Also, as a result of the change in the posture of the arm 11, the position and posture of the object 50 held by the multiple finger portions 13 of the end effector 12 change.
[0015] Camera 15 is fixed to, for example, the end effector 12. Therefore, the shooting range of camera 15 changes according to the position and orientation of the end effector 12. It can also be said that the shooting range of camera 15 changes according to the orientation of arm 11. Camera 15 fixed to the end effector 12 can, for example, photograph the tip side of the finger portion 13 of the end effector 12. Note that camera 15 may also be fixed to arm 11, for example. When camera 15 is fixed to arm 11 or end effector 12, the control device 1 can also control the shooting position of camera 15 by controlling the orientation of robot 10.
[0016] The camera image generated by camera 15 is, for example, a color image. The color image may also be, for example, an RGB image. The color image can also be described as a captured image that shows the area within the shooting range of camera 15. The camera image generated by camera 15 is input to control device 1.
[0017] The camera image may be a grayscale image. The camera 15 may also be a three-dimensional camera. In this case, the camera 15 generates, for example, a color image and a depth image. A depth image is an image containing information about the depth, height, or distance from the camera 15 to the subject. For generating the depth image, a stereo system, a projector system, a combination of the stereo system and the projector system, or other systems may be used.
[0018] The robot 10 can, for example, grasp one object 50 from a plurality of objects 50 housed in a rack 20 with its plurality of fingers 13. Each of the plurality of substrates 50 has, for example, a thickness of 1 mm. The rack 20 can also be described as a storage box for temporarily housing the plurality of objects 50. The object 50 is, for example, a plate-shaped member. The object 50 is, for example, a substrate. The substrate may be, for example, a printed circuit board on which a circuit is formed, or it may be another type of substrate. The outer shape of the substrate may be, for example, a rectangle or a circle. In this example, the case in which the object 50 is a rectangle substrate will be described later. In this case, the object 50 may be referred to as a substrate 50. Also, in this specification, it is possible to interpret the substrate 50 as object 50.
[0019] The rack 20 containing multiple objects 50 may be placed at any position within a predetermined range, provided that the recognition target portion of each object 50 is in a state where it can be photographed. The predetermined range is, for example, the working range of the robot 10. As will be described later, since the control device 1 recognizes the objects 50, the robot 10 can be made to perform work as long as the rack 20 is within the predetermined range. Alternatively, a positioning jig for positioning the rack 20 may be placed within the working range of the robot 10, and the rack 20 may be placed at a predetermined position.
[0020] The finger portion 13 has a shape that allows it to be inserted into, for example, a narrow gap. The finger portion 13 is, for example, flat. As a result, the finger portion 13 can be inserted between multiple objects 50, between the objects 50 and the inner wall surface of the rack 20, or between the objects 50 and other objects. Also, when the finger portion 13 is flat, the contact area between the finger portion 13 and the substrate 50 can be increased when gripping the substrate 50 with multiple finger portions 13. In this disclosure, an example in which multiple finger portions 13 are flat is described, but a flat claw portion may be provided at the tip of each of the multiple finger portions 13. In this case, for example, the finger portion 13 is columnar, and the thickness of the claw portion provided on the finger portion 13 may be less than the thickness of the finger portion 13. Also, in this specification, multiple finger portions 13 can be interpreted as multiple claw portions.
[0021] Figure 2 is a schematic diagram showing an example of the inside of the rack 20. As shown in Figures 1 and 2, an opening 22 is provided on the surface of the rack 20. The substrate 50 is housed in the rack 20 through the opening 22. The rack 20 is placed on a workbench or the like, for example, with the opening 22 facing upwards, as shown in Figure 1.
[0022] The rack 20 has a plurality of slots 21 arranged in a row. Each of the plurality of slots 21 can accommodate one substrate 50. Here, the direction in which the plurality of objects 50 are arranged in the rack 20 is defined as the vertical direction, and the direction perpendicular to the vertical direction and the depth direction of the rack 20 is defined as the left-right direction. Each of the plurality of slots 21 may have grooves extending in the depth direction located on the left and right inner wall surfaces of the rack 20. Also, each of the plurality of slots 21 may have grooves extending in the left-right direction located on the bottom surface of the rack 20. Furthermore, each of the plurality of slots 21 may have grooves located on the bottom surface and inner wall surface of the rack 20. In addition, grooves may be formed by each of the plurality of slots 21 having a strip-shaped protrusion located on the inner wall surface or bottom surface of the rack 20. Note that each of the plurality of slots 21 may have a partition plate that separates each of the plurality of slots 21. On the other hand, in each of the plurality of slots 21, the plurality of protrusions located on the left and right inner wall surfaces may be spaced apart from each other. In other words, the multiple slots 21 may be arranged such that the multiple recesses are aligned along the vertical direction of the rack 20. As a result, it becomes easier to insert the multiple finger portions 13.
[0023] The thickness or width of the protrusions may be greater than the thickness of each of the multiple flattened finger portions 13 or each of the multiple claw portions. Alternatively, the spacing between multiple protrusions (width of the groove), or the distance between the vertically opposing protrusions of the rack 20 and the inner wall surface of the rack 20, may be less than the thickness of the protrusions. Furthermore, the depth of the rack 20 may be less than the length of the object 50. In this case, the tip of the object 50 will protrude from the opening of the rack 20. On the other hand, the depth of the rack 20 may be equal to or greater than the length of the substrate 50. The depth of the rack 20 may be less than twice the length of the flattened portions of the multiple finger portions 13 or multiple claw portions.
[0024] Multiple substrates 50 housed in multiple slots 21 are arranged in a row such that the longitudinal directions of the sides of the substrates 50 are parallel to each other. Also, when viewed from above the rack 20, the multiple substrates 50 are arranged in a row such that the longitudinal directions of the sides of the substrates 50 intersect the longitudinal direction of the rack 20. The substrate 50 has two opposing faces 52 in the thickness direction. Faces 52 can also be called main faces 52. Within the rack 20, one main face 52 of two adjacent substrates 50 faces each other. One side face 51 of each substrate 50 housed in the rack 20 (specifically, one of the four side faces 51 of a rectangular substrate 50) is exposed through the opening 22. The side face of the substrate 50 can also be called, for example, the end face of the substrate 50. Because the substrate 50 is thin, the side face of the substrate 50 is, for example, an elongated rectangle. The side surface of the substrate 50 can be described as, for example, a rectangle that is long in one direction, a long rectangle, or a strip or linear shape. Within the rack 20, the substrate 50 is housed with the side surface 51 exposed from the opening 22 facing upwards. The multiple substrates 50 housed in the multiple slots 21 may be of the same type or of different types. For example, the multiple substrates 50 may have different wiring patterns, different mounted components, or different sizes. Furthermore, the multiple substrates 50 may be of different types on a per-slot basis. That is, the first slot may house multiple substrates 50 of the first type, and the second slot may house multiple substrates 50 of the second type.
[0025] Furthermore, among the multiple surfaces constituting the substrate 50, the surfaces on which electronic components will be later mounted and the surfaces opposite to them may be referred to as multiple main surfaces. The multiple main surfaces may be the surfaces with the largest area among the multiple surfaces constituting the substrate 50. Electronic components may be later mounted on both of the multiple main surfaces. Also, among the multiple surfaces constituting the substrate 50, the surfaces connecting the multiple main surfaces may be referred to as side surfaces. In this case, the substrate 50 will have multiple side surfaces, and if the main surfaces are rectangular, it will have four side surfaces. The height of the side surfaces (in other words, the thickness of the substrate 50) is, for example, about 1 mm.
[0026] <Example of robot operation> The robot 10 can grasp multiple objects 50 one by one under the control of the control device 1. The robot 10 can also move the grasped objects 50 under the control of the control device 1. Specifically, the robot 10 grasps the circuit boards 50 in the rack 20 one by one with multiple fingers 13 and moves the grasped circuit boards 50 one by one to another location. The destination of the circuit boards 50 can be anything, for example, a tray, a jig, or a conveyor belt. The robot 10 moves the circuit boards 50 grasped by the multiple fingers 13 to another location by changing the posture of the arm 11. Then, the robot 10 releases the grip of the objects 50 from the multiple fingers 13 and places the objects 50 on the jig or the other at the destination.
[0027] The robot 10 uses its multiple fingers 13 to grasp one substrate 50 at a time, starting from the outermost substrate 50, and moves them to another location. When grasping a substrate 50 in the rack 20, the multiple fingers 13 insert the fingers 13 into the rack 20 from the direction of the side surface 51 exposed from the opening 22 of the rack 20, and press the fingers 13 against the two main surfaces 52 to grasp the substrate 50. In other words, the robot 10 can grasp a substrate 50 by sandwiching it between its multiple fingers 13.
[0028] Furthermore, even if the multiple circuit boards 50 housed in the multiple slots 21 are of different types, the insertion starting positions of the multiple finger portions 13 may be the same, and the insertion depth may be the same. Alternatively, when removing the circuit boards 50 from the multiple slots 21, the insertion starting positions of the multiple finger portions 13 may be the same, and the insertion depth may be the same, but as described later, when regripping the circuit boards 50, the gripping position may be changed according to the type. On the other hand, if the multiple circuit boards 50 housed in the multiple slots 21 are of different types, the insertion starting positions of the multiple finger portions 13 may be changed, and the insertion depth may be changed, according to the type of circuit board 50.
[0029] Figures 3 to 6 are schematic diagrams showing an example of the operation of the robot 10. As shown in Figure 3, the robot 10 grasps the outermost substrate 50 of the multiple substrates 50 in the rack 20 with multiple fingers 13. At this time, the multiple fingers 13 grasp the two main surfaces 52 of the substrate 50 from the outside. Then, the robot 10 changes the posture of the arm 11 and moves the substrate 50 grasped by the multiple fingers 13 upward, as shown in Figures 4 and 5, to the outside of the rack 20 through the opening 22 of the rack 20. After that, the robot 10 changes the posture of the arm 11 and moves the substrate 50 grasped by the multiple fingers 13 to another location.
[0030] Next, as shown in Figure 6, the robot 10 grasps the second circuit board 50 from the end of the rack 20 with its multiple fingers 13. Then, the robot 10 moves the circuit board 50 grasped by the multiple fingers 13 to another location. Subsequently, the robot 10 operates in the same manner, grasping the remaining circuit boards 50 in the rack 20 one by one and moving them to other locations.
[0031] Furthermore, the held object 50 may be held again after it has been held once. That is, the object 50 may be held again by another robot, or the held object 50 may be released and then held again. In this case, for example, a jig or the like may be used to hold it again.
[0032] Specifically, after the substrate 50 is removed from the rack 20 by the end effector 12, it is placed on a jig fixed at a predetermined position within the working range of the robot 10, and the end effector 12 may readjust the substrate 50 to maintain an appropriate position in preparation for a subsequent process. For example, multiple substrates 50 housed in the rack 20 are housed with some degree of variation, but by being placed once on the jig at a fixed position, the size or coordinates of the substrate 50 can be accurately determined, enabling gripping at an appropriate gripping position. The jig fixed at a predetermined position may be a frame-shaped jig having two grooves along two sides of the substrate 50 that intersect at right angles, for example, if the substrate 50 is rectangular. In this case, by inserting the substrate 50 into the two intersecting grooves of the jig, the coordinates of the substrate 50 can be accurately determined. In this case, the position at which the substrate is readjusted may be changed depending on the size or type of the substrate 50, or the subsequent process may be changed, such as by changing the placement location.
[0033] Furthermore, if the jig is frame-shaped, when the substrate 50 is inserted into the groove of the jig, the edges of the substrate 50 are supported by the jig, so the main surface of the substrate 50 is exposed. In this case, for example, the main surface of the substrate 50 can be photographed with the camera 15 to inspect the substrate 50. Also, in this case, the subsequent process may be changed depending on the wiring pattern or the type of mounted components on the main surface of the substrate 50.
[0034] Furthermore, the openings of the grooves in the jig may face upwards so that the two orthogonally intersecting grooves of the jig are located at the lowest point in the direction of gravity. In this case, the end effector 12, which has lifted the substrate 50 from the rack 20, changes its orientation so that the sides of the substrate 50 are inserted into the two grooves, and places the substrate 50 into the jig from above. Then, based on the position of the jig, the end effector 12 moves to the appropriate position and performs the substrate 50 holding operation, thereby holding the substrate 50 in the appropriate position.
[0035] The robot 10 may stop its operation once it has completed its work on multiple objects 50. Specifically, the robot 10 may stop its operation after removing all of the multiple circuit boards 50 that it has determined to be graspable from the rack 20 and completing the subsequent work. The subsequent work refers to placing the circuit boards 50 in another location, loading them into another device, or inspecting them. The robot system 100, the control device 1, or the robot 10 may also notify the robot to replace the rack 20 after it has completed its work on the multiple circuit boards 50 in the rack 20. The robot may then stop its operation once it has completed its work on at least one of the racks 20 that requires work. The robot 10 may also stop its operation once there are no more recognizable objects 50 left. Specifically, the robot 10 may stop its operation once there are no more recognizable circuit boards 50 left in the rack 20.
[0036] <Example of Control Device Configuration> Figure 7 is a schematic diagram showing an example of the configuration of the control device 1. As shown in Figure 7, the control device 1 includes, for example, a control unit 2, a storage unit 3, an interface 4, and an interface 5. The control device 1 is, for example, a computer device. The control device 1 can also be called, for example, a control circuit.
[0037] Interface 4 can communicate with camera 15. Interface 4 acquires camera images 150 generated by camera 15 from camera 15. Control unit 2 can acquire camera images 150 through interface 5. Interface 4 can also be called, for example, an interface circuit, a communication unit, or a communication circuit. Interface 4 may communicate with camera 15 via wired communication or wireless communication.
[0038] Interface 5 can communicate with the robot 10. The control unit 2 can control the robot 10 through interface 5. Interface 5 can also be called, for example, an interface circuit, a communication unit, or a communication circuit. Interface 5 may communicate with the robot 10 via wired communication or wireless communication.
[0039] The control unit 2 can comprehensively manage the operation of the control device 1 by controlling other components of the control device 1. The control unit 2 can also be called, for example, a control circuit. The control unit 2 includes at least one processor to provide control and processing capabilities for performing various functions, as will be described in more detail below.
[0040] According to various embodiments, at least one processor may be implemented as a single integrated circuit (IC) or as a plurality of communicably connected integrated circuits IC and / or discrete circuits. At least one processor can be implemented according to various known techniques.
[0041] In one embodiment, the processor includes one or more circuits or units configured to perform one or more data computation procedures or processes by, for example, executing instructions stored in associated memory. In other embodiments, the processor may be firmware (e.g., discrete logic components) configured to perform one or more data computation procedures or processes.
[0042] According to various embodiments, the processor may include one or more processors, controllers, microprocessors, microcontrollers, application-specific integrated circuits (ASICs), digital signal processing devices, programmable logic devices, field-programmable gate arrays, or any combination of these devices or configurations, or other known combinations of devices and configurations, and may perform the functions described below.
[0043] The control unit 2 may include, for example, a CPU (Central Processing Unit) as a processor. The control unit 2 can control the arm 11 and the end effector 12 through the interface 5. The control unit 2 can control the posture of the arm 11 through the interface 5. The control unit 2 can control the opening and closing of the plurality of finger portions 13 of the end effector 12 through the interface 5. By controlling the posture of the arm 11, the control unit 2 can control the position and posture of the plurality of finger portions 13. By controlling the posture of the arm 11, the control unit 2 can control the position and posture of the camera 15 fixed to the end effector 12.
[0044] The storage unit 3 may include a non-temporary recording medium readable by the CPU of the control unit 2, such as a ROM (Read Only Memory) and a RAM (Random Access Memory). For example, a program 3a for controlling the control device 1 is stored in the storage unit 3. Various functions of the control unit 2 are realized, for example, when the CPU of the control unit 2 executes the program 3a in the storage unit 3.
[0045] The configuration of the control device 1 is not limited to the above example. For example, the control unit 2 may include a plurality of CPUs. Also, the control unit 2 may include at least one DSP (Digital Signal Processor). Further, all or some of the functions of the control unit 2 may be realized by a hardware circuit that does not require software for realizing the functions. Also, the storage unit 3 may include a non-temporary recording medium readable by a computer other than a ROM and a RAM. The storage unit 3 may include, for example, a small hard disk drive and an SSD (Solid State Drive).
[0046] In addition, the control device 1 may include a display unit controlled by the control unit 2. The display unit may be, for example, a liquid crystal display, an organic EL (electro-luminescence) display, or a plasma display. The display unit of the control device 1 may display, for example, the camera image 150 generated by the camera 15. Also, the display unit of the control device 1 may display various recognition results of the substrate 50.
[0047] In addition, the control device 1 may include an input unit that receives input from the user. The input unit may include, for example, a mouse and a keyboard. Also, the input unit may include a touch sensor that receives the user's touch operation. Also, the input unit may include a microphone that receives the user's voice input.
[0048] In addition, the control device 1 may be composed of a plurality of computer devices. Also, the control device 1 may be a cloud server. In this case, the interface 5 of the control device 1 may communicate with the robot 10 through a network including the Internet or the like.
[0049] <Example of the operation of the control unit> Before starting the work, the control device 1 causes the camera 15 to photograph the object 50. Specifically, before the robot 10 starts gripping the substrate 50 in the rack 20, the camera 15 photographs the plurality of substrates 50 in the rack 20 from above the opening 22 of the rack 20. Note that the camera 15 may photograph one substrate out of the plurality of substrates 50 in the rack 20, or may photograph a plurality of substrates constituting a part of the plurality of substrates 50 in the rack 20. The control unit 2 controls the position and orientation of the camera 15 so that the camera 15 can photograph the plurality of substrates 50 in the rack 20 from above the opening 22 of the rack 20. The camera 15 photographs the plurality of substrates 50 in the rack 20 from the opening 22 side of the rack 20.
[0050] The control device 1 may also photograph the object 50 for each operation. Specifically, the control device 1 may move the camera 15 based on the position of each of the multiple slots 21 for each gripping operation to photograph the object 50. That is, after gripping an object 50 at the end of the rack 20, the control device 1 may move by the size of one of the multiple slots 21 to capture an image of the next object 50. When performing object recognition using the acquired image, the object recognition process may be performed on at least one object 50 located in the center of the image.
[0051] The control unit 2 controls the robot 10 based on the color image and / or depth image of the object 50 generated by the camera 15. The control unit 2 then causes the robot 10 to perform a task on the object 50. Specifically, the control unit 2 controls the robot 10 based on the camera image 150 showing multiple circuit boards 50 in the rack 20, causing the robot 10 to grasp the circuit boards 50 in the rack 20 one by one and move them to another location, as shown in Figures 3 to 6 above. Hereafter, when simply referred to as the camera image 150, it means the camera image 150 showing multiple circuit boards 50 in the rack 20.
[0052] The camera image 150 shows the side view 51 of each circuit board 50 in the rack 20. The camera image 150 includes a side view image 151 of each circuit board 50, showing the side view 51 of that circuit board 50.
[0053] Figure 8 is a schematic diagram showing an example of a camera image 150. As shown in Figure 8, in the camera image 150, the side surface 51 of each substrate 50 is captured such that its longitudinal direction is parallel to, for example, the x-direction (left-right direction in Figure 8) of the image coordinate system of the camera image 150. In other words, in the camera image 150, the side surface 51 of each substrate 50 is captured such that its longitudinal direction is aligned with the x-direction of the image coordinate system of the camera image 150. Furthermore, in the camera image 150, the side surfaces 51 of multiple substrates 50 are captured so that they are lined up in a row along the y-direction (up-down direction in Figure 8) of the image coordinate system of the camera image 150. The longitudinal direction of each side surface image 151 included in the camera image 150 is parallel to the x-direction of the image coordinate system of the camera image 150.
[0054] Hereafter, the x-direction of the image coordinate system of an image such as camera image 150 will be referred to as the image x-direction. Similarly, the y-direction of the image coordinate system of an image such as camera image 150 will be referred to as the image y-direction. The image x-direction can also be called the row direction of the image, and the image y-direction can also be called the column direction of the image.
[0055] The control unit 2 controls the position and orientation of the camera 15 so that the longitudinal direction of the side surface 51 of each substrate 50 is captured parallel to the x-direction of the image in the camera image 150. Based on the camera image 150, the control unit 2 recognizes the side surface 51 of each substrate 50 in the rack 20. The side surface 51 can be described as an elongated recognition target. In this specification, the side surface 51 can be described as, for example, an elongated recognition target, an elongated recognition target, a strip-shaped recognition target, or a linear recognition target. Also in this specification, the recognition target can be described as the recognition target object.
[0056] The side view image 151 in the camera image 150, which shows the side view 51, can also be said to be, for example, the recognition target image in the camera image 150, which shows the recognition target. Based on the recognition result of the side view 51 of the substrate 50, the control unit 2 causes the robot 10 to hold the substrate 50. Based on the recognition result of the side view 51 of the substrate 50, the control unit 2 causes the robot 10's multiple fingers 13 to grasp the two main surfaces 52 of the substrate 50. The recognition of the side view 51 in the camera image 150 can also be said to be the recognition or identification of the side view image 151 included in the camera image 150.
[0057] <Example of a method for recognizing the side surface of a substrate> The control unit 2 uses, for example, a trained model 120, which is a trained learning model, to recognize the long portion of the object 50 that appears in the camera image 150. In this embodiment, the control unit 2 uses the trained model 120 to recognize the side surface 51 of each substrate 50 that appears in the camera image 150. The learning model of the control unit 2 is composed of, for example, a neural network that realizes instance segmentation. The learning model of the control unit 2 is composed of, for example, a Mask Scoring R-CNN. R-CNN is an abbreviation for Region Based Convolutional Neural Networks. The learning model can also be called a machine learning model. The trained model 120 is generated when the original learning model undergoes machine learning processing, i.e., is trained. The original learning model may be a pre-training model, which is a learning model before training, or a trained learning model. If the original learning model is a trained model, the trained model 120 is generated when the original learning model is retrained.
[0058] As shown in Figure 9, the trained model 120 in the control unit 2 receives a camera image 150, which is a color image and / or depth image, as an inference image 150. Based on the input inference image 150, the trained model 120 recognizes the long portion of the object 50 that is the target of recognition. Specifically, based on the inference image 150, the trained model 120 recognizes the side surface 51 of each substrate 50 in the rack 20. In this embodiment, the trained model 120 can also be said to be a recognition unit that recognizes the side surface 51 of the substrate 50. The trained model 120 outputs a recognition result 130 of the target of recognition.
[0059] The recognition result 130 output by the trained model 120 includes, for example, the mask and bounding box of a long recognition target part, such as the side surface 51 of each circuit board 50 in the rack 20. Specifically, the recognition result 130 includes positional and shape information of the mask and bounding box of the side surface 51 of each circuit board 50. The position and shape of the mask and bounding box of the side surface 51 represent the position and shape of the side surface 51. It can also be said that the trained model 120 recognizes the position and shape of the side surface 51 of the circuit board 50.
[0060] Furthermore, the recognition result 130 includes, for example, a classification score for each bounding box that represents the accuracy of the classification of the region within that bounding box (in other words, the level of confidence in the classification, the likelihood of the classification, or the certainty of the classification). The classification score indicates the likelihood that the region within the bounding box is the recognition target part (e.g., side 51). The classification score is, for example, a numerical value between 0 and 1. The larger the classification score, the higher the likelihood that the region within the bounding box is the recognition target part (e.g., side 51). The score can also be called an evaluation value.
[0061] Furthermore, the recognition result 130 includes, for example, a mask score for each mask generated by the trained model 120, which indicates the accuracy of the mask generation. The mask score is, for example, a numerical value between 0 and 1. The larger the mask score, the higher the accuracy of the mask generation. The mask score can be said to indicate the level of confidence in the positional and shape information of the mask generated by the trained model 120, or it can be said to indicate the likelihood of the positional and shape information of the mask, or it can be said to indicate the certainty of the positional and shape information of the mask.
[0062] The control unit 2 causes the robot 10 to grasp the object 50 based on the recognition result 130 output by the trained model 120. In this embodiment, the control unit 2 causes each of the circuit boards 50 in the rack 20 to grasp based on the recognition result 130. Here, the circuit board 50 being described may be called the target circuit board 50. The target circuit board 50 can also be said to be the circuit board 50 of interest. The control unit 2 causes the robot 10 to grasp the target circuit board 50 based on the recognition result of the side surface (also called the target side surface) of the target circuit board 50 included in the recognition result 130.
[0063] The control unit 2 may, for example, cause the robot 10 to grasp the target substrate 50 based on the bounding box of the target side included in the recognition result 130. In this case, the control unit 2 identifies, for example, a position in real space (in other words, the actual work space in which the robot 10 performs the work) corresponding to the center of the bounding box of the target side as the insertion start position for the multiple finger parts 13. The control unit 2 then controls the posture of the arm 11 so that the reference positions of two finger parts 13 are located vertically above the identified position. The reference positions of the two finger parts 13 are also called, for example, TCP (Tool Center Point). The reference positions of the two finger parts 13 are set, for example, at the midpoint between the tips of the two finger parts 13. Subsequently, the control unit 2 moves the two finger parts 13 vertically downward with the two finger parts 13 open, and then closes the two finger parts 13 to cause the target substrate 50 to be grasped by the two finger parts 13.
[0064] In the above example, the control unit 2 sets the insertion start positions of the multiple finger portions 13 based on the bounding box, but they may also be set based on a mask. In this case, the control unit 2 may, for example, use the position in real space corresponding to the center of the mask as the insertion start position. Alternatively, for each of the multiple slots 21, the position of the center of the slot 21 may be set as a fixed value for the insertion start position of the multiple finger portions 13. Alternatively, for each of the multiple slots 21, a position closer to the left and right inner walls of the rack 20 than the center of the slot 21 may be set as a fixed value for the insertion start position of the multiple finger portions 13.
[0065] The opening width of the multiple finger portions 13 may be set to be equal to the length of the mask. Alternatively, the opening width of the multiple finger portions 13 may be set to be equal to the length of the bounding box. Alternatively, the opening width of the multiple finger portions 13 may be set based on the size of each of the multiple slots 21. Alternatively, the opening width of the multiple finger portions 13 may be set to the spacing between the multiple protrusions that each of the multiple slots 21 has. Furthermore, the opening width of the multiple finger portions 13 may be changed depending on the type of object 50, or it may be a fixed value.
[0066] The insertion depth of the multiple finger portions 13 may be set according to the shape or size of the mask. Alternatively, the insertion depth of the multiple finger portions 13 may be set according to the shape or size of the bounding box. Alternatively, the insertion depth of the multiple finger portions 13 may be set based on the size of the rack 20. Alternatively, the insertion depth of the multiple finger portions 13 may be set to half the depth of the rack 20. Alternatively, the insertion depth of the multiple finger portions 13 may be changed according to the type of object 50, or it may be a fixed value. Alternatively, the insertion depth of the multiple finger portions 13 may be set based on the depth image.
[0067] The control unit 2 may, for example, cause the robot 10 to grasp the target substrate 50 based on the mask of the target side included in the recognition result 130. In this case, the control unit 2 sets the pixel values of all pixels in the inference image 150 except for the area at the same position as the mask of the target side to zero. That is, the control unit 2 sets the pixel values of the background area, which is the area other than the target side, to zero in the inference image 150. Then, the control unit 2 determines the gripping position in which the robot 10 grasps the target substrate 50 based on the inference image 150 (also called the modified inference image 150) in which the pixel values of the background area are set to zero. The control unit 2 may, for example, have a second learning model that infers the gripping position based on the modified inference image 150. The control unit 2 controls the arm 11 and the end effector 12 so that the robot 10 grasps the substrate 50 at the gripping position inferred by the second learning model.
[0068] The control unit 2 may also identify a position in real space corresponding to the center of the mask on the target side, and control the posture of the arm 11 so that the reference positions of the two finger portions 13 are located vertically above the identified position. The control unit 2 may then move the two finger portions 13 vertically downward to grip the target substrate 50 with the two finger portions 13.
[0069] Furthermore, the control unit 2 may determine the gripping position in which the robot 10 grips the target substrate 50 based on the mask and bounding box of the target side included in the recognition result 130.
[0070] The control unit 2 may determine which of the multiple substrates 50 in the rack 20 to allow the robot 10 to grasp, based on the classification score and mask score for the side surface 51 of each substrate 50 included in the recognition result 130. For example, the control unit 2 calculates a total score (also called a confidence score) for each substrate 50 by multiplying the classification score and mask score for the side surface 51 of the substrate 50 together. Then, the control unit 2 allows the robot 10 to grasp only the substrates 50 in the rack 20 whose total score is equal to or greater than a threshold (e.g., 0.5).
[0071] <Example of a Learning Method> Next, we will explain an example of a method for training a learning model to generate a trained model 120. Hereafter, the original learning model from which the trained model 120 was trained will be referred to as the learning model 220. The trained model 120 is generated when the learning model 220 is trained.
[0072] Figure 10 is a schematic diagram illustrating an example of a learning method for the learning model 220. The learning of the learning model 220 may be performed in the control unit 2 of the control device 1, or in a device other than the control device 1 (for example, a cloud server). In other words, the generation of the trained model 120 may be performed in the control device 1, or in a device other than the control device 1. The device that trains the learning model 220 can also be called a learning device. If the trained model 120 is generated in a device other than the control device 1, the control device 1 can use the trained model 120 once it is installed in the control device 1.
[0073] The learning device has a learning model 220. The learning device also stores learning data 250 used to train the learning model 220. If the control device 1 is the learning device, for example, the control unit 2 has the learning model 220 and the storage unit 3 stores the learning data 250. When the control unit 2 executes the program 3a in the storage unit 3, a learning unit is formed as a functional block in the control unit 2 that trains the learning model 220 and generates a trained model 120.
[0074] The learning model 220 is trained based on the training data 250. The training data 250 includes multiple datasets consisting of training images 260 and corresponding annotation data 270.
[0075] The training data 250 may include, for example, training images 260 which are computer graphics (CG) images. CG images can also be called composite images. Hereafter, the training images 260 which are CG images may be referred to as training CG images 260.
[0076] Figure 11 is a schematic diagram showing an example of a training CG image 260. As shown in Figure 11, the training CG image 260 shows, for example, only a plurality of circuit boards 50. The training CG image 260 shows the side surfaces 51 of the plurality of circuit boards 50. The circuit boards 50 shown in the training CG image 260 are not real objects, but CG images. The training data 250 includes, for example, a plurality of training CG images 260 in which the number of circuit boards 50 shown differs from one another. In addition, the training CG image 260 may have patterns on the circuit boards 50 that are not present on the real objects.
[0077] For example, 3DCG technology is used to generate the training CG image 260. The training CG image 260 is generated on a computer. 3DCG is an abbreviation for 3-Dimensional Computer Graphics. When the training CG image 260 is generated, 3D CAD data of the circuit board 50 is used. CAD is an abbreviation for Computer Aided Design. Then, based on the 3D CAD data of the circuit board 50, 3DCG technology is used to model the circuit board 50, and a 3DCG model of the circuit board 50 (also called a circuit board model) is generated. Next, multiple circuit board models are arranged in a virtual 3D space. After that, the virtual image obtained by photographing the arranged multiple circuit board models from above with a virtual camera is used as the training CG image 260. However, the method of generating the training CG image 260 is not limited to this. In this example, an image showing multiple circuit boards 50 is prepared as the learning CG image 260, but an image showing only one circuit board 50 may also be prepared as the learning CG image 260. Furthermore, an object corresponding to a rack may be included in the learning CG image 260.
[0078] Thus, when the training images 260 are computer-generated images, it becomes easier to generate a variety of training images 260. This makes it easy to improve the training accuracy of the learning model 220. The training data 250 may also include camera images 150 (in other words, real camera images) as training images 260.
[0079] In training images 260, such as the training CG image 260, the side surfaces 51 of each substrate 50 are depicted so that their longitudinal direction is parallel to the image x-direction, similar to the inference image 150. Also, in training images 260, the side surfaces 51 of multiple substrates 50 are depicted so that they are aligned in a line along the image y-direction. A side view image 261 (see Figure 11) in which the side surfaces 51 are depicted in training images 260 can also be said to be, for example, an image of the recognition target in which the recognition target is depicted in training images 260. The longitudinal direction of each side view image 261 included in training images 260 is parallel to the image x-direction.
[0080] Furthermore, the training data 250 may include training images 260 in which the sides 51 of the multiple circuit boards 50 depicted are not aligned in a line along the y-direction of the image. Also, the number of multiple circuit boards 50 depicted in the training image 260 may or may not match the number of multiple circuit boards 50 depicted in the inference image 150.
[0081] The annotation data 270 includes masks, bounding boxes, and classifications for each side 51 of each substrate 50 (in other words, each substrate model) as seen in the corresponding training image 260. The masks and bounding boxes included in the annotation data 270 are set for each side image 261 (in other words, each image of the part to be recognized) included in the training image 260. Since the training CG image 260 is generated on a computer, the annotation data 270 corresponding to the training CG image 260 can also be easily generated on a computer.
[0082] As described above, the trained model 120 is generated when the trained model 220 is trained based on training data 250, which includes a training image 260 showing the side surface 51 of the substrate 50 (in other words, the part to be recognized) and annotation data 270 relating to the side surface 51 shown in the training image 260.
[0083] Hereafter, unless there is a need to specifically distinguish between the trained model 220 and the pre-trained model 120, they will simply be referred to as the trained model. Furthermore, the masks included in the annotation data 270 may be referred to as annotation masks. Also, the bounding boxes included in the annotation data 270 may be referred to as annotation BBOXes.
[0084] <Example of Annotation Mask and Annotation BBOX> Figure 12 is a schematic diagram showing an example of an annotation mask 280 and an annotation BBOX 290. Figure 12 shows an example of a side view image 261 included in the training image 260, and the mask 280 and annotation BBOX 290 set on the side view image 261. In Figure 12, the side view image 261 is shown with a dashed line. Also in Figure 12, the annotation BBOX 290 is shown with a relatively thick line, and the annotation mask 280 is shown with diagonal lines.
[0085] The annotation mask 280 has a shape corresponding to the side image 261. The annotation mask 280, for example, is a rectangle that is elongated in one direction, similar to the side image 261. The longitudinal direction of the annotation mask 280 set on the side image 261 is parallel to the image x direction (left-right direction in Figure 12), similar to the side image 261.
[0086] The annotation mask 280 includes, for example, the side image 261 and is set to be larger than the side image 261. The size of the annotation mask 280 in the short direction (up and down direction in Figure 12) is, for example, larger than the size of the side image 261 in the short direction.
[0087] The annotation mask 280 has two sides 280a and 280b that face each other in the short direction of the annotation mask 280. The side image 261 has two sides 261a and 261b that face each other in the short direction of the side image 261. The two sides 280a and 280b of the annotation mask 280 face the two sides 261a and 261b of the side image 261, respectively. The two sides 280a and 280b of the annotation mask 280 and the two sides 261a and 261b of the side image 261 are parallel to each other in the x direction of the image.
[0088] Here, the distance between one side 280a of the annotation mask 280 along its longitudinal direction and one side 261a of the side image 261 along its longitudinal direction is called distance d1. Distance d1 is the distance along the image y-direction between side 280a and side 261a. Also, the distance between the other side 280b of the annotation mask 280 along its longitudinal direction and the other side 261b of the side image 261 along its longitudinal direction is called distance d2. Distance d2 is the distance along the image y-direction between side 280b and side 261b. In this example, distance d1 and distance d2 are equal.
[0089] For example, consider the case where the size (in other words, the number of pixels) of the side image 261 in the short direction is 4 pixels. In this case, the size (in other words, the number of pixels) of the annotation mask 280 in the short direction may be, for example, 6 pixels, 8 pixels, or 10 pixels. For example, if the size of the annotation mask 280 in the short direction is 6 pixels, then distances d1 and d2 will each be 1 pixel. Also, if the size of the annotation mask 280 in the short direction is 10 pixels, then distances d1 and d2 will each be 3 pixels.
[0090] The longitudinal size of the annotation mask 280 (the left-right direction in Figure 12) matches, for example, the longitudinal size of the side image 261. Both longitudinal ends (in other words, both sides) of the annotation mask 280 are in contact with both longitudinal ends of the side image 261, respectively.
[0091] The frame-shaped annotation BBOX 290, like the side image 261 and annotation mask 280, forms a rectangle that is elongated in one direction. The longitudinal direction of the annotation BBOX 290 is set parallel to the x-direction of the image, and the short direction of the annotation BBOX 290 is set parallel to the y-direction of the image.
[0092] Annotation BBOX 290 includes, for example, a side image 261 and an annotation mask 280, and is set to be larger than the side image 261 and the annotation mask 280. The size of the annotation BBOX 290 in the short direction is larger than, for example, the size of the side image 261 and the annotation mask 280 in the short direction.
[0093] The annotation BBOX 290 has two sides 290a and 290b that face each other in the short direction of the annotation BBOX 290. One side 290a of the annotation BBOX 290 faces one side 261a of the side image 261 and one side 280a of the annotation mask 280. The other side 290b of the annotation BBOX 290 faces the other side 261b of the side image 261 and the other side 280b of the annotation mask 280. The two sides 290a and 290b of the annotation BBOX 290 are parallel to the x-direction of the image.
[0094] Here, the distance between one side 290a of the annotation BBOX 290 and one side 261a of the side image 261 is called distance d11. Distance d11 is the distance along the image y-direction between side 290a and side 261a. Also, the distance between the other side 290b of the annotation BBOX 290 and the other side 261b of the side image 261 is called distance d12. Distance d12 is the distance along the image y-direction between side 290b and side 261b. In this example, distance d11 and distance d12 are equal.
[0095] For example, consider the case where the size of the side image 261 in the short direction is 4 pixels. In this case, the size of the annotation BBOX 290 in the short direction may be, for example, 14 pixels or 22 pixels. For example, if the size of the annotation BBOX 290 in the short direction is 14 pixels, then distances d11 and d12 will each be 5 pixels. Also, if the size of the annotation mask 280 in the short direction is 22 pixels, then distances d11 and d12 will each be 9 pixels.
[0096] The longitudinal size of the annotation BBOX 290 matches, for example, the longitudinal size of the side image 261 and the annotation mask 280. Both longitudinal ends of the annotation BBOX 290 are in contact with both longitudinal ends of the side image 261, respectively. Also, both longitudinal ends of the annotation BBOX 290 are in contact with both longitudinal ends of the annotation mask 280, respectively.
[0097] As described above, in this example, the size of the annotation mask 280 in the short direction is set to be larger than the size of the side image 261 included in the training image 260 in the short direction. This makes it less likely for the mask (also called the inference mask) of the elongated side 51 of the substrate 50, which is included in the recognition result 130 output from the trained model 120, to be interrupted. In other words, the trained model 120 can appropriately recognize the side 51 that is visible in the inference image 150. Therefore, when the control unit 2 instructs the robot 10 to hold the substrate 50 based on the inference mask, the robot 10 is less likely to fail to hold the substrate 50.
[0098] For example, unlike this example, consider a case where the size and shape of the annotation mask 280 match the size and shape of the side image 261 included in the training image 260. Then, consider a case where, for example, the camera 15's orientation when acquiring the inference image 150 deviates from its original orientation, causing the side 51 of the substrate 50 to be unintentionally obliquely positioned in the inference image 150, so that its longitudinal direction is not parallel to the image x-direction. In such a case, the inference mask 135 of the obliquely positioned, elongated side 51 in the inference image 150 is likely to be interrupted midway, as shown in Figure 13. In Figure 13, the side image 151 included in the inference image 150 is shown with a dashed line, and the inference mask 135 of the side image 151 is shown with diagonal lines.
[0099] In contrast, in this example, the size of the annotation mask 280 in the short direction is set to be larger than the size of the side image 261 included in the training image 260 in the short direction. As a result, for example, even if the elongated side 51 of the substrate 50 is captured at an angle in the inference image 150, the inference mask 135 of that side 51 is less likely to be cut off.
[0100] Figure 14 is a schematic diagram showing an example of an inference mask 135 when the annotation mask 280 of this example is used. In Figure 14, the inference mask 135 for the side surface 51 of the substrate 50 is shown together with the side image 151 in which the side surface 51 is visible, when the side surface 51 of the substrate 50 is unintentionally shown at an angle in the inference image 150. In the example in Figure 14, the inference mask 135 is generated without interruption.
[0101] In this example, as shown in Figure 12, distances d1 and d2 coincide. That is, the annotation mask 280 is enlarged by the same amount on both sides in the short direction relative to the side image 261. Therefore, the inference mask is less likely to deviate from the side 51 that is captured in the inference image 150. Thus, when the control unit 2 instructs the robot 10 to hold the substrate 50 based on the inference mask, the robot 10 is less likely to fail to hold the substrate 50.
[0102] If the inference mask 135 contains portions whose size in the short direction is larger than the size in the short direction of the side image 261 included in the training image 260 (also called the side image size), the inference mask 135 may be modified so that the size in the short direction of those portions is reduced to the size of the side image. The control unit 2 may then cause the robot 10 to hold the substrate 50 based on the modified inference mask 135.
[0103] Furthermore, an annotation mask 280 whose size and shape match those of the side view image 261 may be used to train the learning model 220.
[0104] In this example, as described above, the annotation BBOX 290 is set to include the elongated side image 261 and to be larger than the side image 261. This allows the training model 220 to be properly trained. As a result, the trained model 120 can properly recognize the side 51 that appears in the inference image 150. This point will be explained below.
[0105] The learning model in this example, which is composed of Mask Scoring R-CNN, includes a Backbone network as the input layer, an RPN as the hidden layer, and a Head network as the output layer. RPN is an abbreviation for Region Proposal Network. During inference, the inference image 150 is input to the Backbone network, and the recognition result 130 is output from the Head network. In the training of the learning model 220, the Backbone network, RPN, and Head network are trained.
[0106] RPN is an anchor-based model. RPN uses rectangular anchor boxes. RPN moves the anchor boxes on the feature map output from the Backbone network. The Backbone network generates a feature map based on the inference image 150 during inference and on the training image 260 during training. The feature map is a type of image. The image size of the feature map is a reduced version of the image sizes of the inference image 150 and the training image 260. When the rectangular anchor box moves on the feature map as an image, one side of the anchor box is set parallel to the x-direction of the image, and the other side of the anchor box is set parallel to the y-direction of the image.
[0107] Regarding anchor boxes, for example, there are five different sizes and three different shapes, resulting in a total of 15 types of anchor boxes used in RPN. Figure 15 is a schematic diagram showing an example of three different shapes of anchor boxes.
[0108] As shown in Figure 15, anchor boxes come in three shapes: one with an aspect ratio of 1:1 (leftmost), one with an aspect ratio of 1:0.5 (center), and one with an aspect ratio of 1:2 (rightmost). Furthermore, there are five different vertical sizes for the anchor boxes.
[0109] In RPN training, for each of the 15 types of anchor boxes, the anchor box is moved by a predetermined number of pixels on the feature map. The amount of movement (in other words, the shift amount) of the anchor box increases as the vertical size of the anchor box increases. At each position of the anchor box that has been moved by a predetermined number of pixels on the feature map, the degree of overlap between the region within the anchor box and the region within the part corresponding to annotation BBOX290 on the feature map (the BBOX equivalent part) is calculated. The degree of overlap is calculated as IoU. IoU is an abbreviation for Intersection Over Union. IoU is a value obtained by dividing the area of the logical AND of the region within the anchor box and the region within the BBOX equivalent part by the area of the logical OR of the region within the anchor box and the region within the BBOX equivalent part. If the anchor box and the BBOX equivalent part on the feature map perfectly coincide, IoU is 1. In the training image 260, the region corresponding to the BBOX portion on the feature map becomes the annotation BBOX.
[0110] Here, a side view image 261 exists within the annotation BBOX. Therefore, in the feature map, within the BBOX-equivalent portion, there exists a region corresponding to the side view 51 of the substrate 50 (also called the side view equivalent region). The side view equivalent region can also be said to be the region corresponding to the side view image 261 included in the training image 260.
[0111] In RPN training, subregions within anchor boxes at positions where the IoU (Identification Unit) is greater than or equal to a predetermined value (e.g., 0.3) are determined to be regions where a corresponding lateral region exists (also called an existing region). Then, each existing region in the feature map obtained during RPN training is input to the Head network. The Head network is trained based on each existing region from the RPN.
[0112] Thus, in RPN training, a sub-region within an anchor box at a position in the feature map where the IoU is greater than or equal to a predetermined value is determined to be an existing region where a side-equivalent region exists. Furthermore, because the side-equivalent region in the feature map is elongated, none of the various types of anchor boxes used in RPN have an anchor box that is close in size to the side-equivalent region. As a result, if the annotation BBOX set for the side image 261, which shows the elongated side 51, is set to the same size as the side image 261, as in this example, there is a possibility that a sub-region within the anchor box will not be determined to be an existing region, even though a side-equivalent region actually exists within that anchor box. Hereafter, for the sake of explanation, the annotation BBOX in this example (see Figure 12) may be referred to as the first annotation BBOX, and the annotation BBOX set to the same size as the side image 261 may be referred to as the second annotation BBOX.
[0113] Figure 16 is a schematic diagram showing an example of the relationship between the anchor box 300 and the BBOX equivalent portion (also called the second BBOX equivalent portion 420) corresponding to the second annotation BBOX. In the example of Figure 16, since the second BBOX equivalent portion 420 exists within the anchor box 300, the side view equivalent region 500 exists within the anchor box 300. Since the second annotation BBOX has the same size as the side view image 261, the second BBOX equivalent portion 420 has the same size as the side view equivalent region 500. In the example of Figure 16, since the size of the second BBOX equivalent portion 420, which has the same size as the side view equivalent region 500, is significantly different from the size of the anchor box 300, the IoU does not become 0.3 or greater. Therefore, the subregion within the anchor box 300 in the feature map is not determined to be an existing region where the side view equivalent region 500 exists.
[0114] Figure 17 is a schematic diagram showing an example of the relationship between the anchor box 300 and the BBOX equivalent portion (also called the first BBOX equivalent portion 410) corresponding to the first annotation BBOX. Similar to the example in Figure 16, in the example in Figure 17, since the first BBOX equivalent portion 410 exists within the anchor box 300, the side equivalent region 500 exists within the anchor box 300. Since the size of the first BBOX equivalent portion 410 is larger than the size of the second BBOX equivalent portion 420, the size of the first BBOX equivalent portion 410 is close to the size of the anchor box 300. Therefore, the IoU is 0.3 or greater. Thus, the subregion within the anchor box 300 in the feature map is appropriately determined to be an existing region.
[0115] As described above, because the side-equivalent region 500 in the feature map is elongated, even if there is no anchor box with a size close to the size of the side-equivalent region 500 among the multiple types of anchor boxes used in RPN, by making the size of the annotation BBOX larger than the size of the side image 261, as in this example, an anchor box with a size close to the BBOX equivalent portion corresponding to the annotation BBOX will be available among the multiple types of anchor boxes. Therefore, even without changing the anchor box, if a side-equivalent region exists within the anchor box, the sub-region within that anchor box can be reliably determined to be an existing region. Thus, the Head network can be trained appropriately. In other words, the training model 220 can be trained appropriately. As a result, the recognition accuracy of the side 51 in the trained model 120 can be improved. Therefore, when the control unit 2 has the robot 10 hold the substrate 50 based on the recognition result of the side 51 of the substrate 50, the robot 10 is less likely to fail to hold the substrate 50.
[0116] Furthermore, in this example, as shown in Figure 12, distances d11 and d12 coincide. In other words, the annotation BBOX 290 is equally large on both sides in the short direction relative to the side image 261. Therefore, the bounding box (also called the inference bounding box) included in the recognition result 130 is less likely to deviate from the side 51 shown in the inference image 150. Thus, when the control unit 2 instructs the robot 10 to hold the substrate 50 based on the inference bounding box, the robot 10 is less likely to fail to hold the substrate 50.
[0117] In the example above, the longitudinal direction of the side surface 51 of the substrate 50 as seen in the training image 260 is parallel to the image x-direction, but it may also be oblique to the image x-direction. Figure 18 is a schematic diagram showing an example of the training image 260 in this case.
[0118] In the example shown in Figure 18, the longitudinal direction of the side view 51 in the training image 260 is inclined with respect to the image x-direction and is non-parallel to both the image x-direction and the image y-direction. Therefore, the longitudinal direction of the side view image 261 included in the training image 260 is inclined with respect to the image x-direction and is non-parallel to both the image x-direction and the image y-direction. The inclination angle of the longitudinal direction of the side view 51 in the training image 260 with respect to the image x-direction, in other words, the inclination angle of the longitudinal direction of the side view image 261 with respect to the image x-direction, is, for example, less than 45 degrees.
[0119] Figure 19 is a schematic diagram showing an example of an annotation mask 280 and an annotation BBOX 290 when the longitudinal direction of the side surface 51 of the substrate 50, as seen in the training image 260, is inclined with respect to the image x-direction.
[0120] In the example in Figure 19, the relationship between the annotation mask 280 and the side image 261 is the same as in the example in Figure 12 above. The longitudinal direction of the annotation mask 280 is inclined with respect to the image x direction, similar to the side image 261, and is non-parallel to the image x and y directions.
[0121] Annotation BBOX 290 is a rectangle that is elongated in one direction, similar to the example in Figure 12 above. Annotation BBOX 290 has two sides 290a and 290b that face each other in the short direction of Annotation BBOX 290, similar to the example in Figure 12. Annotation BBOX 290 also has two sides 290c and 290d that face each other in the long direction of Annotation BBOX 290. The two sides 290a and 290b of Annotation BBOX 290 are parallel to the x direction of the image. The two sides 290c and 290d of Annotation BBOX 290 are parallel to the y direction of the image.
[0122] Annotation BBOX 290 touches the four corners of the rectangular annotation mask 280. Edge 290a of Annotation BBOX 290 passes through the point of the annotation mask 280 having the minimum y coordinate. Edge 290b of Annotation BBOX 290 passes through the point of the annotation mask 280 having the maximum y coordinate. Edge 290c of Annotation BBOX 290 passes through the point of the annotation mask 280 having the minimum x coordinate. Edge 290d of Annotation BBOX 290 passes through the point of the annotation mask 280 having the maximum x coordinate. The longitudinal direction of the side view 51 shown in the training image 260, that is, the longitudinal direction of the side view image 261, is non-parallel to each of the four edges 290a, 290b, 290c, and 290d of Annotation BBOX 290. Hereafter, annotation BBOX290 shown in Figure 19 may be referred to as the third annotation BBOX290.
[0123] Figure 20 is a schematic diagram showing an example of the relationship between the anchor box 300 and the BBOX equivalent portion (also called the third BBOX equivalent portion 430) corresponding to the third annotation BBOX 290. As shown in Figure 20, the size of the third BBOX equivalent portion 430 is closer to the size of the anchor box 300 than the size of the second BBOX equivalent portion 420 (see Figure 16), similar to the size of the first BBOX equivalent portion 410 (see Figure 17). In the example in Figure 20, the IoU is 0.3 or greater, and the subregion within the anchor box 300 in the feature map is appropriately determined to be an existing region.
[0124] In this way, by tilting the longitudinal direction of the side image 261 from the image x direction, and making the longitudinal direction of the side image 261 non-parallel to each of the four sides of the annotation BBOX 290, an anchor box with a size close to the BBOX equivalent portion corresponding to the annotation BBOX will exist in multiple types of anchor boxes. Therefore, even without changing the anchor box, if a side-equivalent region exists within the anchor box, the sub-region within that anchor box can be reliably determined to be an existing region. Thus, the Head network can be trained appropriately. In other words, the learning model 220 can be trained appropriately.
[0125] If the side surface 51 of the substrate 50 is tilted with respect to the image x-direction in the training image 260, then an inference image 150 is used in which the side surface 51 of the substrate 50 is tilted with respect to the image x-direction in the same way as in the training image 260. Figure 21 is a schematic diagram showing an example of an inference image 150 in which the side surface 51 of the substrate 50 is tilted with respect to the image x-direction. The tilt angle of the side surface 51 with respect to the image x-direction in the inference image 150 is the same as the tilt angle of the side surface 51 with respect to the image x-direction in the training image 260. The control unit 2 controls the position and orientation of the camera 15 so that the side surface 51 of each substrate 50 is captured in the camera image 150, which is the inference image 150, as shown in Figure 21.
[0126] The control unit 2 may perform rotation processing or other operations on the camera image 150 captured as shown in Figure 8 above to generate the inference image 150 shown in Figure 21.
[0127] In the example above, the elongated recognition target was the side surface of the substrate 50, but it is not limited to this. For example, the elongated recognition target may be the side surface of an elongated rod-shaped member as the object 50 held by the robot 10, or it may be a string-like member such as a cord.
[0128] As described above, control devices and robotic systems have been explained in detail, but the above explanations are illustrative in all respects, and this disclosure is not limited thereto. Furthermore, the various examples described above can be combined and applied insofar as they do not contradict each other. And it is understood that countless examples not illustrated can be conceived without falling outside the scope of this disclosure.
[0129] This disclosure includes the following:
[0130] In one embodiment, (1) the learning method involves training a learned model to recognize an elongated recognition target portion depicted in an inference image based on an inference image, and generating the learned model based on a learning image and annotation data, wherein the annotation data includes a bounding box set for the recognition target portion image in the learning image, and the bounding box includes the recognition target portion image and is larger than the recognition target portion image.
[0131] (2) The learning method of (1) above, wherein the bounding box is in contact with both ends in the longitudinal direction of the recognition target image.
[0132] (3) The learning method of (2) above, wherein the recognition target image has a first side and a second side that are opposite to each other in the short direction of the recognition target image, the bounding box has a third side and a fourth side that are opposite to the first side and the second side, and the first distance between the first side and the third side and the second distance between the second side and the fourth side are equal to each other.
[0133] (4) The learning method of (1) above, wherein each of the four sides of the bounding box is non-parallel to the longitudinal direction of the recognition target image.
[0134] In one embodiment, (5) the learning method generates a trained model that recognizes an elongated recognition target portion depicted in an inference image based on the inference image, by training the trained model based on training data including a training image and annotation data, wherein the annotation data includes a mask set on the recognition target portion image in the training image in which the recognition target portion is depicted, and the size of the mask in the short direction is larger than the size of the recognition target portion image in the short direction.
[0135] (6) The learning method of (5) above, wherein the recognition target image has a first side and a second side that are opposite to each other in the short direction of the recognition target image, and the mask has a third side and a fourth side that are opposite to the first side and the second side, respectively, and the first distance between the first side and the third side and the second distance between the second side and the fourth side are equal to each other.
[0136] In one embodiment, (7) the learning device performs one of the learning methods described in (1) to (6) above.
[0137] In one embodiment, (8) the control device includes a control unit that recognizes the side surface of the plate-shaped member based on an image showing the side surface of the plate-shaped member, and causes the robot to hold the plate-shaped member based on the recognition result of the side surface.
[0138] (9) The control device of (8) above, wherein the robot is capable of grasping an object, the plate-shaped member has a first surface and a second surface facing each other in the thickness direction, and the control unit causes the robot to grasp the first surface and the second surface of the plate-shaped member.
[0139] In one embodiment, the (10) robot system comprises the control device described in (8) or (9) above, and a robot controlled by the control device.
[0140] In one embodiment, (11) the program is a program that causes a computer device to execute any one of the learning methods (1) to (6) described above.
[0141] In one embodiment, (12) the program is a program that causes the computer device to function as the control device described in (8) or (9) above.
[0142] 1 Control device (learning device) 2 Control unit 3a Program 10 Robot 50 Object 51 Side view (recognition target area) 52 Face 100 Robot system 120 Trained model 130 Recognition result 150 Inference image 250 Training data 260 Training image 261 Side view image (recognition target area image) 261a, 261b, 290a, 290b, 290c, 290d Edge 270 Annotation data 280 Mask 290 Bounding box
Claims
1. A learning method comprising: generating a trained model that recognizes an elongated object to be recognized in an inference image based on an inference image by training the trained model based on training data including a training image and annotation data, wherein the annotation data includes a bounding box set for the image of the object to be recognized in the training image, and the bounding box includes the image of the object to be recognized and is larger than the image of the object to be recognized.
2. A learning method according to claim 1, wherein the bounding box is in contact with both ends of the longitudinal direction of the recognition target image.
3. A learning method according to claim 2, wherein the recognition target image has a first side and a second side that are opposite to each other in the short direction of the recognition target image, the bounding box has a third side and a fourth side that are opposite to the first side and the second side, and the first distance between the first side and the third side and the second distance between the second side and the fourth side are equal to each other.
4. A learning method according to claim 1, wherein each of the four sides of the bounding box is non-parallel to the longitudinal direction of the recognition target image.
5. A learning method comprising: generating a trained model that recognizes an elongated object to be recognized in an inference image based on an inference image by training the model based on training data including a training image and annotation data, wherein the annotation data includes a mask set on the image of the object to be recognized in the training image, and the size of the mask in the short direction is larger than the size of the image of the object to be recognized in the short direction.
6. A learning method according to claim 5, wherein the recognition target image has a first side and a second side that are opposite to each other in the short direction of the recognition target image, the mask has a third side and a fourth side that are opposite to the first side and the second side, and the first distance between the first side and the third side and the second distance between the second side and the fourth side are equal to each other.
7. A learning device that performs the learning method described in any one of claims 1 to 6.
8. A control device comprising a control unit that recognizes the side surface of a plate-shaped member based on an image showing the side surface of the plate-shaped member, and causes a robot to hold the plate-shaped member based on the recognition result of the side surface.
9. A control device according to claim 8, wherein the robot is capable of grasping an object, the plate-shaped member has a first surface and a second surface facing each other in the thickness direction, and the control unit causes the robot to grasp the first surface and the second surface of the plate-shaped member.
10. A robot system comprising a control device according to claim 8 or claim 9, and a robot controlled by the control device.
11. A program for causing a computer device to execute the learning method described in any one of claims 1 to 6.
12. A program for causing a computer device to function as the control device described in claim 8 or claim 9.