Information processing method, information processing device, control method, robot system, article manufacturing method, program, and recording medium
The method uses a trained model to distinguish higher-ranking linear members in binary images, addressing the challenge of entanglement during robotic grasping, thereby improving assembly efficiency.
Patent Information
- Application Number
- JP2021171312
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-20
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2041-10-20
AI Technical Summary
Existing systems struggle to accurately determine the intersecting state of multiple linear members in images, leading to potential entanglement during robotic grasping, which complicates automated assembly processes.
A method involving a trained model that processes binary images to identify higher-ranking linear members based on discontinuous portions, using machine learning algorithms like Pix2Pix to distinguish between intersecting linear components, enabling precise robotic handling.
Prevents entanglement of linear members by accurately identifying and grasping the higher-ranking member, enhancing the efficiency and productivity of automated assembly tasks.
Smart Images

Figure 0007790914000001 
Figure 0007790914000002 
Figure 0007790914000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to robotics and image processing technology. [Background technology]
[0002] On production lines, tasks such as removing and assembling wire-like components such as harnesses are difficult processes, and traditionally, these tasks have often been performed manually. Therefore, efforts are being made to automate these tasks using machines that utilize cameras and recognition algorithms.
[0003] As a system for automating such work, Patent Document 1 discloses a system that measures the position and orientation of the tip of a linear member based on an image captured by a camera, and uses the measurement results to have a robot grasp the tip of the linear member.
[0004] Furthermore, Patent Document 2 also discloses that when multiple linear components are mixed and overlapping, the shape of the robot's hand is taken into consideration to select the linear components to be grasped by the hand so that the part and the hand do not interfere with each other and so that the hand does not grasp two linear components at the same time. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Publication No. 2020-112470 [Patent Document 2] International Publication No. 2019 / 098074 Summary of the Invention [Problem to be solved by the invention]
[0006] However, when multiple linear members intersect with each other, it can be difficult to determine the intersecting state in an image of these multiple linear members, and the linear members may be held by the robot in a state where they are entangled with other linear members.
[0007] An object of the present invention is to prevent a linear member from being held by a robot in a state where it is entangled with other linear members. [Means for solving the problem]
[0008] Book Disclosure In the first aspect, the processing unit Using the trained model, a binary image with gray levels capable of detecting multiple linear components is obtained, This information processing method is characterized in that a higher-ranking linear member is identified from among at least two mutually intersecting linear members based on a portion where the linear members are discontinuous in the image.
[0009] Book Disclosure In a second aspect, the processing unit Using the trained model, a binary image with gray levels capable of detecting multiple linear components is obtained, The information processing device is characterized in that it identifies a higher-ranking linear member from among at least two mutually intersecting linear members based on a portion where the linear members are discontinuous in the image.
[0010] Book Disclosure A third aspect of the present invention is a method for controlling a robot, comprising: Using the trained model, a binary image with gray levels capable of detecting multiple linear components is obtained, This control method is characterized by identifying a higher-order linear member from at least two intersecting linear members based on discontinuous portions of the linear members in the image, and controlling the robot to hold the higher-order linear member.
[0011] Book Disclosure A fourth aspect of the present invention is a robot system including a robot and a processing unit, wherein the processing unit: Using the trained model, a binary image with gray levels capable of detecting multiple linear components is obtained, This robot system is characterized by identifying a higher-order linear member from among at least two intersecting linear members based on discontinuous portions of the linear members in the image, and controlling the robot to hold the higher-order linear member. [Effects of the Invention]
[0012] According to the present invention, it is possible to prevent a linear member from being held by a robot in a state where it is entangled with other linear members. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is an explanatory diagram illustrating a schematic configuration of a robot system according to an embodiment. [Figure 2] FIG. 1 is an explanatory diagram of an image processing apparatus according to an embodiment. [Figure 3] FIG. 2 is a block diagram of a computer system in the robot system according to the embodiment. [Figure 4] FIG. 10 is an explanatory diagram of an assembly operation according to the embodiment. [Figure 5] FIG. 10 is an explanatory diagram of an assembly operation in a comparative example. [Figure 6] 1A is a flowchart showing a process in a preparation stage according to an embodiment, and FIG. 1B is a flowchart showing a process in an assembly work stage according to an embodiment. [Figure 7] FIG. 10 is an explanatory diagram of a preparatory stage according to the embodiment. [Figure 8] 1A is an explanatory diagram showing an example of an input image according to an embodiment, FIG. 1B is an explanatory diagram showing an example of an output image according to an embodiment, and FIG. 1C is an explanatory diagram showing an image of a reference example. [Figure 9] 10A to 10C are explanatory diagrams of an assembly work stage according to the embodiment. [Figure 10] 5(a) to 5(c) are diagrams illustrating a method for extracting intersection candidates according to the embodiment. [Figure 11] 10(a) and 10(b) are diagrams illustrating a method for extracting intersection candidates according to an embodiment. [Figure 12] 10A and 10B are diagrams illustrating the intersection detection process according to the embodiment. [Figure 13] 1(a) to 1(d) are explanatory diagrams of an example of three-dimensional measurement according to an embodiment. [Figure 14] 5(a) to 5(f) are explanatory diagrams of an example of a control method according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0014] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. FIG. 1 is an explanatory diagram showing a schematic configuration of a robot system 10 according to an embodiment. The robot system 10 includes a robot 100, an image processing device 200, a robot controller 300 as an example of a control device, and an imaging device 400 as an example of an imaging unit. The robot 100 is an industrial robot that is arranged on a production line and used to manufacture an assembly W20 as an example of an article. The assembly W20 is, for example, a device itself or a final or intermediate product of a component of the device. Examples of the device include electronic devices, electrical devices, and optical devices.
[0015] The robot 100 is a manipulator. The robot 100 is fixed to a stand (not shown). An intermediate product W10 is arranged around the robot 100. The intermediate product W10 includes two linear members W1 and W2, which are examples of a plurality of linear members, and a main body W0 to which the linear members W1 and W2 are connected. The linear member W1 is an example of a first linear member, and the linear member W2 is an example of a second linear member. Each of the linear members W1 and W2 is a flexible linear member, and examples of such members include electric wires such as signal lines and power lines, optical fibers, cables, and wire harnesses. The tip of each of the linear members W1 and W2 is a free end, and the base end of each of the linear members W1 and W2 is a fixed end fixed to the main body W0. Therefore, the robot 100 can determine the position and posture of the tip of each of the linear members W1 and W2 as desired by holding a predetermined portion of each of the linear members W1 and W2 during operation.
[0016] The tip of linear member W1 is formed, for example, by connector A1, and the tip of linear member W2 is formed, for example, by connector A2. Main body W0 has connector A3, which is an assembly object to which connector A1 is to be assembled, and connector A4, which is an assembly object to which connector A2 is to be assembled. Robot 100 assembles connector A1 to connector A3, and connector A2 to connector A4, thereby producing assembly W20.
[0017] While the description will be given taking an example in which the multiple linear members are two linear members W1 and W2, this is not limiting and the multiple linear members may be three or more linear members. Furthermore, the description will be given assuming that each linear member W1 and W2 has a connector A1 or A2, but they do not necessarily have to have a connector. Furthermore, the multiple linear members may be the same or different in color, material, and wire diameter.
[0018] The robot 100 and the robot controller 300 are communicatively connected by wire. The robot controller 300 and the image processing device 200 are communicatively connected by wire. The imaging device 400 and the image processing device 200 are communicatively connected by wire or wirelessly.
[0019] The robot 100 includes a robot arm 101 and a robot hand 102, which is an example of an end effector, i.e., an example of a holding mechanism. The robot arm 101 is, for example, a vertically articulated robot arm. The robot hand 102 is supported by the robot arm 101. The robot hand 102 is attached to a predetermined portion of the robot arm 101, for example, the tip of the robot arm 101. The robot hand 102 is configured to be able to hold each of the linear members W1 and W2. Note that, while a case where the holding mechanism is the robot hand 102 will be described, this is not limiting. For example, the holding mechanism may be a suction mechanism that can hold each of the linear members W1 and W2 by suctioning each of the linear members W1 and W2. In this embodiment, the robot hand 102 is configured to be able to grasp each of the linear members W1 and W2.
[0020] The imaging device 400 is configured as a stereo camera including a camera 401, which is an example of a first imaging unit, and a camera 402, which is an example of a second imaging unit. Each of the cameras 401 and 402 is a digital camera that captures an image of a subject, and outputs the captured image generated by capturing the image to the image processing device 200. Here, the image means digital image data in this embodiment.
[0021] The imaging device 400 is fixed to a frame (not shown). Each camera 401, 402 is disposed at a position where it can capture an image of an area including the plurality of linear members W1, W2. In other words, the intermediate product W10 is transported by a transport mechanism (not shown) to a predetermined position around the robot 100 so that the plurality of linear members W1, W2 are included in the angle of view of the imaging device 400. Therefore, the imaging device 400 can capture images of the plurality of linear members W1, W2.
[0022] Each of the cameras 401 and 402 captures an image of a plurality of linear members W1 and W2 as subjects, thereby generating two-dimensional images I11 and I12 as examples of first images. Each of the cameras 401 and 402 may be a color camera that generates a color image, or a monochrome camera that generates a monochrome image. While the imaging device 400 will be described with respect to a case where it includes two cameras, it may also include one camera, or three or more cameras.
[0023] The image processing device 200 sends an imaging command to each of the cameras 401 and 402 to cause each of the cameras 401 and 402 to capture an image. The image processing device 200 acquires each of the images I11 and I12 generated by each of the cameras 401 and 402 and performs image processing on each of the acquired images I11 and I12. The image processing device 200 calculates the position and orientation of a predetermined portion of each of the linear members W1 and W2 in the coordinate system of the robot 100 by performing image processing on each of the images I11 and I12. The robot controller 300 causes the robot 100 to selectively hold one of the predetermined portions of the linear members W1 and W2 based on the position and orientation of the predetermined portion of each of the linear members W1 and W2 calculated by the image processing device 200. The robot controller 300 controls the operation of the robot 100 so that the connectors of the linear members held by the robot 100 are assembled to the corresponding connectors. A plurality of linear members W1 and W2 are successively assembled to the main body W0 to produce an assembly W20.
[0024] In this embodiment, the image processing device 200 is configured by a computer. Fig. 2 is an explanatory diagram of the image processing device 200 according to the embodiment. The image processing device 200 has a main body 201, a display 202 which is an example of a display device connected to the main body 201, and a keyboard 203 and a mouse 204 which are examples of input devices connected to the main body 201.
[0025] 1 is configured as a computer in this embodiment. The robot controller 300 is configured to be able to control the movement of the robot 100, that is, the posture of the robot 100.
[0026] 3 is a block diagram of a computer system in the robot system 10 according to the embodiment. The main body 201 of the image processing device 200 includes a CPU (Central Processing Unit) 251, which is an example of a processor. The CPU 251 is an example of a processing unit. The main body 201 also includes, as storage units, a ROM (Read Only Memory) 252, a RAM (Random Access Memory) 253, and an HDD (Hard Disk Drive) 254. The main body 201 also includes a recording disk drive 255 and an interface 256, which is an input / output interface. The CPU 251, ROM 252, RAM 253, HDD 254, recording disk drive 255, and interface 256 are connected via a bus so as to be able to communicate with each other.
[0027] The ROM 252 stores a basic program related to the operation of the computer. The RAM 253 is a storage device that temporarily stores various data such as the results of arithmetic processing by the CPU 251. The HDD 254 stores the results of arithmetic processing by the CPU 251 and various data acquired from the outside, as well as a program 261 that causes the CPU 251 to execute various processes described below. The program 261 is application software that enables the CPU 251 to perform the various processes described below. Therefore, the CPU 251 can execute the image processing described below by executing the program 261 recorded in the HDD 254. The recording disk drive 255 can read out various data, programs, etc. recorded on the recording disk 262.
[0028] In this embodiment, the non-transitory computer-readable recording medium is the HDD 254, and the program 261 is recorded on the HDD 254, but this is not limiting. The program 261 may be recorded on any non-transitory computer-readable recording medium. Examples of recording media that can be used to provide the program 261 to a computer include a flexible disk, a hard disk, an optical disk, a magneto-optical disk, a magnetic tape, and a non-volatile memory.
[0029] The robot controller 300 includes a CPU 351, which is an example of a processor. The CPU 351 is an example of a control unit. The robot controller 300 also includes a ROM 352, a RAM 353, and a HDD 354 as storage units. The robot controller 300 also includes a recording disk drive 355 and an interface 356, which is an input / output interface. The CPU 351, ROM 352, RAM 353, HDD 354, recording disk drive 355, and interface 356 are connected via a bus so that they can communicate with each other.
[0030] The ROM 352 stores a basic program related to the operation of the computer. The RAM 353 is a storage device that temporarily stores various data, such as the results of calculations performed by the CPU 351. The HDD 354 records the results of calculations performed by the CPU 351 and various data acquired from the outside, and also records (stores) a program 361 that causes the CPU 351 to execute various processes. The program 361 is application software that enables the CPU 351 to perform various processes, which will be described later. Therefore, the CPU 351 executes the program 361 recorded in the HDD 354 to perform control processing and control the operation of the robot 100 in FIG. 1. The recording disk drive 355 can read out various data, programs, etc. recorded on the recording disk 362.
[0031] In this embodiment, the computer-readable non-transitory recording medium is the HDD 354, and the program 361 is recorded on the HDD 354, but this is not limiting. The program 361 may be recorded on any computer-readable non-transitory recording medium. Examples of recording media that can be used to provide the program 361 to a computer include a flexible disk, a hard disk, an optical disk, a magneto-optical disk, a magnetic tape, and a non-volatile memory.
[0032] In this embodiment, the image processing and control processing are executed by multiple computers, i.e., multiple CPUs 251 and 351, but this is not limiting. The image processing and control processing may be executed by one computer, i.e., one CPU. In this case, it is sufficient to configure the one CPU to function as both a processing unit and a control unit.
[0033] A case in which linear members W1 and W2 are assembled to a main body W0 in this order in such a robot system 10 will be described with reference to FIG. 4. Based on three-dimensional coordinate information indicating the holding position of linear member W1 calculated by the image processing device 200, the robot controller 300 controls the robot arm 101 to move the robot hand 102 above the holding position. Then, the robot controller 300 causes the robot hand 102 to hold the linear member W1. Next, the robot controller 300 controls the robot arm 101 to assemble the connector A1 of the linear member W1 to the connector A3 of the main body W0. After completing the assembly of the connector A1 to the connector A3, the robot controller 300 repeats the same processing operation for the linear member W2. This completes the assembly process, and an assembly W20 is produced.
[0034] Here, a flexible linear member is soft and easily changes its posture. Therefore, at least two of the multiple linear members, for example, two linear members W1 and W2 as shown in FIG. 5, may be arranged around the robot 100 in an intersecting manner. At the intersection of the linear member W1 and the linear member W2, the linear member W2 may be positioned higher than the linear member W1. Here, the higher linear member refers to the linear member that is positioned highest among the at least two linear members at the intersection of the at least two linear members. Furthermore, the lower linear member refers to the linear member that is positioned lower than the linear member that is positioned highest among the at least two linear members at the intersection of the at least two linear members. In the example of FIG. 5, at the intersection of the linear member W1 and the linear member W2 before assembly, the linear member W2 is the higher linear member and the linear member W1 is the lower linear member.
[0035] In this case, if the robot 100 holds the linear member W1 before the linear member W2 and assembles the connector A1 to the connector A3, the linear member W2 may become entangled with the linear member W1. In this state, it may be difficult for the robot 100 to hold the linear member W2 next, or even if the robot 100 can hold the linear member W2, it may be difficult to assemble the connector A2 to the connector A4. Furthermore, linear members are flexible and prone to change position, and the linear members themselves may be small (i.e., have a thin wire diameter). Therefore, a method for measuring the intersection state of multiple linear members with high accuracy in three dimensions, such as a spatial encoding method using a projector, may be considered. However, this measurement method requires a long measurement process, reducing the productivity of the assembly W20.
[0036] In this embodiment, the CPU 251 of the image processing device 200 determines the intersecting state of the linear members W1 and W2 using at least one of the two-dimensional images I11 and I12 captured by the imaging device 400 in order to avoid a situation in which the linear members become entangled during assembly work.
[0037] 6(a) and 6(b) are flowcharts of processing according to an embodiment. The flowcharts include each process of a control method for the robot 100, including an image processing method. The flowchart shown in FIG. 6(a) shows the advance preparation stage, and the flowchart shown in FIG. 6(b) shows the assembly work stage. Here, the CPU 251 executes the program 261 to perform the following processing. Furthermore, the CPU 351 executes the program 361 to perform the following processing.
[0038] First, the advance preparation stage will be described with reference to the flowchart in Fig. 6(a). Fig. 7 is an explanatory diagram of the advance preparation stage according to the embodiment. In the advance preparation stage, a trained model 260 is created as an example of a predetermined model.
[0039] The subject prepared in this preliminary preparation stage is an intermediate product having a similar configuration to the intermediate product W10 that will be the target in the assembly work stage. That is, in the assembly work stage, an assembled product W20 is manufactured from the intermediate product W10, but an intermediate product other than the intermediate product W10 used in the assembly work stage may be prepared in the preliminary preparation stage as long as it has a similar configuration to the intermediate product W10 used in the manufacture in the assembly work stage. Hereinafter, the intermediate product used in the preliminary preparation stage will be described using the same reference numerals as the intermediate product W10 used in the assembly work stage, since it has a similar configuration to the intermediate product W10 used in the assembly work stage.
[0040] Furthermore, it is preferable that the exposure conditions, such as the imaging time and lens aperture of the imaging device 400, in the advance preparation stage be set to the same as those in the assembly work stage. It is also preferable that the installation position of the imaging device 400 and the usage environment, such as the illumination light irradiated onto the multiple linear members W1 and W2, be set to the same as those in the assembly work stage in the advance preparation stage. Note that, although a case where the imaging device 400 used in the assembly work stage is used in the advance preparation stage will be described, this is not limiting, and an imaging device different from the imaging device 400 may also be used. Furthermore, since no assembly work is performed in this advance preparation stage, the robot 100 does not need to be present around the multiple linear members W1 and W2.
[0041] In step S11, CPU 251 causes imaging device 400 to capture images of multiple linear members W1 and W2, and acquires image I1A, which is the captured image, from imaging device 400. Image I1A acquired in the processing of step S11 becomes part of training data D1 used for learning. Therefore, the more images I1A acquired in step S11, the better.
[0042] In step S11, the CPU 251 causes the imaging device 400 to capture images of the linear members W1 and W2 in several patterns in which the crossing states of the linear members W1 and W2 are different, acquires a plurality of images I1A, and stores these images I1A in a storage unit such as the HDD 254. The posture of the linear members W1 and W2 may be changed by an operator or a mechanism such as a robot.
[0043] The multiple images I1A may include images in which the linear members W1 and W2 do not intersect, or may include images in which either the linear members W1 or W2 does not appear. FIG. 8(a) illustrates an example of one image I1A among the multiple images I1A. The image I1A shown in FIG. 8(a) is a grayscale image. The image I1A includes images IW1A and IW2A corresponding to the linear members W1 and W2, and an image IB1A corresponding to the background.
[0044] Next, in step S12, under the operator's command, the CPU 251 creates training data D1 based on the image I1A obtained in step S11. In this embodiment, a machine learning algorithm is used in the learning process, and among machine learning algorithms, a "supervised learning" algorithm is particularly used. Therefore, in step S12, training data D1 to be used in "supervised learning" is created. The training data D1 includes multiple images I1A and multiple images I2A that have a one-to-one correspondence with the multiple images I1A. A trained model 260 is created by machine learning so that each image I2A is output from each image I1A.
[0045] Each image I1A is an example of an input image. Each image I2A is an example of an output image. Each image I2A is created based on the corresponding image I1A by an operator operating the image processing device 200 in accordance with the following steps 1 to 3.
[0046] Step 1: An image corresponding to the background other than the linear members W1 and W2 (other than the linear members) is painted with a single color to create a background image. Step 2: The images corresponding to the linear members W1 and W2 are made into linear images filled with a single color, which is different from the color of the background image. Step 3: At the intersection of the linear image corresponding to linear member W1 and the linear image corresponding to linear member W2, the linear image corresponding to the lower linear member is discontinuous. Also, the linear image corresponding to the higher linear member is continuous. The colors of the linear images need only be different from the color of the background image, and may be the same color or different colors.
[0047] FIG. 8(b) shows an example of image I2A. Image I2A shown in FIG. 8(b) is created based on image I1A shown in FIG. 8(a). In image I2A, a linear image L1A corresponding to an upper linear element and a linear image L2A corresponding to a lower linear element are distinguished by whether they are continuous or interrupted at an intersection C1A. Here, the linear image L2A being discontinuous, or interrupted, at the intersection C1A refers to a state in which the linear image L2A is divided into two, sandwiched between the linear image L1A, so that the linear image L2A does not contact the linear image L1A. Thus, at the intersection C1A, the linear image L2A corresponding to the lower linear element is intermittent. In image I2A shown in FIG. 8(b), the linear images L1A and L2A are a first color, e.g., white, and the background image B1A is a second color different from the first color, e.g., black.
[0048] Note that when a computer performs image processing such as binarization on image I1A, the processed image obtained by the image processing does not necessarily become image I2A as shown in FIG. 8(b). For example, in the processed image, there may be no boundary between two intersecting linear images, and the two linear images may be continuous. In this embodiment, in step 3, the operator may create image I2A directly from image I1A. Alternatively, in step 3, the computer may first perform image processing such as binarization on image I1A, and then the operator may modify the processed image to create image I2A.
[0049] By going through the above steps 1 to 3, a pair of one image I1A and one image I2A is created. By performing the above steps 1 to 3 for all images I1A, multiple pairs of images I1A and I2A are created. These multiple pairs make up the training data D1.
[0050] Next, in step S13, CPU 251 performs machine learning to associate image I1A with image I2A. The machine learning algorithm used in step S13 is, for example, Pix2Pix (Pixel to Pixel). Pix2Pix is a type of "supervised learning" and is an algorithm that performs machine learning based on training data D1 to infer an output value for each pixel of an input image. If the learning progresses smoothly, the linear image corresponding to the upper linear member at the intersection will be filled in with a continuous solid color, and the linear image corresponding to the lower linear member at the intersection will be filled in with a discontinuous color at the intersection.
[0051] Pix2Pix is used in this type of learning because it can automatically extract features for determining whether linear components are continuous or discontinuous from the shading information on the image of the intersecting linear components. Features include, for example, the shading formed on the linear image corresponding to the lower linear component at the intersection and the color continuity of the linear image corresponding to the upper linear component at the intersection. Therefore, in this embodiment, if the vertical relationship between two intersecting linear components does not appear as a shading characteristic in a captured image, as shown in FIG. 8(c), the captured image is excluded from the training data D1. For example, if light is irradiated uniformly from all directions, such as with dome lighting, and an image is captured with an equal amount of light, no shadows are formed near the intersection, which can lead to the situation shown in FIG. 8(c). Therefore, images obtained by such capture are excluded.
[0052] Furthermore, the algorithm used for machine learning is not limited to Pix2Pix, and any algorithm other than Pix2Pix may be used as long as it has the function of extracting the above-mentioned features.
[0053] The trained model 260 obtained by the training in step S13 is stored in a storage unit of the image processing device 200, for example, in the HDD 254.
[0054] Next, the assembly work stage will be described with reference to the flowchart of Fig. 6(b). Fig. 9 is an explanatory diagram of the assembly work stage according to the embodiment. Note that the intermediate product W10 is sequentially transported near the robot 100. In step S21, the CPU 251 causes the imaging device 400 to capture images of the linear members W1 and W2 of the intermediate product W10 transported near the robot 100. As described above, it is preferable that the imaging conditions, such as exposure time and lens aperture, and the usage environment during this imaging are the same as those in step S11.
[0055] Through the imaging operation of step S21, the CPU 251 acquires an image I1 as shown in FIG. 9 from the imaging device 400. The image I1 is an example of a first image. The image I1 is a grayscale image. In this embodiment, the imaging device 400 is a stereo camera, and two images I11 and I12 are obtained through the imaging operation. The image I1 shown in FIG. 9 is one of the two images I11 and I12. As shown in FIG. 9, two linear members W1 and W2 are captured in the image I1. That is, the image I1 includes images IW1 and IW2 corresponding to the linear members W1 and W2 and an image IB1 corresponding to the background. In this embodiment, the linear members W1 and W2 are captured so that their tips, which are an example of predetermined portions to be held by the robot 100, are captured. In the image I1, the end point IE1 of the image IW1 corresponds to the end of the linear member W1, and the end point IE2 of the image IW2 corresponds to the end of the linear member W2.
[0056] Next, in step S22, the CPU 251 infers the intersection state of the linear members W1 and W2 from the image I1 captured in step S21. For the inference, the trained model 260 stored in the HDD 254 is used. That is, in step S22, the CPU 251 performs image processing on the image I1 based on the trained model 260 to generate an image I2. In other words, the CPU 251 uses the image I1 as an input image and processes the image I1 based on the trained model 260 to generate an image I2 as an output image. The image I2 is an example of a second image.
[0057] Image I2 includes linear images L1 and L2 corresponding to linear members W1 and W2. In FIG. 9, linear images L1 and L2 are a first color, such as white, and background image B1 is a second color different from the first color, such as black. At the intersection of two linear images L1 and L2, the linear image corresponding to the higher linear member is a continuous monochromatic image, while the linear image corresponding to the lower linear member is a discontinuous monochromatic image. That is, at the intersection, the linear image corresponding to the lower linear member is discontinuous, sandwiched between the linear image corresponding to the higher linear member.
[0058] 9, image I2 includes multiple intersections C1 and C2. At intersection C1, the first linear image corresponding to the higher linear member is linear image L1, and the second linear image corresponding to the lower linear member is linear image L2. At intersection C2, the first linear image corresponding to the higher linear member is linear image L2, and the second linear image corresponding to the lower linear member is linear image L1.
[0059] A method for detecting intersection parts C1 and C2 in image I2 will be described below. In steps S23 to S24, CPU 251 extracts intersection candidates that are candidates for intersection parts from image I2, performs a predetermined calculation on the intersection candidates, and performs intersection determination processing to determine whether the intersection candidates are intersection parts based on the results of the predetermined calculation.
[0060] The process of step S23 will be described. First, the CPU 251 extracts intersection candidates in the image 12. Hereinafter, the method of extracting intersection candidates will be described in detail with reference to Figs. 10(a) to 11(b).
[0061] First, as shown in FIG. 10(a), the CPU 251 extracts line segment images L11, L12, L21, and L22 from the image I2. Next, as shown in FIG. 10(b), the CPU 251 extracts the endpoints of each of the line segment images L11, L12, L21, and L22. Various endpoint extraction algorithms, such as the Harrison corner detection method, are known as methods for extracting endpoints, but the method is not particularly limited. The endpoints extracted in this manner are shown in FIG. 10(b) as endpoints L11-a, L11-b, L12-a, L21-a, L21-b, and L22-a. Next, the CPU 251 defines a predetermined region including each endpoint and sets it as an endpoint region for each endpoint. In Figure 10(c), these are shown as endpoint regions T11-a, T11-b, T12-a, T21-a, T21-b, and T22-a. Each endpoint region is, for example, a rectangular region of a predetermined size, with each endpoint located at the center of the endpoint region. Note that the shape of each endpoint region is not limited to a rectangle, and may be various shapes, such as a circle.
[0062] If there are at least two endpoint regions in image I2, CPU 251 determines whether the positional relationship between the positions and orientations of the line segment images included in the endpoint regions in the combination of all endpoint regions is within a certain threshold. Based on the result of this determination, if it is within the certain threshold, the region including the endpoint region of the combination is extracted as an intersection candidate.
[0063] A specific description will be given below using two line segment images L21 and L22 as examples. Line segment image L21 is an example of a first line segment image, and line segment image L22 is an example of a second line segment image. Furthermore, endpoint L21-b of line segment image L21 is an example of a first endpoint, and endpoint L22-a of line segment image L22 is an example of a second endpoint. An endpoint region T21-b including endpoint L21-b is an example of a first region, and endpoint region T22-a including endpoint L22-a is an example of a second region.
[0064] 11(a), the CPU 251 calculates a regression line LR1 that follows the line segment image L21 in the endpoint region T21-b, and a regression line LR2 that follows the line segment image L22 in the endpoint region T22-a. The regression line LR1 is calculated using the portion of the line segment image L21 that is included in the endpoint region T21-b. The regression line LR2 is calculated using the portion of the line segment image L22 that is included in the endpoint region T22-a.
[0065] The CPU 251 calculates the distance Da between a perpendicular line Lp1 drawn from the regression line LR1 to the end point L21-b of the line segment image L21 and a perpendicular line Lp2 drawn from the regression line LR2 to the end point L22-a of the line segment image L22. The CPU 251 also calculates the intersection angle θb between the regression lines LR1 and LR2.
[0066] In this embodiment, the CPU 251 determines whether predetermined conditions are satisfied, that is, whether the distance Da is equal to or smaller than a threshold value THa and the intersection angle θb is equal to or smaller than a threshold value THb. If the predetermined conditions are satisfied between the line segment image L21 and the line segment image L22, the CPU 251 extracts a region including the end point region T21-b and the end point region T22-a as the intersection candidate C11. The intersection candidate C11 is, for example, a rectangular minimum region including the end point region T21-b and the end point region T22-a. Note that the shape of the region indicating the intersection candidate is not limited to a rectangular shape and may be various shapes, for example, a circular shape.
[0067] In Figure 11(b), the area including the combination of endpoint regions T21-b and T22-a satisfies the predetermined condition and is therefore an intersection candidate C11, and the area including the combination of endpoint regions T11-b and T12-a satisfies the predetermined condition and is therefore an intersection candidate C12. This makes it possible to infer the discontinuous area between the two line segment images. Note that in Figure 11(c), for example, endpoint region T11-a and endpoint region T21-a do not satisfy the predetermined condition and are therefore not intersection candidates.
[0068] Next, for the intersection candidate C11, the CPU 251 interpolates the line segment image L21 and the line segment image L22 by a predetermined interpolation. The predetermined interpolation is preferably linear interpolation. By linearly interpolating the line segment image L21 and the line segment image L22, a linear interpolated area R1 shown in FIG. 11(a) is obtained. An interpolated area is similarly obtained for the intersection candidate C12 shown in FIG. 11(b).
[0069] By the above interpolation process, an interpolation area R1 for the intersection candidate C11 and an interpolation area R2 for the intersection candidate C12 are obtained, as shown in Fig. 12(a). The interpolation area R1 is a pixel area connecting the line segment image L21 and the line segment image L22. The interpolation area R2 is a pixel area connecting the line segment image L11 and the line segment image L12. Note that the method is not limited to linear interpolation as long as it is a method that can infer the discontinuous area between two line segment images.
[0070] Next, in step S24, the CPU 251 determines whether the line segment image L11 and the interpolated region R1 overlap in the intersection candidate C11. If the line segment image L11 and the interpolated region R1 overlap, the CPU 251 determines the intersection candidate C11 as the intersection portion C1. Similarly, the CPU 251 determines whether the line segment image L22 and the interpolated region R2 overlap in the intersection candidate C12. If the line segment image L22 and the interpolated region R2 overlap, the CPU 251 determines the intersection candidate C12 as the intersection portion C2. Note that depending on the state of the linear members W1 and W2, there may be one intersection portion, or there may be multiple intersection portions C1 and C2 as shown in FIG. 9. Furthermore, if one line segment image and the interpolated region do not overlap in the intersection candidate, the intersection candidate is determined not to be an intersection portion.
[0071] 12(b), at the intersection C1, the continuous line segment image L11 is the linear image L1 corresponding to the linear member W1, and the line segment images L21 and L22 connected in the interpolation area R1 are the linear image L2 corresponding to the linear member W2. At the intersection C1, the CPU 251 identifies the linear member W1 corresponding to the linear image L1 as the upper linear member WU, and identifies the linear member W2 corresponding to the linear image L2 as the lower linear member WL.
[0072] Next, in step S25, the CPU 251 determines the position of a predetermined part of the upper linear member WU that the robot 100 is to hold, i.e., the holding position. In this embodiment, the CPU 251 calculates the holding position by three-dimensional measurement using images captured by the imaging device 400. As shown in FIG. 1, the imaging device 400 has multiple cameras 401 and 402 arranged at intervals. The parallax between the camera 401 and the camera 402 enables three-dimensional measurement of the linear members W1 and W2 that are the subjects.
[0073] 13(a) to 13(d) are explanatory diagrams of an example of three-dimensional measurement according to an embodiment. In step S21 described above, the CPU 251 acquires an image I11 shown in FIG. 13(a) generated by the imaging operation of the camera 401, and an image I12 shown in FIG. 13(a) generated by the imaging operation of the camera 402. Each of the images I11 and I12 is a grayscale image. The image I11 includes an image IW11 corresponding to the linear member W1 and an image IW21 corresponding to the linear member W2. The end point IE11 of the image IW11 corresponds to the tip of the linear member W1, and the end point IE21 of the image IW21 corresponds to the tip of the linear member W2. The image I12 includes an image IW12 corresponding to the linear member W1 and an image IW22 corresponding to the linear member W2. An end point IE12 of the image IW12 corresponds to the tip of the linear member W1, and an end point IE22 of the image IW22 corresponds to the tip of the linear member W2.
[0074] Then, in the above-mentioned step S22, the CPU 251 generates the images I21 and I22 shown in FIG. 13(b) from the images I11 and I12 by inference based on the trained model 260. The images I21 and I22 are binarized images. The image I21 includes a linear image L11 corresponding to the linear member W1, i.e., the image IW11, and a linear image L21 corresponding to the linear member W2, i.e., the image IW21. The image I22 includes a linear image L12 corresponding to the linear member W1, i.e., the image IW12, and a linear image L22 corresponding to the linear member W2, i.e., the image IW22. Then, in the above-mentioned steps S23 and S24, the CPU 251 performs intersection determination processing. As a result, an intersection C11 between linear images L11 and L21 is determined in image I21, and an intersection C12 between linear images L12 and L22 is determined in image I22. At the intersection C11, it is determined that linear image L11 corresponds to the higher linear member WU, and it is determined that linear image L21 corresponds to the lower linear member WL. Furthermore, at the intersection C12, it is determined that linear image L12 corresponds to the higher linear member WU, and it is determined that linear image L22 corresponds to the lower linear member WL.
[0075] Next, in each image I21, I22, the end points E11, E12 of each linear image L11, L12 corresponding to the upper linear member WU are determined. Here, it is assumed that the position of the main body W0 around the robot 100 and the extension direction of the linear members W1, W2 relative to the main body W0 are predetermined. This determines the approximate area in each image I21, I22 where the end points E11, E12 corresponding to the connectors A1, A2 at the ends of the linear members W1, W2 are located. As shown in FIG. 13(c), the CPU 251 interlaces scans each image I21, I22 from the opposite position to the position where the main body W0 is located to search for the linear images L11, L12 corresponding to the upper linear member WU. Then, the CPU 251 determines the first encountered points in each image I21, I22 as the end points E11, E12.
[0076] Note that even if the linear member W1 is a linear member without a connector, the end points E11, E12 may be searched for in a similar manner. Furthermore, the holding position of the portion of the linear member that the robot 100 is to hold may be a position offset from the tip by a predetermined length from the tip toward the base end of the linear member. In this case, as shown in FIG. 13(d), after detecting the end points E11, E12, the CPU 251 determines points F11, F12 that are offset by a predetermined length from the end points E11, E12 along the linear images L11, L12.
[0077] Based on the positions of the endpoints E11, E12 or the points F11, F12 on the images I21, I22 obtained by the above process, the three-dimensional position of the holding position can be measured by the principle of triangulation. The CPU 251 transmits information about the holding position to the CPU 351 of the robot controller 300. In step S26, the CPU 351 of the robot controller 300 controls the robot 100 to hold a predetermined portion of the linear member W1, which is the upper linear member WU, based on the received information about the holding position. The CPU 351 then controls the robot 100 to assemble the connector A1 of the linear member W1 to the connector A3. This prevents the linear members W1, W2 from being held by the robot 100 in an entangled state, allowing the assembly work to be performed smoothly.
[0078] Although the image capturing device 400 is a stereo camera in the above description, the present invention is not limited to this. For example, the image capturing device 400 may be a monocular camera. In this case, the measurement of the holding position may be performed using a separate device.
[0079] In the above example, a case where there is one intersection between the linear members W1 and W2 has been described, but there may be multiple intersections between the linear members W1 and W2. In such a case, the robot 100 may be operated to eliminate the intersection between the linear members W1 and W2. The processing for such a case will be described with reference to FIGS. 14(a) to 14(f). Note that FIGS. 14(a) to 14(f) illustrate only an image I1 obtained from one of the multiple cameras 401 and 402 shown in FIG. 1, or an image I2 obtained by processing the image I1. The image obtained from the other camera is also processed in the same way as the image I1 obtained from the one camera, and therefore its description will be omitted.
[0080] First, the CPU 251 acquires the image I1 shown in Fig. 14(a) from the imaging device 400. In the image I1, two images IW1 and IW2 corresponding to two linear members W1 and W2 are twisted so as to intersect twice. In the image I1, the end point IE1 of the image IW1 corresponds to the tip of the linear member W1, and the end point IE2 of the image IW2 corresponds to the tip of the linear member W2.
[0081] 14(b) from the image I1 based on the trained model 260, and performs intersection detection processing, thereby detecting multiple, for example, two intersections C1 and C2 between the linear image L1 corresponding to the linear member W1 and the linear image L2 corresponding to the linear member W2.
[0082] Next, the CPU 251 selects only one intersection CC of the intersections C1 and C2 as a focus, as shown in FIG. 14(c). The twisted linear members W1 and W2 need to be unwound from their tips. Therefore, the CPU 251 searches for an intersection by interlaced scanning the image I2 from a position opposite to the position where the main body W0 is located, and selects the first intersection C1 as the focus intersection CC. Alternatively, the CPU 251 may trace the linear image L1 from the endpoint E1 of the linear image L1 that is first encountered by interlaced scanning and select the first intersection C1 as the focus intersection CC.
[0083] Once the unique intersection CC of interest is determined in this manner, the upper linear member WU and the lower linear member WL are defined at the intersection CC by the above-described process, as shown in FIG. 14(d). In this manner, the CPU 251 identifies the upper linear member WU at the first intersection point by tracing the tips of the linear members W1 and W2. Note that, at the intersection C1, the linear image L1 corresponds to the upper linear member WU, and the linear image L2 corresponds to the lower linear member WL. Then, the endpoint E1 is searched for by interlaced scanning, and the grip position is determined.
[0084] Because there are two intersections C1 and C2, the CPU 351 controls the robot 100 to hold the linear member W1, which is the upper linear member WU, at the intersection C1. The CPU 351 then controls the robot 100 to eliminate the intersection of the upper linear member WU with the lower linear member WL. The position or location to which the upper linear member WU is retracted is determined based on the positional relationship between the linear members W1 and W2 on the two-dimensional image related to the intersection C1. By repeating the above operations, the number of intersections is reduced by one.
[0085] The CPU 251 again causes the imaging device 400 to capture images of the linear members W1 and W2, and performs image processing on the captured image based on the trained model 260 to generate a new image I2 as shown in FIG. 14(e). The CPU 251 then performs the intersection determination process again. In the example shown in FIG. 14(e), only one intersection C2 is detected. In the intersection C2, the linear image L2 corresponds to the upper linear member WU, and the linear image L1 corresponds to the lower linear member WL. Then, as shown in FIG. 14(f), interlaced scanning is performed to search for the end point E2 of the linear image L2, and the gripping position is determined. Thereafter, the linear members W2 and W1 are assembled into the main body W0 in the order of W2 and W1. The above operations eliminate the entanglement of the linear members W1 and W2, facilitating the assembly of the linear members W1 and W2.
[0086] As described above, according to this embodiment, the overlapping state of the linear members W1 and W2 can be determined with high accuracy using the image I1 captured of the linear members W1 and W2, so that the upper linear member of the linear members W1 and W2 can be held by the robot 100. That is, the robot 100 is prevented from holding a linear member in a state where it is entangled with other linear members. For example, when assembling the connector A2 to the connector A4 and the connector A1 to the connector A3, it is possible to avoid a situation in which the linear members W1 and W2 are entangled as shown in FIG. 5. This improves the productivity of the assembly W20.
[0087] Furthermore, since the image I2 is generated using the trained model 260 to determine the intersection state of the linear members W1 and W2, there is no need to use a high-pixel camera for the imaging device 400. Furthermore, there is no need to prepare a separate projector to determine the holding position of the linear member, nor is there a need to prepare a highly accurate 3D measurement system that measures the thickness of a thin wire. Furthermore, in this embodiment, since the intersection state of multiple linear members W1 and W2 can be determined using the image I2 generated based on the trained model 260, it is possible to have the robot 100 hold the linear members without tangling them with a simple and portable configuration. Therefore, the productivity of the assembly W20 is improved.
[0088] It should be noted that the present invention is not limited to the above-described embodiments, and many modifications are possible within the technical concept of the present invention. Furthermore, the effects described in the embodiments are merely a list of the most preferable effects resulting from the present invention, and the effects of the present invention are not limited to those described in the embodiments.
[0089] In the above embodiment, the image I2A used in the training data D1 is a binary image, but this is not limiting. In the image I2A, each of the linear images corresponding to the linear components may be in a different color. In this case, in the image I2 generated based on the trained model 260, the linear images will be in a different color.
[0090] Furthermore, in the above-described embodiment, the robot arm 101 is described as a vertically articulated robot arm, but the present invention is not limited to this. The robot arm 101 may be, for example, a horizontally articulated robot arm, a parallel link robot arm, an orthogonal robot, or any of various other robot arms. Furthermore, instead of a robot arm, the present invention may be applied to a machine that can automatically perform operations such as extension and contraction, bending and stretching, vertical movement, horizontal movement, or rotation, or a combination of these operations, based on information stored in a storage device provided in a control device.
[0091] In addition, in the above embodiment, the imaging device 400 is fixed to a frame other than the robot 100, but this is not limiting, and the imaging device 400 may be fixed to the robot 100, for example.
[0092] (Other Examples) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions. [Explanation of symbols]
[0093] 10... robot system, 100... robot, 200... image processing device, 251... CPU (processing unit), 351... CPU (control unit), 400... imaging device (imaging unit)
Claims
1. The processing unit Using the trained model, a binary image with gray levels capable of detecting multiple linear components is obtained, identifying a higher-ranking linear member from among at least two mutually intersecting linear members based on a portion where the linear members are discontinuous in the image; 1. An information processing method comprising:
2. A portion where at least two linear members intersect in the image is discontinuous.
2. The information processing method according to claim 1,
3. the images include a first linear image corresponding to an upper linear member and a second linear image corresponding to a lower linear member intersecting the upper linear member; The processing unit determining that the linear member corresponding to the second linear image is discontinuous when the second linear image is divided into two parts with the first linear image sandwiched between them so as not to come into contact with the first linear image; Identifying a linear member corresponding to the second linear image that is discontinuous as a lower linear member, and identifying a linear member corresponding to the first linear image that is not discontinuous as a higher linear member.
3. The information processing method according to claim 1, wherein:
4. The processing unit extracting an area including a first area in which a first end point of a first line segment image exists and a second area in which a second end point of a second line segment image exists as an intersection candidate if a predetermined condition is satisfied between the first area and the second area in the image; If an area obtained by interpolating the first line segment image and the second line segment image using a predetermined interpolation method overlaps with another line segment image, the intersection candidate is determined to be an intersection portion.
4. The information processing method according to claim 1, wherein the first and second inputs are input to the first and second inputs.
5. The predetermined condition is a condition that a distance between a perpendicular line drawn from a first regression line along the first line segment image to the first endpoint and a perpendicular line drawn from a second regression line along the second line segment image to the second endpoint is within a threshold range, and an intersection angle between the first regression line and the second regression line is within a threshold range.
5. The information processing method according to claim 4.
6. The predetermined interpolation is a linear interpolation.
6. The information processing method according to claim 4 or 5.
7. Each of the linear members has a free end, taking an image of the linear members so that the tips of the linear members can be detected from the image; 7. The information processing method according to claim 1, wherein:
8. the at least two linear members include a first linear member and a second linear member that intersect with each other; When there are a plurality of points where the first linear member and the second linear member intersect, the processing unit traces back from the tip of each of the first linear member and the second linear member to identify a linear member that is higher than the first intersecting point.
8. The information processing method according to claim 1, wherein:
9. The processing unit having the robot retain the first detected end point of the identified upper linear member relative to the first intersection; 9. The information processing method according to claim 8.
10. The image is binarized by making the color of the linear members different from the color of the background of the linear members.
10. The information processing method according to claim 1, wherein:
11. The processing unit: acquiring a first image of the linear member; converting the first image into the image for detecting linear members using the trained model; 11. The information processing method according to claim 1.
12. The trained model is generated by machine learning using the first image as an input image and a second image in which the shades of light and shade that can detect multiple linear members are binarized as training data as an output image; The processing unit converting the first image into the image for detecting linear members using the trained model; 12. The information processing method according to claim 11.
13. The second image used as the training data can be modified by a user.
13. The information processing method according to claim 12.
14. The processing unit controlling the robot to hold the upper linear member, and then controlling the robot to assemble the tip of the upper linear member to the assembly target; 14. The information processing method according to claim 1,
15. The processing unit controlling the robot to hold the upper linear member, and then controlling the robot to prevent the upper linear member from crossing the lower linear member; 14. The information processing method according to claim 1,
16. The first image is acquired by a stereo camera.
14. The information processing method according to claim 11, wherein:
17. The processing unit Using the trained model, a binary image with gray levels capable of detecting multiple linear components is obtained, identifying a higher-ranking linear member from among at least two mutually intersecting linear members based on a portion where the linear members are discontinuous in the image; 1. An information processing device comprising:
18. A method for controlling a robot, comprising: Using the trained model, a binary image with gray levels capable of detecting multiple linear components is obtained, Identifying a higher-ranking linear member from among at least two mutually intersecting linear members based on a portion where the linear members are discontinuous in the image; controlling the robot to hold the upper linear member; A control method comprising:
19. A robot system including a robot and a processing unit, The processing unit Using the trained model, a binary image with gray levels capable of detecting multiple linear components is obtained, Identifying a higher-ranking linear member from among at least two mutually intersecting linear members based on a portion where the linear members are discontinuous in the image; controlling the robot to hold the upper linear member; A robot system characterized by:
20. A method for manufacturing an article by manipulating a linear member using the information processing method according to any one of claims 1 to 16.
21. A program for causing a computer to execute the information processing method according to any one of claims 1 to 16 or the control method according to claim 18.
22. A computer-readable recording medium on which the program according to claim 21 is recorded.
Citation Information
Patent Citations
Image processor, image processing method, and computer- readable storage medium with image processing program stored therein, and diagnosis support system using them
JP2002269539A
Image processing method and image processor
JP2007207009A
Workpiece holding method
JP2015196208A
Manufacturing device for electronic apparatus and cable shape estimation program
JP2018014362A
Information processing apparatus, information processing method, and program
JP2019207535A