Robot control system

By employing a learned machine learning model for instance segmentation, the apparatus effectively improves the identification accuracy of multiple linear objects in an image, overcoming challenges related to same-color intersections and varied patterns.

JP7692269B2Active Publication Date: 2025-06-13KURABO INDUSTRIES LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021005576
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-01-18
Publication Date
2025-06-13
Estimated Expiration
2041-01-18

AI Technical Summary

Technical Problem

Existing methods for identifying multiple linear objects in an image face challenges when objects of the same color intersect, leading to non-unique corresponding points and difficulties in accurately recognizing each object, especially in scenarios with random and varied patterns of linear objects.

Method used

An apparatus equipped with an image acquisition unit and an estimation unit that utilizes a learned machine learning model to individually identify linear objects in an image, employing instance segmentation to extract feature amounts from unit regions, thereby improving identification accuracy.

Benefits of technology

The proposed solution significantly enhances the identification accuracy of multiple linear objects in an image by leveraging machine learning and instance segmentation, effectively addressing the challenges of same-color intersections and varied patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007692269000001
    Figure 0007692269000001
  • Figure 0007692269000002
    Figure 0007692269000002
  • Figure 0007692269000003
    Figure 0007692269000003
Patent Text Reader

Abstract

To provide an apparatus and a robot control system that improve an identification accuracy of a linear object included in an image.SOLUTION: In a robot control system 10 including a robot 20, a camera 30 (imaging unit), and a control unit 40, an apparatus 40 includes an image acquisition unit 45, and a processor 411, a learning model M1, and a linear object identification program 423 functioning as an estimation unit. The image acquisition unit 45 acquires an image including at least one of linear objects W1 to W3. The estimation unit inputs the image to the trained learning model M1 with machine learning for estimating to individually identify the linear objects W1 to W3 in the image to obtain an estimation result in which at least one of the linear objects W1 to W3 included in the image are identified individually from the trained learning model M1.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an apparatus for identifying a linear object, an apparatus for learning the identification of a linear object, and a robot control system.

Background Art

[0002] Conventionally, an apparatus for identifying a linear object has been known. For example, International Publication No. 2019 / 017360 (Patent Document 1) discloses a three-dimensional measuring apparatus for a linear object. According to the three-dimensional measuring apparatus, by extracting the color of the linear object to be measured as a line image from among a plurality of linear objects, the load of the matching process in a stereo three-dimensional measuring method for measuring the three-dimensional position of a measurement point using the parallax of two cameras is reduced. As a result, it becomes possible to speed up the matching process of finding corresponding points on two images with different viewpoints.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the method described in Patent Document 1, when there are multiple linear objects of the same color in the image, multiple intersections with the epipolar lines may be detected, and corresponding points cannot be uniquely determined, resulting in a situation where each linear object may not be recognized.

[0006] In addition, various patterns are assumed for the arrangement of linear objects handled in an actual manufacturing site. In particular, in a stacked state where multiple linear objects are randomly placed, it is assumed that the linear objects intersect in a large number of patterns. It is difficult for the user to manually adjust the huge amount of data used for pattern matching so as to conform to all of the large number of patterns.

[0007] Therefore, according to the three-dimensional measurement device or the conventional pattern matching method disclosed in Patent Document 1, it may be difficult to accurately identify each of the multiple linear objects included in the image.

[0008] The present disclosure has been made to solve the above-described problems, and an object thereof is to improve the identification accuracy of multiple linear objects included in an image.

Means for Solving the Problems

[0009] An apparatus according to an aspect of the present disclosure includes an image acquisition unit and an estimation unit. The image acquisition unit acquires an image including at least one linear object. The estimation unit inputs the image into a learned learning model in which machine learning for individually identifying the linear objects in the image is performed, and obtains an estimation result of individually identifying at least one of the linear objects included in the image from the learned learning model.

[0010] An apparatus according to another aspect of the present disclosure includes a storage unit and an arithmetic unit. A learning model for performing an estimation of individually identifying the linear objects in an image including at least one linear object is stored in the storage unit. The arithmetic unit makes the learning model a learned model by machine learning.

Effects of the Invention

[0011] According to the device of the present disclosure, the identification accuracy of the linear objects included in an image can be improved by a learning model for making an estimation to individually identify the linear objects in the image including at least one linear object.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Modes for Carrying Out the Invention

[0013] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the drawings, the same or corresponding parts are denoted by the same reference numerals, and the description thereof will not be repeated in principle.

[0014] In the following description, the case where an electric wire is used as an example of the linear object is described, but the linear object is not limited to the electric wire. In the present disclosure, the linear object may be any object as long as it has an elongated shape. Examples of the linear object include an electric wire, a wire harness, solder, a string, a thread, a fiber, a glass fiber, an optical fiber, a tube, or dried noodles. The linear object is not limited to an electric wire in which thin wires are bundled, and includes an electric wire composed of a single wire.

[0015] In a state where a plurality of linear objects are supplied, the directions of the respective linear objects are often indeterminate. In particular, when the linear objects are flexible and their shapes are not fixed, since the shapes of the linear objects are also indeterminate, various patterns are assumed for the rectangles including the linear objects included in the image. In addition, a situation where a plurality of linear objects intersect easily occurs, and it is difficult to detect one linear object by a rectangle in such a situation. Therefore, depending on a framework for performing object detection by a rectangle such as YOLO (You Only Look Once) or SSD (Single Shot MultiBox Detector), it is difficult to improve the identification accuracy of each of at least one linear object included in the image by machine learning. Therefore, in the following embodiments, a configuration for identifying each of at least one linear object by instance segmentation that extracts a plurality of feature amounts corresponding to a plurality of unit regions included in the image will be described. According to instance segmentation, since detection of a linear object by a rectangle is not performed, the identification accuracy of each of at least one linear object included in the image can be improved by machine learning.

[0016] [Embodiment 1] FIG. 1 is a diagram showing the configuration of a robot control system 10 according to Embodiment 1. As shown in FIG. 1, the robot control system 10 includes a robot 20, a camera 30 (imaging unit), and a control unit 40. In the work space, a wire harness W formed of a power line W1 (linear object), a power line W2 (linear object), and a power line W3 (linear object) is arranged.

[0017] The robot 20 includes a robot arm 21 and a robot hand 22. As the robot 20, a known articulated robot can be preferably used. A robot hand 22 is attached to the tip of the robot arm 21. The robot 20 grips any one of the power lines W1 to W3 by a pair of gripping portions 23 of the robot hand 22.

[0018] The camera 30 may be any imaging device as long as it can image the shapes of the electric wires W1 to W3, and is not particularly limited to a 2D camera or a 3D camera. Preferably, the camera 30 uses a stereo camera. When the camera 30 is a stereo camera, after the camera 30 individually recognizes a plurality of linear objects, it is suitable for calculating the three-dimensional positions of measurement points on a predetermined linear object by the principle of triangulation on two images captured from different viewpoints by two stereo cameras. Although two cameras 30 are shown in FIG. 1, the camera 30 included in the robot control system according to the first embodiment may be one camera.

[0019] The control unit 40 includes an arithmetic unit 41, a hard disk 42, a communication unit 43, an input / output unit 44, and an image acquisition unit 45. The image acquisition unit 45 communicates with the camera 30 via the communication unit 43 and acquires an image including the electric wires W1 to W3 from the camera 30. The arithmetic unit 41 identifies each of the electric wires W1 to W3 (instances) included in the image acquired from the image acquisition unit 45 and performs various operations for determining the target linear object to be grasped by the robot hand 22. The arithmetic unit 41 outputs the grasping position (specific information) of the target linear object to be grasped to the robot 20 via the communication unit 43. When another device (for example, a robot controller or a control personal computer, etc.) for controlling the operation of the robot 20 is provided between the control unit 40 and the robot 20, it is not necessary to directly output the grasping position to the robot 20, and the grasping position may be output to another device for controlling the operation of the robot 20.

[0020] The hard disk 42 is a non-volatile storage device. The hard disk 42 stores a learning model M1, a machine learning program 421, a learning data set 422 including a plurality of learning data, and a linear object identification program 423. In addition to the data shown in FIG. 1, the hard disk 42 stores, for example, an operating system program, and settings and outputs of various applications.

[0021] The learning model M1 is a neural network model for estimating and individually identifying linear objects in an image containing at least one linear object. The learning model M1 includes a fully convolutional network (e.g., U-Net).

[0022] The machine learning program 421 is a program for performing supervised learning using the learning dataset 422 for the learning model M1. The machine learning program 421 performs backpropagation on the learning model M1 by deep distance learning that targets minimizing the Discrimitive loss function (see Non-Patent Document 1), and sets the learning model M1 as a learned model. The Discrimitive loss function is an example of a loss function for instance segmentation.

[0023] The learning dataset 422 includes data obtained by performing a process (annotation) of attaching a correct answer to each instance of a linear object in an image containing at least one linear object. The image containing at least one linear object may be an image actually captured by an imaging device such as a camera, or may be an image artificially drawn by an image processing library (e.g., OpenCV (Open Source Computer Vision Library)). The annotation of the image may be a manual operation by an operator, or may be automatically performed by an annotation program.

[0024] To perform segmentation learning, a large amount of image and annotation data is required. Examples of automatic annotation include performing preprocessing to extract the contour lines of an instance by removing color information from the instance included in an image actually captured by an imaging device, and using the color information to attach a correct answer to the instance corresponding to the color information. At this time, by using an image capturing a plurality of linear objects with different colors, the process of automatically attaching instances becomes easier.

[0025] The linear object recognition program 423 uses the learned learning model M1 to identify each of at least one linear object included in the image captured by the camera 30. That is, the linear object recognition program 423 is a program that performs instance segmentation on the image.

[0026] The arithmetic unit 41 includes a processor 411 and a memory 412. The processor 411 includes a CPU (Central Processing Unit). The processor 411 may further include a GPU (Graphic Processing Unit). The memory 412 is a volatile storage device and includes, for example, a DRAM (Dynamic Random Access Memory). The processor 411 reads and executes the program stored in the hard disk 42 in the memory 412 to realize various functions of the robot control system 10. The processor 411 that executes the linear object recognition program 423 functions as an estimation unit.

[0027] The input / output unit 44 receives operations from the user and outputs the processing results of the arithmetic unit 41 to the user. The input / output unit 44 includes, for example, a mouse, a keyboard, a touch panel, a display, and a speaker.

[0028] FIG. 2 is a diagram for explaining the input and output of the learning model M1 in FIG. 1. As shown in FIG. 2, an image Im1 including electric wires W11, W12, and W13 is input to the learning model M1. The learning model M1 extracts a plurality of feature amounts corresponding to a plurality of pixels (a plurality of unit regions) included in the image Im1. Hereinafter, the coordinate space in which each of the plurality of feature amounts extracted by the learning model M1 is distributed is called a feature amount space. In FIG. 2, each of the plurality of feature amounts is a two-dimensional vector, and the case where the feature amount is distributed in the two-dimensional feature amount space Fcd is shown. However, the number of dimensions of each of the plurality of feature amounts is not limited to 2. Also, in FIG. 2, the number of linear objects included in the image Im1 is 3. However, the number of linear objects included in the image (input image) input to the learning model M1 is not limited to 3.

[0029] When two feature quantities included in a plurality of feature quantities extracted by the learning model M1 are derived from the same linear object included in the input image, the distance between the two feature quantities in the feature quantity space Fcd is shortened by machine learning. When the two feature quantities are respectively derived from different objects included in the input image, the distance is extended by machine learning. The learning model M1 is a learned model by machine learning. The learned learning model M1 extracts a plurality of feature quantities so that the feature quantity distribution derived from each of at least one linear object is separated from the feature quantity distribution derived from an object different from the linear object in the feature quantity space where the plurality of feature quantities are distributed.

[0030] The feature quantity space Fcd shown in FIG. 2 includes four feature quantity distributions Dsb1, Dsb2, Dsb3, and Dsb4. Hereinafter, it is assumed that the feature quantity distributions Dsb1 to Dsb3 are feature quantity distributions corresponding to a plurality of pixels included in the electric wires W11 to W13, respectively. The feature quantity distributions Dsb1 to Dsb3 are separated from each other. The feature quantity distribution Dsb4 is a feature quantity distribution corresponding to a plurality of pixels included in the background (region other than the electric wires W11 to W13) of the image Im1. The learned learning model M1 identifies each of a plurality of instances included in the image Im1 and extracts the feature quantity derived from the instance.

[0031] FIG. 3 is a diagram for explaining the clustering (classification) process performed on a plurality of feature amounts extracted by the learned learning model M1 of FIG. 2. As shown in FIG. 3, the plurality of feature amounts are classified into a plurality of clusters Cst1, Cst2, Cst3, Cst4 (a plurality of groups) by a non-hierarchical clustering method (for example, the k-means method, DBSCAN (Density-based spatial clustering of applications with noise), or mean shift). The clusters Cst1 to Cst4 respectively correspond to the feature amount distributions Dsb1 to Dsb4. That is, the clusters Cst1 to Cst3 respectively correspond to the electric wires W11 to W13. Since the distributions of the plurality of feature amounts extracted by the learned learning model M1 are biased in the feature amount space for each instance included in the input image, it is possible to group the feature amount distributions corresponding to the instances included in the input image as one cluster.

[0032] FIG. 4 is a diagram showing a state in which the result of the clustering process of FIG. 3 (the estimation result of the instance) is reflected in the image Im2. As shown in FIG. 4, for each of the clusters Cst1 to Cst4, by specifying a plurality of pixels respectively corresponding to the plurality of feature amounts included in the cluster, the background other than the electric wires, the electric wires W11, W12, and W13 included in the image Im1 of FIG. 2 are identified. In the image Im2, the electric wires W11 to W13 and the background included in the image Im1 are colored in different colors from each other. The image Im2 (specific information) is displayed, for example, on the display included in the input / output unit 44 of FIG. 1. Further, based on the estimation result of each of the electric wires W11 to W13 included in the image Im1, the gripping position of the target linear object to be gripped by the robot 20 of FIG. 1 is determined.

[0033] FIG. 5 is a flowchart showing the flow of the identification process of the linear object performed by the calculation unit 41 in FIG. 1. The process shown in FIG. 5 is called by a main routine (not shown) that integrally controls the robot control system 10 when a gripping operation of a linear object by the robot 20 is requested. Hereinafter, the steps are simply described as S.

[0034] As shown in FIG. 5, in S101, the calculation unit 41 controls the camera 30 to image the linear object arranged in the work space, and proceeds to S102. In S102, the calculation unit 41 uses the learned learning model M1 to extract a plurality of feature amounts corresponding to each of the plurality of pixels included in the image acquired by the camera 30, and proceeds to S103. In S102, before the preprocessing of extracting the contour line of the instance excluding the color information from the instance included in the image acquired by the camera 30 is performed on the image acquired by the camera 30, the extraction process of a plurality of feature amounts may be performed on the image on which the preprocessing has been performed. In S103, the calculation unit 41 identifies the instances included in the image input to the learned learning model M1 by the k-means method, and proceeds to S104. In S104, the calculation unit 41 outputs information (specific information) based on the estimation result of the instance to the robot 20 and the input / output unit 44, and returns the process to the main routine.

[0035] As described above, according to the apparatus for identifying a linear object according to Embodiment 1, the identification accuracy of the linear object included in the image can be improved. Even if there are a plurality of linear objects of the same type that are particularly difficult to distinguish from each other in the image, each linear object can be individually identified.

[0036] [Embodiment 2] In Embodiment 1, an apparatus having both a function (inference function) of identifying a linear object included in an image and a function (learning function) of learning the identification of the linear object has been described. In Embodiment 2, an apparatus having no learning function but having an inference function will be described.

[0037] FIG. 6 is a diagram showing the configuration of the robot control system 10A according to the second embodiment. The configuration of the robot control system 10A is a configuration in which the control unit 40 in FIG. 1 is replaced by 40A. The configuration of the control unit 40A is a configuration in which the machine learning program 421 and the learning dataset 422 are removed from the control unit 40 in FIG. 1, and the learning model M1 is replaced by M1A. Since the rest is the same as in the first embodiment, the description will not be repeated.

[0038] The learning model M1A is a pre-trained model by a learning device different from the control unit 40A. In the robot control system 10A, since it is not necessary to perform machine learning on the learning model M1A, it is not necessary to store the machine learning program and the learning dataset in the hard disk 42.

[0039] As described above, according to the device for identifying a linear object according to the second embodiment, the identification accuracy of the linear object included in the image can be improved.

[0040] [Embodiment 3] In the first and second embodiments, an apparatus having an inference function for identifying a linear object included in an image has been described. In the third embodiment, an apparatus having no such inference function but having a learning function for identifying a linear object included in an image will be described.

[0041] FIG. 7 is a diagram showing the configuration of the robot control system 10B according to the third embodiment. The configuration of the robot control system 10B is a configuration in which the control unit 40 in FIG. 1 is replaced by 40B. The configuration of the control unit 40B is a configuration in which the linear object identification program 423 is removed from the control unit 40 in FIG. 1. Since the rest is the same as in the first embodiment, the description will not be repeated. The robot control system 10B functions as a learning device that makes the learning model M1 a pre-trained model by machine learning.

[0042] As described above, according to the device for learning the identification of a linear object according to the third embodiment, the identification accuracy of the linear object included in the image can be improved.

[0043] Each embodiment disclosed this time is also planned to be implemented by appropriately combining them within a non - conflicting range. It should be considered that all the embodiments disclosed this time are illustrative in all respects and not restrictive. The scope of the present disclosure is indicated by the claims rather than the above description, and it is intended that all changes within the meaning and scope equivalent to the claims are included.

Explanation of Signs

[0044] 10, 10A, 10B robot control system, 20 robot, 21 robot arm, 22 robot hand, 23 gripping part, 30 camera, 40, 40A, 40B control unit, 41 arithmetic unit, 42 hard disk, 43 communication unit, 44 input / output unit, 45 image acquisition unit, 411 processor, 412 memory, 421 machine learning program, 422 learning dataset, 423 linear object identification program, Cst1 - Cst4 cluster, Dsb1 - Dsb4 feature distribution, Im1 image, M1, M1A identification model, W wire harness, W1 - W3, W11 - W13 electric wire.

Claims

1. An imaging unit including a camera for imaging a linear object, a robot having an arm for gripping the linear object, and a control unit for controlling the imaging unit and the robot, wherein the control unit includes an image acquisition unit and a calculation unit, the image acquisition unit acquires an image including a plurality of linear objects imaged by the imaging unit, the calculation unit inputs the image into a learned learning model in which machine learning for individually identifying the linear objects in the image is performed, and obtains an estimation result of individually identifying a linear object including at least one of a flexible linear object and a bundle of thin wires among the linear objects included in the image from the learned learning model, including an estimation unit, the calculation unit performs a calculation for determining a target linear object and a gripping position of the target linear object to be gripped based on the estimation result, the control unit controls the robot based on the gripping position, the plurality of linear objects are of the same color, flexible, and have an indeterminate shape, a robot control system.

2. the learned learning model extracts a plurality of feature amounts respectively corresponding to a plurality of unit regions included in the image acquired by the image acquisition unit, the estimation unit outputs specific information based on a plurality of groups into which the plurality of feature amounts output from the learned learning model are classified, the robot control system according to claim 1.

3. the estimation unit extracts the plurality of feature amounts so that the feature amount distribution derived from each of the plurality of linear objects is separated from the feature amount distribution derived from an object different from the linear object in a feature amount space in which the plurality of feature amounts extracted by the learned learning model are distributed, the robot control system according to claim 2.

4. the estimation unit classifies the plurality of feature amounts into the plurality of groups by a non-hierarchical clustering method, the robot control system according to claim 2 or 3.

5. the machine learning includes deep distance learning for minimizing a loss function of instance segmentation, the robot control system according to claim 1.

6. The control unit further includes a storage unit in which a learning model for performing an estimation for individually identifying the linear objects in the image is stored, the calculation unit further performs a calculation for making the learning model a learned model by machine learning, the robot control system according to claim 1.

7. The robot control system according to claim 1, wherein the image includes an image captured by a stereo camera.

Citation Information

Patent Citations

  • Instance segmentation method and device, electronic device, program, and medium

    JP2021507388A

  • Method and device for three-dimensional measurement of wire-like object

    WO2019017360A1