Robot control system
A learning model-based approach enhances the accuracy of identifying linear objects in images by clustering feature amounts, addressing the challenges of same-color and intersecting objects in conventional methods.
Patent Information
- Application Number
- JP2025030182
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-20
AI Technical Summary
Conventional methods struggle to accurately identify multiple linear objects in an image, especially when they are of the same color or intersect in various patterns, leading to difficulties in high-accuracy recognition.
An apparatus utilizing a trained learning model through machine learning, specifically deep metric learning, to individually identify linear objects by extracting and clustering feature amounts, enabling separation of feature distributions from different objects.
Improves the accuracy of identifying multiple linear objects in an image by separating their feature distributions, allowing for precise recognition even in complex arrangements.
Smart Images

Figure 2025078666000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to an apparatus for identifying a linear object, an apparatus for learning to identify a linear object, and a robot control system. [Background technology]
[0002] Conventionally, devices for identifying linear objects are known. For example, International Publication No. 2019 / 017360 (Patent Document 1) discloses a 3D measurement device for linear objects. According to this 3D measurement device, the load of the matching process in a stereo 3D measurement method that measures the 3D position of a measurement point using the parallax of two cameras is reduced by extracting the color of a linear object to be measured from among multiple linear objects as a line image. As a result, it is possible to speed up the matching process that finds corresponding points on two images with different viewpoints. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2019 / 017360 [Non-patent literature]
[0004] [Non-Patent Document 1] Bert De Brabandere, Davy Neven, Luc Van Gool, Chunhong Pan, "Semantic Instance Segmentation with a Discriminative Loss Function", https: / / arxiv.org / abs / 1708.02551. Summary of the Invention [Problem to be solved by the invention]
[0005] However, in the method described in Patent Document 1, when multiple linear objects of the same color are present in an image, multiple intersections with the epipolar line are detected and corresponding points cannot be uniquely determined, making it impossible to recognize each linear object.
[0006] In addition, various patterns are assumed for the arrangement of linear objects handled in an actual manufacturing site. In particular, in a bulk pile in which a plurality of linear objects are randomly placed, the linear objects are assumed to intersect in a large number of patterns. It is difficult for a user to manually adjust the huge amount of data used for pattern matching so as to match all of the large number of patterns.
[0007] Therefore, when using the three-dimensional measuring device disclosed in Patent Document 1 or a conventional pattern matching method, it may be difficult to identify each of a plurality of linear objects included in an image with high accuracy.
[0008] The present disclosure has been made to solve the above-mentioned problems, and has an object to improve the accuracy of identifying a plurality of linear objects contained in an image. [Means for solving the problem]
[0009] An apparatus according to an aspect of the present disclosure includes an image acquisition unit and an estimation unit. The image acquisition unit acquires an image including at least one linear object. The estimation unit inputs the image to a trained learning model that has been subjected to machine learning for making an estimation for individually identifying linear objects in the image, and acquires an estimation result that individually identifies at least one linear object from among the linear objects included in the image from the trained learning model.
[0010] In the above device, the trained learning model may extract a plurality of feature amounts corresponding to a plurality of unit regions included in the image acquired by the image acquisition unit, and the estimation unit may output specific information based on a plurality of groups into which the plurality of feature amounts output from the trained learning model are classified.
[0011] In the above device, the estimation unit may extract multiple features such that, in a feature space in which the multiple features extracted by the trained learning model are distributed, a feature distribution originating from each of at least one linear object is separated from a feature distribution originating from an object other than the linear object.
[0012] In the above device, the estimation unit may classify the plurality of feature quantities into a plurality of groups by a non-hierarchical clustering method.
[0013] The machine learning may include deep metric learning, which aims to minimize a loss function for instance segmentation.
[0014] According to another aspect of the present disclosure, an apparatus includes a storage unit and a calculation unit. The storage unit stores a learning model for making an estimation to individually identify linear objects in an image including at least one linear object. The calculation unit uses machine learning to make the learning model a trained model.
[0015] A robot control system according to another aspect of the present disclosure includes an imaging unit, a robot, and a control unit. The imaging unit captures an image of at least one linear object. The robot has an arm that grasps the at least one linear object. The control unit controls the imaging unit and the robot. The control unit includes the device described above. The image acquisition unit acquires the image captured by the imaging unit. The control unit controls the robot based on the estimation result. Effect of the Invention
[0016] According to the device of the present disclosure, the accuracy of identifying linear objects contained in an image can be improved by using a learning model for making estimations to individually identify linear objects in an image that contains at least one linear object. [Brief description of the drawings]
[0017] [Figure 1] 1 is a diagram showing a configuration of a robot control system according to a first embodiment. [Diagram 2]2 is a diagram for explaining the input and output of the discrimination model in FIG. 1. [Diagram 3] FIG. 3 is a diagram for explaining a clustering process performed on a plurality of feature amounts extracted by the trained discrimination model of FIG. 2. [Figure 4] FIG. 4 is a diagram showing how the result of the clustering process in FIG. 3 is reflected in an image. [Diagram 5] 2 is a flowchart showing the flow of a linear object identification process performed by a calculation unit in FIG. 1. [Figure 6] FIG. 11 is a diagram showing a configuration of a robot control system according to a second embodiment. [Figure 7] FIG. 11 is a diagram showing a configuration of a robot control system according to a third embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0018] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the drawings, the same or corresponding parts are designated by the same reference characters, and their description will not be repeated in principle.
[0019] In the following description, an electric wire is used as an example of a linear object, but the linear object is not limited to an electric wire. In the present disclosure, the linear object may be any object having an elongated shape. Examples of linear objects include electric wires, wire harnesses, solder, strings, threads, fibers, glass fibers, optical fibers, tubes, and dried noodles. The linear object is not limited to an electric wire made of bundled thin wires, and also includes an electric wire made of a single wire.
[0020] In a state where a plurality of linear objects are supplied, the direction of each linear object is often indefinite. In particular, when the linear object is flexible and the shape is not fixed, the shape of the linear object is also indefinite, so various patterns are assumed for the rectangles including the linear object included in the image. In addition, a situation where a plurality of linear objects intersect is likely to occur, and in such a situation, it is difficult to detect one linear object by a rectangle. Therefore, depending on a framework for performing object detection by a rectangle, such as YOLO (You Only Look Once) or SSD (Single Shot MultiBox Detector), it is difficult to improve the identification accuracy of each of at least one linear object included in the image by machine learning. Therefore, in the following embodiment, a configuration for identifying each of at least one linear object by instance segmentation that extracts a plurality of feature amounts corresponding to a plurality of unit areas included in the image will be described. According to instance segmentation, since detection of a linear object by a rectangle is not performed, the identification accuracy of each of at least one linear object included in the image can be improved by machine learning.
[0021] [Embodiment 1] Fig. 1 is a diagram showing a configuration of a robot control system 10 according to embodiment 1. As shown in Fig. 1, the robot control system 10 has a robot 20, a camera 30 (imaging unit), and a control unit 40. A wire harness W formed of an electric wire W1 (linear object), an electric wire W2 (linear object), and an electric wire W3 (linear object) is arranged in a working space.
[0022] The robot 20 includes a robot arm 21 and a robot hand 22. A known articulated robot can be suitably used as the robot 20. The robot hand 22 is attached to the tip of the robot arm 21. The robot 20 grips any one of the electric wires W1 to W3 with a pair of gripping portions 23 of the robot hand 22.
[0023] The camera 30 may be any imaging device capable of capturing an image of the shapes of the electric wires W1 to W3, and is not particularly limited to a two-dimensional camera or a three-dimensional camera. Preferably, the camera 30 is a stereo camera. When the camera 30 is a stereo camera, it is suitable for calculating the three-dimensional position of a measurement point on a predetermined linear object by the principle of triangulation on two images captured from different viewpoints by the two stereo cameras after the camera 30 individually recognizes a plurality of linear objects. Although two cameras 30 are shown in FIG. 1, the robot control system according to the first embodiment may include only one camera 30.
[0024] The control unit 40 includes a calculation unit 41, a hard disk 42, a communication unit 43, an input / output unit 44, and an image acquisition unit 45. The image acquisition unit 45 communicates with the camera 30 via the communication unit 43 and acquires an image including the electric wires W1 to W3 from the camera 30. The calculation unit 41 identifies each of the electric wires W1 to W3 (instances) included in the image acquired from the image acquisition unit 45 and performs various calculations to determine a target linear object to be grasped by the robot hand 22. The calculation unit 41 outputs a gripping position (specific information) of the target linear object to be grasped to the robot 20 via the communication unit 43. Note that, when another device (for example, a robot controller or a control personal computer, etc.) that controls the operation of the robot 20 is provided between the control unit 40 and the robot 20, it is not necessary to directly output the gripping position to the robot 20, and the gripping position may be output to another device that controls the operation of the robot 20.
[0025] The hard disk 42 is a non-volatile storage device. The hard disk 42 stores a learning model M1, a machine learning program 421, a learning data set 422 including a plurality of learning data, and a linear object recognition program 423. In addition to the data shown in Fig. 1, the hard disk 42 stores, for example, an operating system program, and settings and outputs of various applications.
[0026] The learning model M1 is a neural network model for making an estimation to individually identify linear objects in an image including at least one linear object. The learning model M1 includes a fully convolutional network (e.g., U-Net).
[0027] The machine learning program 421 is a program for performing supervised learning on the learning model M1 using a learning dataset 422. The machine learning program 421 performs backpropagation on the learning model M1 by deep metric learning with a discriminative loss function (see Non-Patent Document 1) as a minimization target, and sets the learning model M1 as a trained model. The discriminative loss function is an example of a loss function for instance segmentation.
[0028] The learning data set 422 includes data in which an image including at least one linear object has been annotated with a correct answer for each instance of the linear object. The image including at least one linear object may be an image actually captured by an imaging device such as a camera, or may be an image artificially drawn by an image processing library (e.g., OpenCV (Open Source Computer Vision Library)). The annotation of the image may be performed manually by an operator, or automatically by an annotation program.
[0029] A large amount of image and annotation data is required to learn segmentation. An example of automatic annotation is a process in which an image actually captured by an imaging device is pre-processed to remove color information from an instance contained in the image and extract the contour of the instance, and the correct answer is assigned to the instance corresponding to the color information using the color information. In this case, the process of automatically assigning instances is facilitated by using an image capturing multiple linear objects of different colors.
[0030] The linear object identification program 423 uses the trained learning model M1 to identify at least one linear object included in an image captured by the camera 30. That is, the linear object identification program 423 is a program that performs instance segmentation on the image.
[0031] The calculation unit 41 includes a processor 411 and a memory 412. The processor 411 includes a CPU (Central Processing Unit). The processor 411 may further include a GPU (Graphic Processing Unit). The memory 412 is a volatile storage device, and includes, for example, a DRAM (Dynamic Random Access Memory). The processor 411 loads a program stored in the hard disk 42 into the memory 412 and executes the program, thereby realizing various functions of the robot control system 10. The processor 411, which executes the linear object identification program 423, functions as an estimation unit.
[0032] The input / output unit 44 receives operations from a user and outputs to the user the processing results of the calculation unit 41. The input / output unit 44 includes, for example, a mouse, a keyboard, a touch panel, a display, and a speaker.
[0033] FIG. 2 is a diagram for explaining the input and output of the learning model M1 in FIG. 1. As shown in FIG. 2, an image Im1 including electric wires W11, W12, and W13 is input to the learning model M1. The learning model M1 extracts a plurality of feature amounts corresponding to a plurality of pixels (a plurality of unit areas) included in the image Im1. Hereinafter, the coordinate space in which each of the plurality of feature amounts extracted by the learning model M1 is distributed is referred to as a feature amount space. Note that FIG. 2 shows a case in which each of the plurality of feature amounts is a two-dimensional vector and the feature amount is distributed in a two-dimensional feature amount space Fcd, but the number of dimensions of each of the plurality of feature amounts is not limited to two. Also, the number of linear objects included in the image Im1 in FIG. 2 is three, but the number of linear objects included in the image (input image) input to the learning model M1 is not limited to three.
[0034] When two features included in the plurality of features extracted by the learning model M1 originate from the same linear object included in the input image, the distance between the two features in the feature space Fcd is shortened by machine learning. When the two features originate from different objects included in the input image, the distance is extended by machine learning. The learning model M1 is made into a trained model by machine learning. The trained learning model M1 extracts a plurality of features such that, in a feature space in which a plurality of features are distributed, a feature distribution originating from at least one linear object is separated from a feature distribution originating from an object different from the linear object.
[0035] The feature space Fcd shown in FIG. 2 includes four feature distributions Dsb1, Dsb2, Dsb3, and Dsb4. In the following, the feature distributions Dsb1 to Dsb3 correspond to a plurality of pixels included in the electric wires W11 to W13, respectively. The feature distributions Dsb1 to Dsb3 are separated from each other. The feature distribution Dsb4 is a feature distribution that corresponds to a plurality of pixels included in the background of the image Im1 (area other than the electric wires W11 to W13). The trained learning model M1 identifies each of a plurality of instances included in the image Im1 and extracts features derived from the instances.
[0036] FIG. 3 is a diagram for explaining a clustering (classification) process performed on a plurality of feature amounts extracted by the trained learning model M1 of FIG. 2. As shown in FIG. 3, the plurality of feature amounts are classified into a plurality of clusters Cst1, Cst2, Cst3, and Cst4 (a plurality of groups) by a non-hierarchical clustering method (for example, k-means, DBSCAN (Density-based spatial clustering of applications with noise), or mean shift). The clusters Cst1 to Cst4 correspond to the feature amount distributions Dsb1 to Dsb4, respectively. That is, the clusters Cst1 to Cst3 correspond to the electric wires W11 to W13, respectively. Since the plurality of feature amounts extracted by the trained learning model M1 are distributed in a feature amount space biased for each instance included in the input image, it becomes possible to group the feature amount distributions corresponding to the instances included in the input image as one cluster.
[0037] FIG. 4 is a diagram showing a state in which the result of the clustering process (instance estimation result) in FIG. 3 is reflected in the image Im2. As shown in FIG. 4, for each of the clusters Cst1 to Cst4, a plurality of pixels corresponding to a plurality of feature amounts included in the cluster are specified, and thereby the background other than the electric wires included in the image Im1 in FIG. 2, and each of the electric wires W11, W12, and W13 are identified. In the image Im2, the electric wires W11 to W13 included in the image Im1 and the background are colored in different colors. The image Im2 (specific information) is displayed, for example, on a display included in the input / output unit 44 in FIG. 1. In addition, the gripping position of the target linear object to be gripped by the robot 20 in FIG. 1 is determined based on the estimation result of each of the electric wires W11 to W13 included in the image Im1.
[0038] Fig. 5 is a flowchart showing the flow of the linear object identification process performed by the calculation unit 41 in Fig. 1. The process shown in Fig. 5 is called when a gripping operation of a linear object by the robot 20 is requested by a main routine (not shown) that comprehensively controls the robot control system 10. Below, each step is simply abbreviated as S.
[0039] As shown in FIG. 5, the calculation unit 41 controls the camera 30 in S101 to capture an image of a linear object arranged in the working space, and proceeds to processing in S102. In S102, the calculation unit 41 uses the trained learning model M1 to extract a plurality of feature amounts corresponding to a plurality of pixels included in the image acquired by the camera 30 from the image, and proceeds to processing in S103. Note that in S102, the image acquired by the camera 30 may be pre-processed to remove color information from the instance included in the image and extract the contour of the instance, and then the extraction processing of a plurality of feature amounts may be performed on the pre-processed image. In S103, the calculation unit 41 identifies the instance included in the image input to the trained learning model M1 by the k-means method, and proceeds to processing in S104. In S104, the calculation unit 41 outputs information (specific information) based on the estimation result of the instance to the robot 20 and the input / output unit 44, and returns the processing to the main routine.
[0040] As described above, according to the device for identifying linear objects of embodiment 1, it is possible to improve the accuracy of identifying linear objects contained in an image, and even if an image contains multiple linear objects of the same type that are particularly difficult to distinguish, it is possible to identify each linear object individually.
[0041] [Embodiment 2] In the first embodiment, a device having both a function of identifying a linear object included in an image (inference function) and a function of learning to identify the linear object (learning function) has been described. In the second embodiment, a device having an inference function but no learning function will be described.
[0042] Fig. 6 is a diagram showing a configuration of a robot control system 10A according to embodiment 2. In the configuration of robot control system 10A, control unit 40 in Fig. 1 is replaced with 40A. In the configuration of control unit 40A, machine learning program 421 and learning data set 422 are removed from control unit 40 in Fig. 1, and learning model M1 is replaced with M1A. Since the rest is the same as in embodiment 1, description will not be repeated.
[0043] The learning model M1A is a model that has already been trained by a learning device different from the control unit 40A. In the robot control system 10A, since there is no need to perform machine learning on the learning model M1A, there is no need for the hard disk 42 to store a machine learning program and a learning data set.
[0044] As described above, the device for identifying a linear object according to the second embodiment can improve the accuracy of identifying a linear object included in an image.
[0045] [Embodiment 3] In the first and second embodiments, a device having an inference function for identifying a linear object included in an image is described. In the third embodiment, a device having a learning function for identifying a linear object included in an image without the inference function is described.
[0046] Fig. 7 is a diagram showing a configuration of a robot control system 10B according to the third embodiment. In the configuration of the robot control system 10B, the control unit 40 in Fig. 1 is replaced with 40B. In the configuration of the control unit 40B, the linear object identification program 423 is removed from the control unit 40 in Fig. 1. Since the rest is the same as in the first embodiment, the description will not be repeated. The robot control system 10B functions as a learning device that uses machine learning to make the learning model M1 into a learned model.
[0047] As described above, according to the device for learning to identify linear objects according to the third embodiment, it is possible to improve the accuracy of identifying linear objects included in an image.
[0048] The embodiments disclosed herein are also intended to be combined appropriately within a range that is not inconsistent. The embodiments disclosed herein should be considered to be illustrative and not restrictive in all respects. The scope of the present disclosure is indicated by the claims, not the above description, and is intended to include all modifications within the meaning and scope of the claims. [Explanation of symbols]
[0049] 10, 10A, 10B robot control system, 20 robot, 21 robot arm, 22 robot hand, 23 gripping unit, 30 camera, 40, 40A, 40B control unit, 41 calculation unit, 42 hard disk, 43 communication unit, 44 input / output unit, 45 image acquisition unit, 411 processor, 412 memory, 421 machine learning program, 422 learning dataset, 423 linear object recognition program, Cst1 to Cst4 clusters, Dsb1 to Dsb4 feature distribution, Im1 image, M1, M1A recognition model, W wire harness, W1 to W3, W11 to W13 electric wires.
Claims
1. An imaging unit that images a linear object; A robot having an arm that grasps a linear object; a control unit that controls the imaging unit and the robot, The control unit is Acquiring an image including at least one linear object captured by the imaging unit; inputting the image into a trained learning model that has undergone machine learning for making an estimation for individually identifying linear objects in the image, thereby obtaining an estimation result from the trained learning model that individually identifies at least one linear object among the linear objects included in the image; determining a target linear object based on the estimation result; A robot control system that controls the robot to grasp the target linear object.
2. The robot control system according to claim 1 , wherein the imaging unit includes a stereo camera.
3. the at least one linear object is a plurality of linear objects, The robot control system according to claim 1 , wherein the plurality of linear objects includes a plurality of intersecting linear objects.
Citation Information
Patent Citations
Image acquisition system for wire group processing
WO2016158282A1
Method and device for three-dimensional measurement of wire-like object
WO2019017360A1