Recognition device, model generation device, and work system
The recognition device enhances object recognition stability and accuracy by detecting feature regions and aligning representative points with predefined reference points, addressing the limitations of conventional technologies in complex environments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- DENSO CORP
- Filing Date
- 2024-10-10
- Publication Date
- 2026-04-22
AI Technical Summary
Conventional recognition technologies are insufficient in terms of recognition stability and accuracy, especially for complex environments and various objects, necessitating improved methods for accurately determining the position and orientation of objects.
A recognition device that includes a feature region detection unit, a representative point extraction unit, and a position and orientation estimation unit, which detects at least three feature regions from imaging data, calculates representative points, and estimates the position and orientation by aligning these points with predefined reference points, eliminating the need for surface selection and reducing computational load.
Enables high-accuracy and stable recognition of the position and orientation of a wide variety of objects, even with complex shapes, by defining feature regions and reference representative points in a trained model, thereby improving recognition stability and accuracy.
Smart Images

Figure 2026068242000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a recognition device, a model generation device, and a work system.
Background Art
[0002] In recent years, with the decline of the labor force population mainly in developed countries, the need for automation of work in, for example, factories and farms has been rapidly increasing. For the automation of work, a technology for recognizing the position and orientation of an object is indispensable, and in particular, the development of a recognition technology that is stable and highly accurate even for complex environments and various objects is required.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, the conventional recognition technology was not sufficient in terms of recognition stability and accuracy, and there was room for improvement.
[0005] The present disclosure has been made in view of the above circumstances, and an object thereof is to provide a recognition device, a model generation device, and a work system that can stably and highly accurately recognize the position and orientation of an object.
Means for Solving the Problems
[0006] The recognition devices (201, 202) of this disclosure include: a feature region detection unit (21) that detects at least three feature regions (62) of an object (61) from imaging data (67) obtained by imaging the object (61); a representative point extraction unit (23) that calculates extracted representative points (69) from the imaging data and the feature regions and extracts at least three extracted representative points and extracted label information (68) for identifying each extracted representative point; and a position and orientation estimation unit (24) that estimates the position and orientation of an object by aligning the extracted representative points with the reference representative points where the extracted label information corresponding to the extracted representative points and the reference label information (64) for identifying at least three reference representative points having a predefined correct positional relationship coincide.
[0007] According to this method, the recognition device estimates the position and orientation of the object itself from the image data of the object, eliminating the need for tasks such as selecting a surface for each scene, and allowing the position and orientation of any location within the object to be determined. Furthermore, since the position and orientation of the entire object can be estimated if at least three feature regions are detected, it is possible to avoid excessive computational load on neural networks, for example, and prevent a decrease in estimation accuracy. In addition, since the surface of the object does not need to have a constant curvature, it can be adapted to objects of any shape. For these reasons, using this recognition device makes it possible to recognize the position and orientation of a wide variety of objects with high accuracy and stability.
[0008] The model generation device (10) of this disclosure generates a trained model (54) used to estimate the position and orientation of an object (61) from imaging data (67) obtained by imaging the object, and includes a definition unit (11) that pre-defines, in object data (52) representing the object, a feature region (62) which is a characteristic area of the object, a reference representative point (66) which is a representative point within the feature region, and reference label information (64) which identifies the feature region or the reference representative point.
[0009] According to this, the model generation device can efficiently and accurately generate a trained model that can recognize the position and orientation of an object by defining feature regions for object data and pre-defining reference representative points corresponding to those feature regions. Furthermore, because the feature regions of the object and their reference representative points are pre-defined in this trained model, it is possible to estimate the position and orientation with high accuracy and robustness even if the shape and orientation of the object are complex.
[0010] Furthermore, the work system (30) of this disclosure comprises the above-mentioned recognition devices (201, 202) and a work device (32) that performs a predetermined operation on an object recognized by the recognition devices.
[0011] According to this, by using the recognition device described above, the accuracy and stability of object recognition will be improved, and as a result, the work performed by the work device will be able to be carried out accurately and stably. [Brief explanation of the drawing]
[0012] [Figure 1] Figure 1 is a block diagram showing an example of a schematic configuration of a model generation device according to an embodiment. [Figure 2] Figure 2 shows an example of the hardware configuration of a model generation device according to an embodiment. [Figure 3] Figure 3 shows an example of the definition process performed in the definition unit of the model generation device according to the embodiment. [Figure 4] Figure 4 shows an example of manually defining feature regions in a model generation device according to an embodiment. [Figure 5] Figure 5 shows an example of how feature regions are automatically defined in the model generation device according to the embodiment. [Figure 6] Figure 6 shows an example of setting a reference representative point in the representative point setting unit of a model generation device according to an embodiment. [Figure 7] Figure 7 shows an example of correct data set in the representative point setting unit in the model generation device according to the embodiment. [Figure 8] FIG. 8 is a diagram showing an example of a dataset generated by a dataset generation unit in a model generation apparatus according to an embodiment. [Figure 9] FIG. 9 is a diagram showing an example of an input and an output in a learning unit in a model generation apparatus according to an embodiment. [Figure 10] FIG. 10 is a block diagram showing an example of a schematic configuration of a recognition apparatus according to an embodiment. [Figure 11] FIG. 11 is a diagram showing an example of a hardware configuration of a recognition apparatus according to an embodiment. [Figure 12] FIG. 12 is a diagram showing an example of an input and an output in a feature region detection unit in a recognition apparatus according to an embodiment. [Figure 13] FIG. 13 shows an example of a process in a recognition apparatus according to an embodiment. (A) is a diagram showing an example in which a reliability is set for extraction label information, (B) is a diagram showing an example in which a feature region is selected based on the reliability, and (C) is a diagram showing an example in which representative points are extracted from the selected feature region. [Figure 14] FIG. 14 is a diagram showing an example in which a reference representative point and an extracted representative point are aligned in a position and orientation estimation processing unit in a recognition apparatus according to an embodiment. [Figure 15] FIG. 15 is a diagram showing an example in which an input point group is determined based on a feature region in a precise estimation unit in a recognition apparatus according to an embodiment. [Figure 16] FIG. 16 is a diagram showing an example in which an input point group is determined based on extraction label information in a precise estimation unit in a recognition apparatus according to an embodiment. [Figure 17] FIG. 17 shows an example when a feature region is detected in a feature region detection unit in a recognition apparatus according to an embodiment. (A) is a diagram showing an example when four or more feature regions are detected, and (B) is a diagram showing an example when some feature regions are hidden by some object. [Figure 18] FIG. 18 is a diagram showing an example when a plurality of representative point groups are defined in a position and orientation estimation unit in a recognition apparatus according to an embodiment. [Figure 19]FIG. 19 is a block diagram showing an example of the schematic configuration of the work system according to the embodiment. [Figure 20] FIG. 20 shows an example of the processing in the definition unit in the recognition device according to the embodiment. (A) is a diagram showing an example when a specific meaning is given among a plurality of reference label information, and (B) is a diagram showing an example when a part of the feature region of the object is hidden by another object. [Figure 21] FIG. 21 is a block diagram showing another example of the schematic configuration of the recognition device according to the embodiment. [Figure 22] FIG. 22 shows an example of the processing in the grouping unit in the recognition device according to the embodiment. (A) is a diagram showing before grouping, and (B) is a diagram showing after grouping.
Mode for Carrying Out the Invention
[0013] Hereinafter, the model generation device, the recognition device, and the work system according to the embodiment will be described with reference to the drawings. In each embodiment, the same components may be denoted by the same reference numerals and the description thereof may be omitted.
[0014] First, the model generation device 10 will be described. The model generation device 10 shown in FIG. 1 is a device for generating a learned model 54 for use in, for example, the first recognition device 201 shown in FIG. 10. As shown in FIG. 1, the model generation device 10 includes, for example, a definition unit 11, a representative point setting unit 12, a dataset creation unit 13, and a learning unit 14.
[0015] The model generation device 10 may be a dedicated computer, but it can also be implemented by installing the model generation program 51 on a general-purpose personal computer or server. The hardware configuration of the model generation device 10 can include a first processor 1a, a first main memory 2a, a first input unit 3a, a first output unit 4a, and a first auxiliary memory 5a, as shown in Figure 2. The first processor 1a is configured to include a microcomputer such as a CPU and performs calculation processing. The first main memory 2a is composed of storage areas such as ROM, RAM, and rewritable flash memory.
[0016] The first input unit 3a is a user interface such as a mouse, keyboard, or touch panel, and accepts input operations from the user. The first output unit 4a is a user interface such as a display, and presents various information to the user. The model generation device 10 can also be configured to communicate with an external computer via a telecommunications line such as the Internet or a LAN.
[0017] The first auxiliary storage device 5a stores, for example, a model generation program 51 and object data 52. The model generation program 51 is a computer program that causes a computer to execute a process to generate a trained model 54. In other words, the model generation program 51 is a computer program that virtually implements the definition unit 11, representative point setting unit 12, dataset creation unit 13, and learning unit 14 shown in Figure 1 on a computer. The model generation device 10 can virtually implement the definition unit 11, representative point setting unit 12, dataset creation unit 13, and learning unit 14 on a computer by having the first processor 1a read the model generation program 51 from the first auxiliary storage device 5a, expand it into the first main memory device 2a, and execute it.
[0018] In other words, the definition unit 11, the representative point setting unit 12, the dataset creation unit 13, and the learning unit 14 can be configured as functional units that are virtually realized by executing the model generation program 51 in, for example, the first processor 1a. The model generation device 10 can be configured with the same or common hardware for the definition unit 11, the representative point setting unit 12, the dataset creation unit 13, and the learning unit 14 shown in Figure 1, or it can be configured with different hardware.
[0019] The first auxiliary storage device 5a consists of a tangible, non-temporary computer-readable medium. Examples of the first auxiliary storage device 5a include, but are not limited to, HDDs (Hard Disk Drives), SSDs (Solid State Drives), magnetic disks, magneto-optical disks, CD-ROMs (Compact Disc Read Only Memory), DVD-ROMs (Digital Versatile Disc Read Only Memory), and semiconductor memory. The first auxiliary storage device 5a may be an internal medium directly connected to the bus of the computer constituting the model generation device 10, or it may be an external medium connected to the model generation device 10 via a telecommunications line such as the Internet or a LAN. Furthermore, when the model generation program 51 is delivered to the model generation device 10 via a telecommunications line, the model generation device 10, upon receiving the delivery, expands the model generation program 51 into the first main storage device 2a and executes it, thereby realizing the definition unit 11, the representative point setting unit 12, the dataset creation unit 13, and the learning unit 14.
[0020] Furthermore, the implementation of the definition unit 11, the representative point setting unit 12, the dataset creation unit 13, and the learning unit 14 is not limited to the combination of the hardware and the model generation program 51 described above. The implementation of the definition unit 11, the representative point setting unit 12, the dataset creation unit 13, and the learning unit 14 may be carried out by a single piece of hardware, such as an integrated circuit implementing the model generation program 51, or some functions may be carried out by dedicated hardware, and the rest by a combination of hardware and the model generation program 51.
[0021] Object data 52 is data for representing the object to be recognized on a computer. Object data 52 can consist of, for example, CAD data of the object to be recognized, and includes 2D and 3D information of the object. In this embodiment, object data 52 is stored in the first auxiliary storage device 5a, but it may also be acquired by the model generation device 10 from, for example, an external data server as appropriate. Furthermore, in the following description, the object represented on the computer based on object data 52 may be referred to as the object model 521.
[0022] As shown in Figure 1, for example, the model generation device 10 receives object data 52 as training data and generates correct answer data 53 by sequentially processing it in the definition unit 11 and the representative point setting unit 12. The object data 52 is a 3D model containing detailed shape information of the object, and preferably includes information on all surface shapes of the object.
[0023] The definition unit 11 executes a definition process. The definition process includes defining three or more feature regions 62 of the object model 521 from the object data 52 shown in Figure 3(A), as shown in Figure 3(B). A feature region 62 is, for example, a bounding box, which is a partial region that surrounds the object model 521 in an image or video. A feature region 62 is, for example, a region arbitrarily selected as a distinctive part of the object model 521, such as a face, edge, or vertex.
[0024] The feature region 62 can be set manually by an operator using, for example, the first input unit 3a and the first output unit 4a. In this case, as shown in Figure 4, the model generation device 10 displays object data 52, including the object model 521 displayed as a 3D image, and a cursor 71 indicating the position for defining the feature region 62, on the first output unit 4a. That is, the object model 521 and the feature region 62 are visualized on the first output unit 4a. The operator operates the first input unit 3a to move the cursor 71 in a three-dimensional space including the object model 521, for example, to specify the position for defining the feature region 62. In this case, the operator can rotate the object model 521, displayed as a 3D image, by operating the first input unit 3a while looking at the object model 521 and cursor 71 displayed on the first output unit 4a. After specification, the operator can define the feature region 62 in the three-dimensional space including the object model 521 by selecting the confirm button 72.
[0025] Furthermore, the setting of the feature region 62 can be performed automatically, without, for example, the operation of the cursor 71 by the operator as described above. In this case, the model generation device 10, for example as shown in Figure 5, divides the object model 521 into multiple voxels 63 in three-dimensional space and performs a process to define three or more randomly selected voxels 63 as the feature region 62.
[0026] The definition process includes assigning reference label information 64 to the defined feature regions 62, as shown in Figure 3(B). The reference label information 64 is used to identify the feature regions 62 and to describe the location and properties of those feature regions 62 when estimating the position and orientation using the trained model 54 later. The reference label information 64 includes identification information to uniquely identify each defined feature region 62, such as a unique number or symbol for each feature region 62. In Figure 3, etc., a common code is assigned to each feature region 62 and each reference label information 64 for the sake of simplicity, but when explaining them separately, they will be referred to as the first reference label information 641, the second reference label information 642, the third reference label information 643, and the fourth reference label information 644.
[0027] The reference label information 64 may include information indicating the type of defined feature area 62, such as "edge portion," "corner portion," or "plane portion." Furthermore, the reference label information 64 may include information indicating the purpose for which the defined feature area 62 is used or what role it plays. For example, in a system intended for use in a picking system, the reference label information 64 may include information such as whether the defined feature area 62 is a graspable area or a non-graspable area. The assignment of the reference label information 64 may be done manually by an operator, or it may be done automatically based on 3D CAD data of the object. In the following description, the reference label information 64 assigned by the definition unit 11 may be referred to simply as the reference label information 64.
[0028] As shown in Figure 1, the definition unit 11 outputs, for example, the defined feature region 62 and reference label information 64 to the representative point setting unit 12. The representative point setting unit 12 receives the feature region 62 and reference label information 64 from the definition unit 11 and executes the representative point setting process. The representative point setting process is a process in which, for example, as shown in Figure 6, for each defined feature region 62, a three-dimensional point that represents that feature region 62 is calculated and set as a reference representative point 66.
[0029] The reference representative point 66 has three-dimensional coordinate values of x, y, and z. The representative point setting unit 12 can set, for example, the center point or centroid of the defined feature region 62 as the reference representative point. The reference representative point 66 is linked to the reference label information 64 assigned to the original feature region 62, as shown in Figure 7, and stored as the correct positional relationship data 53.
[0030] Furthermore, the model generation device 10 receives object data 52 as training data, as shown in Figure 1, for example, and generates a trained model 54 by sequentially executing processing in the definition unit 11, the dataset creation unit 13, and the learning unit 14. The definition unit 11 outputs, for example, the defined feature region 62 and reference label information 64 to the dataset creation unit 13, as shown in Figure 1. The dataset creation unit 13 receives the feature region 62 and reference label information 64 from the definition unit 11 and executes the dataset creation process.
[0031] The dataset creation process includes generating a training dataset 58 based on object data 52, for example, as shown in Figure 8. This dataset consists of 2D and 3D data with different angles, sizes, and viewpoints, i.e., different appearances, along with feature regions 62 and reference label information 64 corresponding to those appearances. The dataset creation unit 13 outputs the created dataset 58 to the training unit 14. In this case, in Figure 8, the "input" data is the object data after scaling and rotation based on the object data 52. The "output" data is data that includes feature regions 62 and reference label information 64 corresponding to the "input" data. In the "output" of Figure 8, the labels "A", "B", "C", and "D" are reference label information, and the areas enclosed by thick lines and rectangles near the reference label information are feature regions.
[0032] The learning unit 14 outputs a trained model 54 by training with the dataset 58 received from the dataset creation unit 13. As shown in Figure 9, the trained model is a neural network that takes 2D and 3D data of various angles and orientations based on object data 52 as input, and outputs feature regions 62 and reference label information 64 at those angles and orientations.
[0033] Thus, the model generation device 10 of this disclosure generates a trained model 54 used for recognizing the position and orientation of an object. The model generation device 10 includes a definition unit 11. The definition unit 11 pre-defines a feature region 62 and a reference representative point 66 corresponding to the feature region 62 for the object data 52.
[0034] According to this, the model generation device 10 can generate a trained model 54 that can efficiently and accurately recognize the position and orientation of an object by defining a feature region 62 for the object data 52 and pre-defining a reference representative point 66 corresponding to that feature region 62. Furthermore, because the feature region 62 of the object and its reference representative point 66 are pre-defined in this trained model 54, it is possible to estimate the position and orientation with high accuracy and robustness even if the shape and orientation of the object are complex.
[0035] Furthermore, the definition unit 11 divides the object data 52 in three-dimensional space using voxels 63 and defines randomly selected voxels 63 as feature regions 62. This allows for the automation of the feature region definition process by randomly defining the feature region 62, thereby reducing manual labor. Additionally, by dividing the object data 52 into voxels 63 in three-dimensional space and randomly selecting these voxels 63, the automation of feature region definition can be achieved with a simple configuration.
[0036] Alternatively, the definition unit 11 may be configured to define the feature region 62 by, for example, a marker attached to an actual object. In this case, the marker refers to, for example, a writing instrument equipped with ink. The operator then marks the area to be defined in the feature region 62 on the object using, for example, a marker of a different color from the object itself. The model generation device 10 then photographs the object with the markers attached using, for example, a camera. The model generation device 10 then acquires the image data of the object as object data 52, recognizes the area with the markers attached within that object data, and defines that area as the feature region 62.
[0037] According to this, the feature region 62 becomes easier to see with the naked eye and easier to detect. Therefore, the definition of the feature region 62 can be simplified, and the burden on the worker required to define the feature region 62 can be reduced.
[0038] Next, the first recognition device 201 will be described with reference to Figures 10 to 18. The first recognition device 201 is an example of a recognition device, and is a device that recognizes the position and orientation of an object using, for example, a trained model 54 generated by the model generation device 10 described above. As shown in Figure 10, for example, the first recognition device 201 includes a feature region detection unit 21, a selection unit 22, a representative point extraction unit 23, a position and orientation estimation unit 24, and a precision estimation unit 25.
[0039] The first recognition device 201 may be a dedicated computer, but it can also be implemented by installing the recognition program 55 on a general-purpose personal computer or server. The first recognition device 201 may be the same computer as the model generation device 10, or it may be a different computer. The hardware configuration of the first recognition device 201 can be the same as that of the model generation device 10, as shown in Figure 11, and may include a second processor 1b, a second main memory 2b, a second input unit 3b, a second output unit 4b, and a second auxiliary storage device 5b. Note that the second processor 1b, second main memory 2b, second input unit 3b, second output unit 4b, and second auxiliary storage device 5b have the same or common configuration as the first processor 1a, first main memory 2a, first input unit 3a, first output unit 4a, and first auxiliary storage device 5a of the model generation device 10, so a detailed explanation of each configuration is omitted. The first recognition device 201 can also be configured to communicate with an external computer via a telecommunications line such as the Internet or a LAN.
[0040] The second auxiliary storage device 5b of the first recognition device 201 stores, for example, a trained model 54 generated by the model generation device 10, and a recognition program 55. The recognition program 55 is a program that causes a computer to execute a process to recognize the position and orientation of an object using the trained model 54. In other words, the recognition program 55 is a computer program that virtually implements the feature region detection unit 21, selection unit 22, representative point extraction unit 23, position and orientation estimation unit 24, and precision estimation unit 25 shown in Figure 10 on a computer. The first recognition device 201 can virtually implement the feature region detection unit 21, selection unit 22, representative point extraction unit 23, position and orientation estimation unit 24, and precision estimation unit 25 on a computer by having the second processor 1b read the recognition program 55 from the second auxiliary storage device 5b, expand it in the second main memory device 2b, and execute it.
[0041] In other words, the feature region detection unit 21, selection unit 22, representative point extraction unit 23, position and orientation estimation unit 24, and precision estimation unit 25 can be configured as functional units that are virtually realized by executing the recognition program 55 in the second processor 1b, for example. The first recognition device 201 can be configured with the same or common hardware for the feature region detection unit 21, selection unit 22, representative point extraction unit 23, position and orientation estimation unit 24, and precision estimation unit 25 shown in Figure 10, or it can be configured with different hardware.
[0042] The second auxiliary storage device 5b, like the first auxiliary storage device 5a of the model generation device 10, is composed of a tangible, non-temporary computer-readable medium. Examples of the second auxiliary storage device 5b include, but are not limited to, HDDs (Hard Disk Drives), SSDs (Solid State Drives), magnetic disks, magneto-optical disks, CD-ROMs (Compact Disc Read Only Memory), DVD-ROMs (Digital Versatile Disc Read Only Memory), and semiconductor memory. The second auxiliary storage device 5b may be an internal medium directly connected to the bus of the computer constituting the first recognition device 201, or it may be an external medium connected to the first recognition device 201 via a telecommunications line such as the Internet or a LAN. Furthermore, when the recognition program 55 is distributed to the first recognition device 201 via a telecommunications line, the first recognition device 201, upon receiving the distribution, expands the recognition program 55 into the second main storage device 2b and executes it, thereby realizing the feature region detection unit 21, the selection unit 22, the representative point extraction unit 23, the position and orientation estimation unit 24, and the precision estimation unit 25.
[0043] Furthermore, the implementation of the feature region detection unit 21, selection unit 22, representative point extraction unit 23, position and orientation estimation unit 24, and precision estimation unit 25 is not limited to the combination of the hardware and recognition program 55 described above. The feature region detection unit 21, selection unit 22, representative point extraction unit 23, position and orientation estimation unit 24, and precision estimation unit 25 may be implemented by hardware alone, such as an integrated circuit implementing the recognition program 55, or some functions may be implemented by dedicated hardware, and the rest by a combination of hardware and the recognition program 55.
[0044] The first recognition device 201, as shown in Figure 10 for example, takes the image data 67 of the object as input and outputs the position and orientation 80 of the object by sequentially processing it in the feature region detection unit 21, selection unit 22, representative point extraction unit 23, position and orientation estimation unit 24, and precision estimation unit 25. The image data 67 of the object is image data obtained by imaging a real object with various sensors such as a camera, depth sensor, or LiDAR (Light Detection And Ranging), and includes position information in the three-dimensional direction of the object's surface. The image data 67 can consist of, for example, two-dimensional RGB image data or video data, or depth image data or video data. In the following description, the image data 67 will be described as image data, but video data will not be excluded.
[0045] The feature region detection unit 21 performs feature region detection processing. Based on the image data 67 of the object, the feature region detection processing detects and outputs at least three feature regions 62 of the object. As shown in Figure 12, the feature region detection unit 21 can perform the process of detecting feature regions 62 by inference using a trained model 54. The feature region detection unit 21 inputs the image data 67 to the trained model 54 and obtains multiple feature regions 62 and extracted label information 68, which is label information corresponding to those feature regions 62, as output values.
[0046] The feature region detection unit 21 can set a confidence level for the detected extracted label information 68, for example, as shown in Figure 13(A). In Figure 13, the information displayed as 95%, 98%, 90%, and 70% represents the confidence level. The confidence level is an indicator of the accuracy of the detected feature region 62 and can be expressed as a multi-level rank or percentage, for example. The higher the confidence level, the higher the probability that the feature region 62 is correct. The feature region detection unit 21 outputs the feature region 62 and extracted label information 68 obtained by the feature region detection process to the selection unit 22.
[0047] The selection unit 22 selects the most reliable feature regions 62 and extracted label information 68 from the feature region detection unit 21 and outputs them to the representative point extraction unit 23. The selection unit 22 selects, for example, three or more predetermined feature regions 62 in order of the highest confidence level of the extracted label information 68 they contain. In the example in Figure 13, as shown in (A) and (B), the selection unit 22 selects three feature regions 62 in order of highest confidence level from among the four feature regions 62, in this case, feature regions 62 having extracted label information 68 of "A: 95%", "B: 98%", and "C: 90%". The selection unit 22 then outputs the three selected feature regions 62 and the corresponding extracted label information 68 to the representative point extraction unit 23 along with the imaging data 67. For example, if the feature region detection unit 21 detects five or more feature regions 62, the selection unit 22 may output four or more feature regions 62 and the corresponding extracted label information 68 to the representative point extraction unit 23 along with the imaging data 67.
[0048] The representative point extraction unit 23 performs a representative point extraction process. The representative point extraction process includes, for example, the calculation of extracted representative points 69 from the imaging data 67 and feature region 62 received from the selection unit 22. For example, if the imaging data 67 is a depth image, the imaging data 67 includes a point cloud with three-dimensional positional information. The representative point extraction unit 23 extracts the point cloud contained within the feature region 62 from the imaging data 67 and calculates the extracted representative points 69 as representative points of the feature region 62 by, for example, calculating the mean value or median of that point cloud. In other words, the extracted representative points 69 are representative points that represent the feature region 62.
[0049] Furthermore, the representative point extraction process includes, for example, the extraction of at least three extracted representative points 69 from the calculated extracted representative points 69, and extracted label information 68 for identifying each extracted representative point 69, as shown in Figure 13(C). In this embodiment, the representative point extraction unit 23 receives input from the three feature regions 62 selected by the selection unit 22, and outputs extracted representative points 69 that are representative points of those three feature regions 62.
[0050] The position and orientation estimation unit 24 performs position and orientation estimation processing. The position and orientation estimation processing includes, for example, as shown in Figure 14, estimating the position and orientation of the object 61 captured in the imaging data 67 by aligning the extracted label information 68 of the extracted representative points 69 extracted from the imaging data 67 with the reference label information 64 of the reference representative points 66 defined in advance by the model generation device 10. In this case, the extracted label information 68 shown in Figure 14(A) corresponds to each extracted representative point 69 extracted by the representative point extraction unit 23. The reference label information 64 shown in Figure 14(B) is for identifying at least three reference representative points 66 that have the correct positional relationship defined in advance by the model generation device 10.
[0051] In other words, the position and orientation estimation unit 24 estimates the position and orientation of the object using the positional relationship between the extracted representative point 69 and the positional relationship between the reference representative point 66. In this case, the position and orientation estimation unit 24 can determine, for example, pairs in which the extracted label information 68 and the reference label information 64 match, and then determine the translation and rotation for alignment based on the covariance matrix created from the coordinate values of the extracted representative point 69 and the reference representative point 66 of that pair. Then, as shown in Figure 10, the position and orientation estimation unit 24 outputs the position and orientation 80 of the object 61 obtained by the estimation process, for example, as a value in a 6-degree-of-freedom coordinate system.
[0052] Conventional object recognition and position / orientation estimation techniques have the following challenges, for example. For instance, methods that require prior information about the type of surface to detect feature regions necessitate selecting the appropriate surface for each scene if there are multiple options within the object. Furthermore, these methods only yield surface equations, making it impossible to determine the object's position within the coordinate system. Additionally, if a pointed area is facing the camera or sensor, it may not be possible to obtain sufficient feature regions.
[0053] Furthermore, in methods using end-to-end neural networks that take point clouds as input and position and orientation as output, the computational load on the neural network can be high, which can lead to a decrease in estimation accuracy. It has also been confirmed that even with rule-based estimation methods, if the features and the shape of the object are poorly matched, the estimation accuracy can deteriorate.
[0054] In contrast, the first recognition device 201 of this disclosure comprises a feature region detection unit 21, a representative point extraction unit 23, and a position and orientation estimation unit 24. The feature region detection unit 21 detects at least three feature regions 62 of the object 61 from imaging data 67 obtained by imaging the object 61. The representative point extraction unit 23 calculates extracted representative points 69 from the imaging data 67 and the feature regions 62, and extracts at least three extracted representative points 69 and extracted label information 68 for identifying each extracted representative point 69. The position and orientation estimation unit 24 then estimates the position and orientation of the object 61 by aligning the extracted representative points 69 with reference representative points 66 that match the extracted label information 68 corresponding to the extracted representative points 69 and reference label information 64 for identifying at least three reference representative points 66 that have a predefined correct positional relationship.
[0055] According to this, the first recognition device 201 estimates the position and orientation of the object 61 itself from the image data of the object 61, eliminating the need for tasks such as selecting a surface for each scene, and also enabling the determination of the position and orientation of any location within the object 61. Furthermore, if at least three feature regions 62 can be detected, the position and orientation of the entire object 61 can be estimated, thus avoiding excessive computational load on neural networks and preventing a decrease in estimation accuracy. In addition, since the surfaces of the object 61 do not need to have a constant curvature, it can be adapted to objects of any shape. For these reasons, the first recognition device 201 makes it possible to recognize the position and orientation of a wide variety of objects 61 with high accuracy and stability.
[0056] Furthermore, the feature region detection unit 21 detects the feature region 62 through inference using the trained model 54. This allows for robust and highly accurate detection of the feature region 62, thereby improving the accuracy of position and orientation recognition.
[0057] Furthermore, the imaging data 67 is a depth image. By using a depth image with three-dimensional positional information for the imaging data 67, it is less affected by light compared to RGB images, etc. As a result, feature regions can be detected more robustly even in bright or dark environments, and consequently, the accuracy of position and orientation recognition can be further improved.
[0058] Here, the first recognition device 201 can estimate the position and orientation of the object 61 with relatively high accuracy using the output from the position and orientation estimation unit 24. Furthermore, by including a precision estimation unit 25, the first recognition device 201 can estimate the position and orientation of the object 61 with even higher accuracy. In this embodiment, the position and orientation estimation unit 24 outputs the position and orientation of the object 61 obtained through the estimation process to the precision estimation unit 25. The precision estimation unit 25 then takes the estimation result from the position and orientation estimation unit 24 as input and performs position and orientation estimation with higher accuracy than the position and orientation estimation unit 24.
[0059] The precision estimation unit 25 performs a precision estimation process. This precision estimation process provides a more accurate position and orientation estimation than the position and orientation estimation unit 24. For example, the precision estimation process uses the estimation result of the position and orientation estimation unit 24 as the initial position and orientation, and based on this initial position and orientation, it includes a process to improve the accuracy of the position and orientation by repeatedly comparing the 3D point cloud of the object with the point cloud of the trained model 54 using methods such as the ICP (Iterative Closest Point) algorithm and performing alignment.
[0060] The precision estimation unit 25 can, for example, select only the feature regions 62 detected by the feature region detection unit 21, specifically the feature regions 62 selected by the selection unit 22, as the input point cloud, that is, the point cloud input to the precision estimation unit 25, specifically, the point cloud input to the ICP. In other words, as shown in Figure 15, the precision estimation unit 25 extracts the point cloud of the region corresponding to the feature region 62 from the 3D point cloud data of the captured image based on the feature region 62 selected by the selection unit 22, and uses the extracted point cloud as the input point cloud 65 input to the ICP.
[0061] Generally, in ICP algorithms, it is preferable to input the point cloud of the entire object. However, if, for example, multiple objects of the same shape exist in the imaging data, it is difficult to accurately extract only the point cloud of the object to be recognized. Even in such situations, by using an input point cloud 65 that consists only of the point cloud extracted from the detected feature region 62 to input to ICP, noise can be removed and highly accurate precise estimation can be performed.
[0062] Furthermore, the precision estimation unit 25 may determine the input point cloud 65 based on the extracted label information 68 of the feature region 62 detected by the feature region detection unit 21, as shown in Figure 16, for example. For example, as shown in Figure 16(A), if there is a feature region 62 among the multiple feature regions 62 detected by the feature region detection unit 21 that has an extremely low reliability compared to the other feature regions 62, it can be estimated that the object 61 has a feature region 62 with a low reliability that is located on the back side, which is difficult to capture with a camera or sensor, and a feature region 62 with a high reliability that is located on the front side, which is easy to capture with a camera or sensor. In this case, as shown in Figure 16(B), the precision estimation unit 25 excludes the feature region 62 with a reliability lower than a predetermined level, and extracts the point cloud of the remaining feature region 62 to be used as the input point cloud for input to the ICP. By excluding the point cloud with a low reliability, noise is reduced, and as a result, more accurate precision estimation can be performed.
[0063] Here, the feature region detection unit 21 may be configured to detect four or more feature regions 62. In this case, the model generation device 10 also generates a trained model 54 using four or more feature regions 62, as shown in Figure 17(A). For example, as shown in Figure 17(B), even if some of the multiple feature regions 62 are hidden by an object 57 in the image data 67, the first recognition device 201 can perform position and orientation estimation processing and precise estimation processing using the other feature regions 62 visible in the image data 67. As a result, the robustness of the recognition of the position and orientation of the object can be improved.
[0064] Furthermore, the position and orientation estimation unit 24 can define multiple representative point groups, each consisting of at least three extracted representative points, as shown in Figure 18, for example. It can then estimate the position and orientation for each representative point group and use the multiple estimated position and orientation information to determine the final position and orientation.
[0065] In this case, the model generation device 10 generates a trained model 54 that detects a large number of feature regions 62. The feature region detection unit 21 detects a large number of feature regions based on this trained model 54. Then, the position and orientation estimation unit 24 defines representative point groups, each having three extracted representative points 69, from the large number of extracted representative points extracted from the feature regions 62. For example, in the example of Figure 18, it defines a first representative point group 701 and a second representative point group 702.
[0066] Then, for example, the position and orientation estimation unit 24 determines the final position and orientation estimate by discretizing and voting on the multiple position and orientation estimation results estimated from multiple representative point groups, such as the multiple position and orientation estimation results estimated by the first representative point group 701 and the second representative point group 702 in the example of Figure 18. This makes the system more robust and reduces the probability of recognition failure, as position and orientation estimation can be performed by detecting other feature regions 62 even if detection of one feature region 62 fails.
[0067] Next, with reference to Figures 19 to 22, the work system 30 and the second recognition device 202 suitable for the work system will be described. As shown in Figure 19, the work system 30 comprises the second recognition device 202, an imaging device 31, a work device 32, and a control device 33. The work device 32 is a device that performs a predetermined operation on the object 61 recognized by the second recognition device 202.
[0068] The imaging device 31 is for imaging the object 61 and acquiring imaging data 67 for use by the second recognition device 202. The imaging device 31 is composed of, for example, an RGB camera or a depth camera, and captures still images or videos of the object 61 at predetermined intervals. The work system 30 may be equipped with multiple types of imaging devices 31. The imaging device 31 outputs the imaging data 67 of the object 61 to the second recognition device 202.
[0069] The working device 32 is, for example, a multi-joint robot and has, for example, a working unit 321 that performs work on an object 61. The working device 32 can, for example, grasp, move, and place any object 61 from among a plurality of objects 61. In this case, the working unit 321 can be made up of, for example, a chuck capable of grasping parts, etc. That is, the working device 32 can be made up of a picking device that, for example, grasps any object 61 from among a plurality of objects 61 with the working unit 321 and picks it up. In this case, the area of the object 61 that is grasped by the working unit 321 is called the work-to-work area. The control device 33 receives position and orientation information of the object 61 from the second recognition device 202 and controls the operation of the working device 32.
[0070] In this scenario, where multiple objects 61 of the same shape are picked up in a so-called random stacking manner, multiple objects 61 are captured in the image data 67. As a result, the feature region detection unit 21 detects similar feature regions 62 from multiple objects 61. However, in order for the position and orientation estimation unit 24 to estimate the position and orientation of the objects 61, it is necessary to group the feature regions 62 for each object 61 and to determine which object's position and orientation will be estimated.
[0071] Therefore, as shown in Figure 21, the second recognition device 202 further includes a grouping unit 26 and a target determination unit 27 in addition to the feature region detection unit 21, selection unit 22, representative point extraction unit 23, position and orientation estimation unit 24, and precision estimation unit 25. The feature region detection unit 21, selection unit 22, representative point extraction unit 23, position and orientation estimation unit 24, and precision estimation unit 25 have the same configuration as the first recognition device 201 described above, so their explanation is omitted. In this case, the recognition program 55 further implements the grouping unit 26 and the target determination unit 27 in addition to the feature region detection unit 21, selection unit 22, representative point extraction unit 23, position and orientation estimation unit 24, and precision estimation unit 25.
[0072] The grouping unit 26 performs a grouping process. The grouping process, as shown in Figure 22 for example, groups the multiple feature regions 62 detected by the feature region detection unit 21 into groups belonging to each object 61, such as the first group 81, the second group 82, and the third group 83. The grouping unit 26 may use methods such as grouping using a neural network or grouping using distance information between the feature regions 62.
[0073] For example, when a feature region 62 is defined in the definition unit 11, the distance information between feature regions 62 of the trained model 54 is already known. The grouping unit 26 uses this distance information between feature regions 62 to group the multiple feature regions 62 detected from the imaging data 67 together, based on the distance relationships between the feature regions 62 of the trained model 54. This allows for the correspondence between the detected feature regions 62 and each object 61, even if the regions of each object 61 have not been separated by instance segmentation or other means.
[0074] Furthermore, the grouping unit 26 groups the image data 67 based on information obtained by segmenting the image data 67 at the pixel level. That is, the grouping unit 26 associates a label or category, such as what is depicted, with each pixel of the image data 67, and groups the feature regions 62 based on that information. By using segmented information for each object in this way, highly reliable grouping becomes possible.
[0075] The target determination unit 27 performs a determination process. The determination process includes determining which of the multiple targets 61 will have its position and orientation recognized, using the detection results of the feature region detection unit 21. In other words, the target determination unit 27 determines which target 61 will be worked on by the work device 32. Using the detection results of the feature region detection unit 21, the target determination unit 27 prioritizes recognition of the target 61 in which, for example, more feature regions 62 have been detected. This allows the work to be performed on the target whose position and orientation can be reliably recognized, thereby increasing the success rate of the work performed by the work system 30.
[0076] Here, the definition unit 11 of the model generation device 10 may define lines or surfaces that have a specific meaning by combining two or more feature regions 62. Each feature region 62 has reference label information 64, and by giving meaning to combinations of this reference label information 64, it can be used in a work system using the first recognition device 201, which will be described later.
[0077] For example, as shown in Figure 20(A), the definition unit 11 defines the positional relationship of, for example, three reference label information 64, in this case the first reference label information 641, the second reference label information 642, and the third reference label information 643, as a plane. If the number of first reference label information 641, second reference label information 642, and third reference label information 643 does not match, it can be estimated that the feature region to which the extracted label information belongs is hidden and not visible due to other objects, etc. That is, for example, in Figure 20(B), two first reference label information 641 and two second reference label information 642 are detected, while only one third reference label information 643 is detected, so it can be estimated that the feature region to which the other third reference label information 643 belongs is hidden and not visible due to other objects, etc. This can be useful in a work system, for example, for determining the order of tasks.
[0078] Furthermore, the reference label information 64 of the feature region 62 includes information regarding whether the feature region 62 is a part of the object 61 that is gripped by the work unit 321, i.e., a part of the object 61 that is worked on by the work unit 321. For example, as shown in Figure 7, each reference label information 64 of the feature region 62 is accompanied by the information "gripping possible" if the corresponding feature region 62 is a part to be gripped, and by the information "not gripping possible" if it is not a part to be gripped. The first recognition device 201 and the second recognition device 202 can determine from the extracted label information 68 of the detected feature region 62 whether the feature region 62 is a part to be gripped. When the first recognition device 201 and the second recognition device 202 detect a feature region 62 that has information about a part to be gripped, they can determine that the part to be gripped is visible. According to this, the work system 30 can determine from the extracted label information 68 of the detected feature region 62 whether the part to be worked on is visible, that is, whether work can be performed on the part to be worked on, and this can be used as a basis for deciding whether or not to make it a picking target.
[0079] (Other embodiments) This disclosure is not limited to the embodiments described above and shown in the drawings, and can be modified, combined, or expanded at will without departing from its essence. The numerical values and other details shown in the embodiments are illustrative and not limiting.
[0080] This disclosure is described in accordance with the embodiments, but it is understood that this disclosure is not limited to such embodiments or structures. This disclosure also includes various modifications and variations within the equivalence. In addition, various combinations and forms, as well as other combinations and forms that include only one, more, or fewer of those elements, fall within the scope and concept of this disclosure.
[0081] The control unit and its method described herein may be implemented by a dedicated computer provided by configuring a general-purpose processor and memory programmed to perform one or more functions embodied by a computer program. Alternatively, the control unit and its method described herein may be implemented by a dedicated computer provided by configuring a processor with one or more dedicated hardware logic circuits. Alternatively, the control unit and its method described herein may be implemented by one or more dedicated computers configured by a combination of a processor and memory programmed to perform one or more functions and a processor configured with one or more hardware logic circuits. Furthermore, the computer program may be stored as instructions executed by the computer on a computer-readable non-transitional tangible recording medium.
[0082] This disclosure includes, in addition to, the inventions described in the claims, the following inventions: [1] A feature region detection unit (21) detects at least three feature regions (62) of an object (61) from imaging data (67) obtained by imaging the object (61), A representative point extraction unit (23) calculates extracted representative points (69) from the aforementioned imaging data and the aforementioned feature region, and extracts at least three of the extracted representative points and extracted label information (68) for identifying each of the extracted representative points. A position and orientation estimation unit (24) estimates the position and orientation of an object by aligning the extracted representative point with the extracted representative point, where the extracted label information corresponding to the extracted representative point and the reference label information (64) for identifying at least three reference representative points having a predefined correct positional relationship coincide. Recognition devices (201, 202) equipped with the following. [2] The feature region detection unit detects the feature region by inference using a trained model (54). The recognition device described in [1]. [3] The aforementioned imaging data is a depth image. The recognition device described in [1] or [2]. [4] The system further includes a precision estimation unit (25) that takes the estimation result from the position and attitude estimation unit as input and performs position and attitude estimation with higher accuracy than the position and attitude estimation unit. The recognition device described in any one of [1] to [3]. [5] The precision estimation unit uses only the feature regions detected by the feature region detection unit as input point clouds. The recognition device described in [4]. [6] The precision estimation unit determines the input point cloud based on the extracted label information of the feature region detected by the feature region detection unit. The recognition device described in [4] or [5]. [7] The feature region detection unit detects four or more feature regions. A recognition device as described in any one of [1] to [6]. [8] The position and orientation estimation unit defines multiple representative point groups, each consisting of at least three of the extracted representative points, estimates the position and orientation for each of the representative point groups, and uses the estimated position and orientation information to determine the final position and orientation. The recognition device described in [7]. [9] The aforementioned imaging data includes multiple objects, The feature region detection unit further comprises a grouping unit that groups the plurality of feature regions detected by the feature region detection unit into groups belonging to each of the objects. The recognition device described in any one of [1] to [8].
[10] The grouping unit groups the objects based on information segmented at the pixel level. The recognition device described in [9].
[11] The aforementioned imaging data includes multiple objects, The system further includes an object determination unit that determines which of the multiple objects to perform position and orientation analysis on using the detection results of the feature region detection unit. The recognition device described in any one of [1] to
[10] .
[12] This method generates a trained model used to estimate the position and orientation of an object from imaging data obtained by imaging the object, The object data representing the aforementioned object includes a definition unit that pre-defines a characteristic region of the object, a reference representative point that is a representative point within the characteristic region, and reference label information that identifies the characteristic region or the reference representative point. Model generation device.
[13] The definition unit divides the object data into voxels in three-dimensional space and defines randomly selected voxels in the feature region.
[12] The model generation device described above.
[14] The defining unit defines the feature region using markers added to the object afterwards.
[12] The model generation device described above.
[15] The defining unit defines a line or surface that has a specific meaning by two or more combinations of the feature regions. A model generation device as described in any one of
[12] to
[14] .
[16] The aforementioned reference label information includes information regarding whether the feature region is a workpiece that the work device of the work system is working on. A model generation device as described in any one of
[12] to
[14] .
[17] A recognition device described in any one of [1] to
[11] , The device comprises a work device that performs a predetermined operation on an object recognized by the recognition device, Work system. [Explanation of Symbols]
[0083] 10...Model generation device, 11...Definition unit, 201...First recognition device, recognition device, 202...Second recognition device, recognition device, 21...Feature region detection unit, 23...Representative point extraction unit, 24...Position and orientation estimation unit, 25...Precise estimation unit, 26...Grouping unit, 27...Target determination unit, 30...Work system, 32...Work device, 321...Work unit, 52...Target data, 54...Trained model, 61...Target, 62...Feature region, 63...Voxel, 64...Reference label information, 641...First reference label information, reference label information, 642...Second reference label information, reference label information, 643...Third reference label information, reference label information, 66...Reference representative point, 67...Imaging data, 68...Extracted label information, 69...Extracted representative point, 70...Input point cloud, 701...First representative point group, representative point group, 702...Second representative point group, representative point group
Claims
1. A feature region detection unit (21) detects at least three feature regions (62) of an object (61) from imaging data (67) obtained by imaging the object (61), A representative point extraction unit (23) calculates representative points (69) from the aforementioned imaging data and the feature region, and extracts at least three of the aforementioned representative points and extraction label information (68) for identifying each of the aforementioned representative points. A position and orientation estimation unit (24) estimates the position and orientation of an object by aligning the extracted representative point with the extracted representative point, where the extracted label information corresponding to the extracted representative point and the reference label information (64) for identifying at least three reference representative points having a predefined correct positional relationship match. Recognition devices (201, 202) equipped with the following.
2. The feature region detection unit detects the feature region by inference using a trained model (54). The recognition device according to claim 1.
3. The aforementioned imaging data is a depth image. The recognition device according to claim 1.
4. The system further includes a precision estimation unit (25) that takes the estimation result from the position and attitude estimation unit as input and performs position and attitude estimation with higher accuracy than the position and attitude estimation unit. The recognition device according to claim 1.
5. The precision estimation unit uses only the feature regions detected by the feature region detection unit as input point clouds. The recognition device according to claim 4.
6. The precision estimation unit determines the input point cloud based on the extracted label information of the feature region detected by the feature region detection unit. The recognition device according to claim 4.
7. The feature region detection unit detects four or more feature regions. The recognition device according to claim 1.
8. The position and orientation estimation unit defines a plurality of representative point groups (701, 702) in which at least three or more of the extracted representative points form one group, estimates the position and orientation for each of the representative point groups, and uses the multiple estimated position and orientation information to determine the final position and orientation. The recognition device according to claim 7.
9. The aforementioned imaging data includes multiple objects, The feature region detection unit further comprises a grouping unit (26) that groups the plurality of feature regions detected by the feature region detection unit into groups belonging to each of the objects. The recognition device according to claim 1.
10. The grouping unit groups the objects based on information segmented at the pixel level. The recognition device according to claim 9.
11. The aforementioned imaging data includes multiple objects, The system further includes an object determination unit (27) that determines which of the multiple objects to perform position and orientation analysis on using the detection results of the feature region detection unit. The recognition device according to claim 1.
12. This method generates a trained model (54) used to estimate the position and orientation of an object (61) from imaging data (67) obtained by imaging the object, The object data (52) representing the object includes a definition unit (11) that pre-defines a feature region (62) which is a characteristic area of the object, a reference representative point (66) which is a representative point within the feature region, and reference label information (64) which identifies the feature region or the reference representative point. Model generation device (10).
13. The definition unit divides the object data into voxels (63) in three-dimensional space and defines randomly selected voxels in the feature region. The model generation apparatus according to claim 12.
14. The defining unit defines the feature region using markers added to the object afterwards. The model generation apparatus according to claim 12.
15. The defining unit defines a line or surface that has a specific meaning by combining two or more of the feature regions. The model generation apparatus according to claim 12.
16. The aforementioned reference label information includes information regarding whether the feature region is a workpiece that the work device of the work system is working on. The model generation apparatus according to claim 12.
17. A recognition device (201, 202) according to any one of claims 1 to 11, The device comprises a work device (32) that performs a predetermined operation on an object recognized by the recognition device, Work system (30).
Citation Information
Patent Citations
Package
JP2019085168A