Object recognition apparatus and object recognition method

By generating two-dimensional images from parallel projection of three-dimensional data and matching them with a single-size template, the problem of excessive template generation load and processing time in existing technologies is solved, achieving high-speed and reliable three-dimensional object recognition.

CN113939852BActive Publication Date: 2025-11-21OMRON CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201980096937.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-06-12
Publication Date
2025-11-21
Estimated Expiration
2039-06-12

AI Technical Summary

Technical Problem

Existing template matching methods, when recognizing 3D objects, lead to increased template generation load, larger data volume, and longer processing time as resolution increases, and also require excessive memory, making it difficult to achieve high-speed recognition.

Method used

A two-dimensional image is generated by parallel projection of three-dimensional data, and a template of a single size is used for matching. Combining the three-dimensional data acquisition unit and the parallel projection conversion unit, the generated two-dimensional image is used for template matching to identify the pose and position of the object.

Benefits of technology

It enables high-speed object detection at various depth distances, reduces the number of templates and data volume, lowers memory requirements, and improves processing speed and recognition reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113939852B_ABST
    Figure CN113939852B_ABST
Patent Text Reader

Abstract

An object recognition device includes a three-dimensional data acquisition section that acquires three-dimensional data composed of a plurality of points each having three-dimensional information, a parallel projection conversion section that generates a two-dimensional image by parallel projecting each point of the three-dimensional data onto a projection plane, and a recognition processing section that detects a target object from the two-dimensional image by template matching.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a technique of recognizing a three-dimensional object by template matching. BACKGROUND

[0002] As one of methods of recognizing (detecting) an object from an image, there is template matching. Template matching is a method of preparing a model (template) of an object to be recognized in advance, and detecting an object included in an input image by evaluating the degree of coincidence of image features between the input image and the model. Object recognition based on template matching is used in a plurality of fields such as inspection or sorting in FA (Factory Automation), robot vision, monitoring cameras, and the like.

[0003] In recent years, a technique of applying template matching to recognition of a three-dimensional position and attitude of an object has been attracting attention. The basic principle is to prepare a plurality of templates having different views (appearances) by changing the viewpoint position with respect to a target object, and determine the three-dimensional position and attitude of the target object with respect to a camera by selecting a template having the most matched view with respect to the target object in an input image from among the templates. However, this method has problems that the load of template generation increases, the data amount of the templates increases, and the processing time of template matching increases if the resolution of recognition is to be improved because the resolution of recognition is proportional to the variation of the templates.

[0004] As a countermeasure against such problems, in Patent Literature 1, there is disclosed an idea of measuring the longitudinal distance of a target object by a depth sensor, and scaling (enlarging / reducing) a template (a two-dimensional grid in which feature values are sampled) according to the longitudinal distance.

[0005] PRIOR ART DOCUMENTS

[0006] PATENT LITERATURE

[0007] Patent Literature 1: U.S. Patent No. 9659217 Specification SUMMARY

[0008] PROBLEMS TO BE SOLVED BY THE INVENTION

[0009] According to the method of Patent Document 1, since the templates of a plurality of views differing only in the depth distance can be generalized, it is expected that the load of template generation is reduced and the number of templates is reduced, and the like. However, in the search of template matching, since the process of enlarging or reducing the template in accordance with the depth distance of each pixel occurs, there is a disadvantage that the processing speed becomes slow. In order to reduce the time required for the enlargement or reduction of the template, it is technically possible to generate templates of a plurality of scales in accordance with the distance range in which the object can exist and the resolution necessary before the template matching process and store them in the work memory, but a very large memory capacity is required, and thus it is not practical.

[0010] The present application has been achieved in view of the above-described actual situation, and an object thereof is to provide a practical technology capable of detecting an object that can exist at various depth distances at high speed by template matching.

[0011] Means for solving the technical problem

[0012] One aspect of the present application provides an object recognition device characterized by comprising: a three-dimensional data acquisition section that acquires three-dimensional data composed of a plurality of points each having three-dimensional information; a parallel projection conversion section that generates a two-dimensional image by parallel projecting each point of the three-dimensional data onto a projection plane; and a recognition processing section that detects an object from the two-dimensional image by template matching.

[0013] The three-dimensional data can be data obtained by three-dimensional measurement. The method of three-dimensional measurement can be any method, and can be an active measurement method or a passive measurement method. Template matching is a method of judging whether or not a partial image in a region of interest in the two-dimensional image is an image of the object by evaluating the degree of coincidence (similarity) of the image features between the template (model) of the object and the region of interest. If a plurality of templates of the object differing in the view (appearance) are used for template matching, the posture of the object can also be recognized.

[0014] In the present application, the two-dimensional image generated by parallel projecting the three-dimensional data is used for template matching. In parallel projection, the object is projected at the same size regardless of the distance from the projection plane to the object. Therefore, in the two-dimensional image generated by parallel projection, the image of the object (regardless of the depth distance thereof) is always taken at the same size. Therefore, matching can be performed using only a single size of template, and high-speed processing can be performed compared to the conventional method (a method of scaling the template in accordance with the depth distance). In addition, the number of templates and the amount of data can be reduced, and the amount of work memory required is also small, and thus the practicality is excellent.

[0015] It can also be that the recognition processing section uses a template generated from an image obtained by parallel projection of the object to be recognized as the template of the object to be recognized. By generating a template from a parallel projection image, the matching accuracy of the template with the object image in a two-dimensional image is improved, and thus the reliability of the object recognition processing can be improved.

[0016] The projection surface can be arbitrarily set, but it is preferable to set the projection surface in such a manner that the projection points of the points constituting the three-dimensional data are distributed over as wide a range as possible on the projection surface. For example, it can also be that, in the case where the three-dimensional data is data generated using an image captured by a camera, the parallel projection conversion section sets the projection surface in a manner orthogonal to the optical axis of the camera.

[0017] It can also be that, in the case where a first point in the three-dimensional data is projected onto a first pixel in the two-dimensional image, the parallel projection conversion section associates depth information obtained from the three-dimensional information of the first point with the first pixel. In the case where the points of the three-dimensional data have luminance information, in the case where a first point in the three-dimensional data is projected onto a first pixel in the two-dimensional image, the parallel projection conversion section associates the luminance information of the first point with the first pixel. In the case where the points of the three-dimensional data have color information, in the case where a first point in the three-dimensional data is projected onto a first pixel in the two-dimensional image, the parallel projection conversion section associates the color information of the first point with the first pixel.

[0018] It can also be that, in the case where there is no point projected onto a second pixel in the two-dimensional image, the parallel projection conversion section generates information associated with the second pixel on the basis of information associated with the pixels around the second pixel. For example, it can also be that the parallel projection conversion section obtains information associated with the second pixel by interpolating the information associated with the pixels around the second pixel. By such processing, the amount of information of the two-dimensional image is increased, and it is expected that the accuracy of template matching is improved.

[0019] It can also be that the three-dimensional data is data generated using an image captured by a camera, and, in the case where a plurality of points in the three-dimensional data are projected onto the same position on the projection surface, the parallel projection conversion section generates the two-dimensional image using the point closest to the camera among the plurality of points. By such processing, a parallel projection image that takes into account the overlap (occlusion) of objects with each other when viewed from the projection surface side (i.e., only the points visible from the camera are mapped to the two-dimensional image) is generated, and thus the object recognition processing based on template matching can be performed with high accuracy.

[0020] The present application can be understood as an object recognition device having at least a part of the above-described mechanism or structure, or as an image processing device that performs the above-described parallel projection conversion. In addition, the present application can be understood as an object recognition method, an image processing method, a template matching method, a control method of an object recognition device, or the like that includes at least a part of the above-described processing, or as a program for realizing such a method or a recording medium that non-temporally records the program. Note that each of the above-described mechanisms and processing can be combined with each other as much as possible to constitute the present application.

[0021] According to the present application, a practical technique can be provided that can detect an object that can exist at various longitudinal distances at high speed through template matching. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 is a diagram that schematically shows processing by an object recognition device.

[0023] Figure 2 is a diagram that schematically shows the overall structure of an object recognition device.

[0024] Figure 3 is a block diagram showing the structure of an image processing device.

[0025] Figure 4 is a flowchart of a template generation processing.

[0026] Figure 5 is a diagram showing a setting example of a viewpoint position.

[0027] Figure 6 is a diagram showing an example of a parallel projection image in the template generation processing.

[0028] Figure 7 is a flowchart of an object recognition processing.

[0029] Figure 8 is a flowchart of a parallel projection conversion in the object recognition processing.

[0030] Figure 9 is a diagram showing a setting example of a camera coordinate system and a projection image coordinate system.

[0031] Figure 10 is a flowchart of a projection point supplement processing. DETAILED DESCRIPTION

[0032] <APPLICATION EXAMPLE>

[0033] Figure 1 schematically shows processing by an object recognition device as one of application examples of the present application. Figure 1The symbol 10 shows a case where three objects 102a, 102b, 102c on the stage 101 are measured (photographed) from an oblique upper side by the camera 103. The objects 102a, 102b, 102c are objects of the same shape (cylindrical) and the same size, but the longitudinal distance from the camera 103 is farther in the order of the object 102a, the object 102b, and the object 102c.

[0034] The symbol 11 is an example of three-dimensional data generated on the basis of an image photographed by the camera 103. The three-dimensional data 11 is data composed of a plurality of points each having three-dimensional information. The form of the three-dimensional data 11 can be any form, for example, can be data of a form in which each point has a three-dimensional coordinate value, or can be data of a form in which a depth value (information of a longitudinal distance) is associated with each point (each pixel) of a two-dimensional image. The three-dimensional coordinate value can be a coordinate value of a camera coordinate system, or a coordinate value of a global coordinate system, or a coordinate value of a coordinate system other than these. Figure 1 The three-dimensional data 11 is an example of a depth image, and for convenience, the depth values are expressed by shades (the darker a point is, the farther it is from the camera 103). In a general optical system, the farther an object is from the camera 103, the smaller the object is imaged, and thus the size on the image becomes smaller in the order of the object 102a, the object 102b, and the object 102c.

[0035] In the conventional template matching, in order to cope with objects of various sizes, a plurality of templates of different sizes are used, or the size of the template is scaled according to the depth value as in Patent Literature 1. However, these conventional methods have the drawbacks that the processing speed is reduced and the memory capacity is increased, as described above.

[0036] Therefore, in the embodiment of the present application, the three-dimensional data 11 is subjected to parallel projection conversion, and a two-dimensional image 12 is generated, and this two-dimensional image 12 is used for template matching. By performing the parallel projection conversion, objects of the same actual size have the same size on the two-dimensional image 12. Therefore, all the objects 102a, 102b, 102c contained in the two-dimensional image 12 can be detected by applying only a single size of template 13. The symbol 14 shows an example of a recognition result.

[0037] According to the method of the present embodiment, a high-speed processing can be performed compared with the conventional method. In addition, it has the advantages that the number of templates and the amount of data can be reduced, and the amount of working memory required is small, and thus the practicality is excellent. Note that, for convenience of explanation, in the Figure 1 The posture of the objects 102a, 102b, 102c is the same in the example shown in the symbol 10, but in the case where the shape changes depending on the posture of the object (i.e., the angle at which the object is observed), it is only necessary to prepare the template 13 in the posture that is intended to be recognized.

[0038] <Embodiments>

[0039] (Overall configuration of object recognition device)

[0040] Reference Figure 2 An object recognition device according to an embodiment of the present application will be described.

[0041] The object recognition device 2 is a system provided on a production line where assembly or processing of articles and the like is performed, and recognizes (three-dimensional object recognition) the position and posture of an object 27 loaded on a pallet 26 by template matching using data taken in from a sensor unit 20. The object (hereinafter also referred to as "target object") 27 to be recognized is bulked on the pallet 26.

[0042] The object recognition device 2 is roughly composed of the sensor unit 20 and an image processing device 21. The sensor unit 20 and the image processing device 21 are connected by wire or wirelessly, and the output of the sensor unit 20 is taken in to the image processing device 21. The image processing device 21 is a device that performs various processes using data taken in from the sensor unit 20. As the process of the image processing device 21, for example, distance measurement (ranging), three-dimensional shape recognition, object recognition, scene recognition, and the like can be included. The recognition result of the object recognition device 2 is output, for example, to a PLC (Programmable Logic Controller) 25 or a display 22 or the like. The recognition result is used, for example, for control of a picking robot 28, control of a processing device or a printing device, inspection or measurement of the target object 27, and the like.

[0043] (Sensor unit)

[0044] The sensor unit 20 has at least a camera for taking an optical image of the target object 27. In addition, the sensor unit 20 can also include structures (sensors, lighting devices, light projecting devices, and the like) necessary for three-dimensional measurement of the target object 27. For example, in the case of measuring a longitudinal distance by stereo matching (also referred to as stereo vision, stereo camera method, and the like), a plurality of cameras are provided in the sensor unit 20. In the case of active stereo, a light projecting device that projects pattern light to the target object 27 is also provided on the sensor unit 20. In the case of three-dimensional measurement by a spatially encoded pattern projection method, a light projecting device that projects pattern light and a camera are provided on the sensor unit 20. In addition to these, any method can be used as long as it is a method capable of acquiring three-dimensional information of the target object 27, such as an illumination difference stereo method, a TOF (Time of Flight) method, a phase shift method, and the like.

[0045] (Image processing device)

[0046] The image processing apparatus 21 is constituted by, for example, a computer provided with a CPU (processor), a RAM (memory), a nonvolatile storage device (hard disk, SSD, or the like), an input device, an output device, and the like. In this case, the CPU expands a program stored in the nonvolatile storage device in the RAM and executes the program, whereby the various structures described later are realized. However, the structure of the image processing apparatus 21 is not limited to this, and all or a part of the structures described later can be realized by a dedicated circuit such as an FPGA or an ASIC, or can be realized by cloud computing or distributed computing.

[0047] Figure 3 is a block diagram showing the structure of the image processing apparatus 21. The image processing apparatus 21 has the structure of the template generation apparatus 30 and the structure of the object recognition processing apparatus 31. The template generation apparatus 30 is a structure for generating a template used in object recognition processing, and has a three-dimensional CAD data acquisition section 300, a parallel projection parameter setting section 301, a viewpoint position setting section 302, a two-dimensional projection image generation section 303, a feature extraction section 304, and a template generation section 305. The object recognition processing apparatus 31 is a structure for performing object recognition processing by template matching, and has a three-dimensional data acquisition section 310, a parallel projection parameter setting section 311, a parallel projection conversion section 312, a feature extraction section 313, a template storage section 314, a template matching section 315, and a recognition result output section 316. In the present embodiment, the feature extraction section 313, the template storage section 314, and the template matching section 315 constitute the "recognition processing section" of the present application.

[0048] (template generation processing)

[0049] Referring to Figure 4 , an example of template generation processing performed by the template generation apparatus 30 will be described.

[0050] In step S400, the three-dimensional CAD data acquisition section 300 acquires three-dimensional CAD data of the target object 27. The CAD data can be read from an internal storage device of the image processing apparatus 21, or can be acquired from an external CAD system or a storage, or the like via a network. Note that, instead of the CAD data, three-dimensional shape data measured by a three-dimensional sensor or the like can be acquired.

[0051] In step S401, the viewpoint position setting section 302 sets a viewpoint position at which a template is to be generated. Figure 5Setting example of the viewpoint position. In this example, the viewpoints (indicated by black dots) are set at 42 vertices of the octahedron that contains the object 27. Note that the number and arrangement of the viewpoints can be appropriately set depending on the resolution required, the shape of the object 27, and the posture that can be assumed, and the like. The number and arrangement of the viewpoints can be specified by the user or can be automatically set by the viewpoint position setting section 302.

[0052] In step S402, the parallel projection parameter setting section 301 sets the parallel projection parameters used in the template generation. Here, two parameters res x , res y are used as the parallel projection parameters. (res x , res y ) is the size (in mm) of one pixel of the projection image. Note that the parallel projection parameters are also used in the parallel projection conversion in the object recognition processing described later, and the same value of the parameters can be used at the time of the template generation and at the time of the object recognition processing. This is because, by making the values of the parallel projection parameters consistent, the size of the object 27 in the template and the size of the object 27 in the projection image generated by the object recognition processing are consistent, and thus there is no need to adjust the scale of the template or the image at the time of the template matching.

[0053] In step S403, the two-dimensional projection image generation section 303 generates a two-dimensional projection image obtained by projecting the three-dimensional CAD data in parallel. Figure 6 Example of a two-dimensional projection image. By projecting each point on the surface of the object 27 in parallel onto the projection plane 62 through the viewpoint VP, a two-dimensional projection image 60 corresponding to the viewpoint VP is generated.

[0054] In step S404, the feature extraction section 304 extracts the image features of the object 27 from the two-dimensional projection image 60 generated in step S403. As the image features, for example, the luminance, the color, the luminance gradient direction, the quantized gradient direction, the HoG (Histogram of Oriented Gradients), the normal direction of the surface, the HAAR-like, the SIFT (Scale-invariant feature transform), and the like can be used. The luminance gradient direction is a direction in which the luminance gradient direction (angle) in a local area centered on a feature point is expressed in a continuous value, and the quantized gradient direction is a direction in which the luminance gradient direction in a local area centered on a feature point is expressed in a discrete value (for example, 8 directions are held in 1 byte of information from 0 to 7). The feature extraction section 304 can obtain the image features of all the points (pixels) of the two-dimensional projection image 60, or can obtain the image features of a part of the points sampled according to a prescribed rule. The points from which the image features are obtained are referred to as feature points.

[0055] In step S405, the template generation section 305 generates a template corresponding to the viewpoint VP based on the image feature extracted in step S404. The template is, for example, a data set containing coordinate values of each feature point and the extracted image feature.

[0056] The processing of steps S403 to S405 is performed for all the viewpoints set in step S401 (step S406). When the generation of the template is completed for all the viewpoints, the template generation section 305 stores the data of the template in the template storage section 314 of the object recognition processing apparatus 31 (step S407). Thus far, the template generation processing ends.

[0057] (Object recognition processing)

[0058] Reference Figure 7 An example of the object recognition processing performed by the object recognition processing apparatus 31 will be described with reference to the flowchart of Fig. 8.

[0059] In step S700, the three-dimensional data acquisition section 310 generates three-dimensional data within the field of view based on the image captured by the sensor unit 20. In the present embodiment, three-dimensional information of each point within the field of view is obtained by an active stereo method of capturing a stereo image by two cameras in a state where pattern light is projected from a light projecting device, and calculating a depth distance based on the parallax between the images.

[0060] In step S701, the parallel projection parameter setting section 311 sets a parallel projection parameter used for parallel projection conversion. Here, four parameters of res x , res y , c x , and c y are used as the parallel projection parameter. (res x , res y ) is the size of one pixel of the projection image (unit: mm), and can be set to an arbitrary value. For example, using the focal length (f x , f y ) of the camera of the sensor unit 20, the following expression can be obtained:

[0061] res x = d / f x

[0062] res y = d / f y

[0063] where d is a constant set in accordance with the longitudinal distance at which the object 27 can exist. For example, the average, minimum, or maximum value of the longitudinal distance from the sensor unit 20 to the object 27 can be set as the constant d. Note that, as described above, with respect to (res x , res y ), the same values as when the template is generated are preferably used.(c x , c y ) are the center coordinates of the projection image.

[0064] In step S702, the parallel projection conversion section 312 generates a two-dimensional projection image by parallel projecting each point (hereinafter referred to as "three-dimensional point") in the three-dimensional data onto a prescribed projection plane.

[0065] The details of the parallel projection conversion will be described with reference to Figure 8 and Figure 9 . In step S800, the parallel projection conversion section 312 calculates the image coordinate value when a three-dimensional point is parallel projected. Let the camera coordinate system be (X, Y, Z), and let the image coordinate system of the projection image be (x, y). In the example shown in Figure 9 , the camera coordinate system is set so that the origin O coincides with the center (principal point) of the lens of the camera of the sensor unit 20, the Z axis overlaps the optical axis, and the X and Y axes are parallel to the horizontal and vertical directions of the imaging element of the camera, respectively. In addition, the image coordinate system is set so that the image center (c x , c y ) is on the Z axis of the camera coordinate system, and the x and y axes are parallel to the X and Y axes of the camera coordinate system, respectively. The xy plane of the image coordinate system is the projection plane. That is, in the present embodiment, the projection plane of the parallel projection conversion is set so as to be orthogonal to the optical axis of the camera. In the case where the coordinate systems are set as shown in Figure 9 , the image coordinate value (x i , y i ) corresponding to a three-dimensional point (X i , Y i , Z i ) after parallel projection conversion can be found by the following expressions:

[0066] x i = ROUND(X i / res x +c x )

[0067] y i = ROUND(Y i / res y +c y )

[0068] where ROUND is a rounding operator that rounds off after the decimal point.

[0069] In step S801, the parallel projection conversion section 312 checks whether or not a three-dimensional point projected onto the image coordinate values (x i , y i ) already exists. Specifically, it is checked whether or not information of a three-dimensional point has already been associated with the pixel (x i , y i ) of the projection image. In the case where there is no associated three-dimensional point ("NO" in step S801), the parallel projection conversion section 312 associates the information of the three-dimensional point (X i , Y i , Z i ) with the pixel (x i , y i ) (step S803). In the present embodiment, the coordinate values (X i , Y i , Z i ) of the three-dimensional point are associated with the pixel (x i , y i ), however, this is not limiting, and it can be associated with depth information (for example, the value of Z i ), color information (for example, the RGB value), luminance information, and the like of the three-dimensional point. In the case where there is an associated three-dimensional point ("YES" in step S801), the parallel projection conversion section 312 compares the value of Z i with the value of Z already associated, and if the value of Z i is smaller ("YES" in step S802), it overwrites the information associated with the pixel (x i , y i ) with the information of the three-dimensional point (X i , Y i , Z i ) (step S803). Through such processing, in the case where a plurality of three-dimensional points are projected onto the same position on the projection surface, the information of the three-dimensional point closest to the camera among the plurality of three-dimensional points is used for the generation of the projection image. If the processing of steps S800 to S803 is performed for all three-dimensional points, it proceeds to step S703 of FIG. 7 (step S804). Figure 7

[0070] ​In step S703, the feature extraction section 313 extracts image features from the projection image. The image features extracted here are the same as those used in the template generation. In step S704, the template matching section 315 reads in the template from the template storage section 314, and detects the object object from the projection image by template matching processing using the template. At this time, by using templates of different viewpoints, it is also possible to recognize the posture of the object object. In step S705, the recognition result output section 316 outputs the recognition result. Thus far, the object recognition processing ends.

[0071] (Advantages of the Present Embodiment)

[0072] In the above-described structure and processing, a two-dimensional image generated by parallel projection of three-dimensional data is used for template matching. In parallel projection, the object object is projected at the same size regardless of the distance from the projection plane to the object object. Therefore, in the two-dimensional image generated by parallel projection, the image of the object object (regardless of the distance in depth) is always taken at the same size. Therefore, matching can be performed using only a single size of template, and high-speed processing can be performed compared to conventional methods. In addition, it is also possible to reduce the number of templates and the amount of data, and the amount of required work memory is also small, so the practicality is excellent.

[0073] In addition, in the present embodiment, the template is also generated from a parallel projection image, so the matching accuracy of the template with the image of the object object in the image converted by parallel projection is improved. Thus, it is possible to improve the reliability of the object recognition processing.

[0074] In addition, in the present embodiment, the projection plane is set orthogonal to the optical axis of the camera, so it is possible to simplify the calculation of the conversion from the camera coordinate system to the image coordinate system, to achieve high-speed parallel projection conversion processing, and further to achieve high-speed object recognition processing based on template matching. In addition, by setting the projection plane orthogonal to the optical axis of the camera, it is also possible to suppress the deformation of the image of the object object after parallel projection conversion.

[0075] In addition, in the case where a plurality of three-dimensional points are projected onto the same pixel, only the information of the three-dimensional point closest to the camera is used, so a parallel projection image that takes into account the overlap (occlusion) of objects with each other when viewed from the camera is generated, and it is possible to perform high-precision object recognition processing based on template matching.

[0076] <Other>

[0077] The above-described embodiments are merely illustrative of the structure of the present application. The present application is not limited to the above-described specific modes, and various modifications can be made within the scope of the technical idea thereof.

[0078] For example, it is also possible to perform the parallel projection conversion processing (Figure 7 After step S702) Figure 10 The projection point supplementation process is shown. Specifically, the parallel projection conversion unit 312 checks whether the information of the three-dimensional points matches the pixel (x, y) of the projection image generated in step S702. i y i (Step S100) In cases where the information of a 3D point is not associated with a pixel (i.e., there is no projection point), based on the information of the pixel (x) i y i Information associated with the pixels surrounding the pixel (e.g., 4 neighboring pixels or 8 neighboring pixels, etc.) is used to generate information for the pixel (x). i y i Information for the pixel (x) (step S101). For example, interpolation for the pixel (x) can also be generated using nearest neighbor methods, bilinear methods, bicubic methods, etc. i y i The information generated in step S101 is then compared with the information of the pixel (x). The parallel projection conversion unit 312 then combines the information generated in step S101 with the information of the pixel (x). i y i (Step S102) The processing of steps S100 to S102 is performed on all pixels of the projected image. Through this processing, the amount of information in the projected image (the number of projection points) increases, thus improving the accuracy of template matching.

[0079] Furthermore, the setting of the projection surface is not limited to... Figure 9 For example, the projection plane can be positioned behind the origin O of the camera coordinate system (image side). Alternatively, the projection plane can be positioned so that it intersects the optical axis (Z-axis) at an angle (i.e., the projection direction is not parallel to the optical axis).

[0080] <Postscript>

[0081] (1) An object recognition device (2), characterized in that it has:

[0082] The three-dimensional data acquisition unit (310) acquires three-dimensional data composed of multiple points, each of which has three-dimensional information;

[0083] The parallel projection conversion unit (312) generates a two-dimensional image by projecting each point of the three-dimensional data onto a projection plane in parallel; and

[0084] The recognition processing unit (313, 314, 315) detects object from the two-dimensional image by template matching.

[0085] Symbol Explanation

[0086] 2 … object recognition device, 20 … sensor unit, 21 … image processing device, 22 … display, 27 … object, 30 … template generation device, 31 … object recognition processing device.

Claims

1. An object recognition device, characterized in that, have: The three-dimensional data acquisition unit acquires three-dimensional data composed of multiple points, each with three-dimensional information. The parallel projection conversion unit generates a two-dimensional image by projecting each point of the three-dimensional data onto a projection plane in parallel. as well as The recognition processing unit detects objects from the two-dimensional image through template matching. In the case where the three-dimensional data includes multiple identical objects, the parallel projection conversion unit generates the two-dimensional image in such a way that the images of the multiple objects in the two-dimensional image are of the same size regardless of their depth distance, and the recognition processing unit performs the template matching using a template of a single size.

2. The object recognition device according to claim 1, characterized in that, The recognition processing unit uses a template generated from an image obtained by parallel projection of the object as the template for the object.

3. The object recognition device according to claim 1, characterized in that, The three-dimensional data is generated using images captured by a camera. The parallel projection conversion unit sets the projection surface in a manner orthogonal to the optical axis of the camera.

4. The object recognition device according to claim 1, characterized in that, When a first point in the three-dimensional data is projected onto a first pixel in the two-dimensional image, the parallel projection conversion unit associates the depth information obtained from the three-dimensional information of the first point with the first pixel.

5. The object recognition device according to claim 1, characterized in that, Each point in the three-dimensional data contains brightness information. When a first point in the three-dimensional data is projected onto a first pixel in the two-dimensional image, the parallel projection conversion unit associates the brightness information of the first point with the first pixel.

6. The object recognition device according to claim 1, characterized in that, Each point in the three-dimensional data has color information. When a first point in the three-dimensional data is projected onto a first pixel in the two-dimensional image, the parallel projection conversion unit associates the color information of the first point with the first pixel.

7. The object recognition device according to any one of claims 4 to 6, characterized in that, In the absence of a point projected onto the second pixel in the two-dimensional image, the parallel projection conversion unit generates information associated with the second pixel based on information associated with the pixels surrounding the second pixel.

8. The object recognition device according to claim 7, characterized in that, The parallel projection conversion unit obtains information associated with the second pixel by interpolating information associated with pixels surrounding the second pixel.

9. The object recognition device according to any one of claims 1 to 6, characterized in that, The three-dimensional data is generated using images captured by a camera. When multiple points in the three-dimensional data are projected onto the same position on the projection plane, the parallel projection conversion unit uses the point among the multiple points that is closest to the camera to generate the two-dimensional image.

10. An object recognition method, characterized in that, have: The steps to obtain three-dimensional data consisting of multiple points, each with its own three-dimensional information; The generation step involves generating a two-dimensional image by projecting each point of the three-dimensional data onto a projection plane in parallel. as well as The detection step involves detecting objects from the two-dimensional image using template matching. In the case where the three-dimensional data includes multiple identical objects, in the generation step, the two-dimensional image is generated in such a way that the images of the multiple objects in the two-dimensional image are of the same size regardless of their depth distance, and in the detection step, template matching is performed using a template of a single size.

11. A program product, characterized in that, Including programs, The program is used to cause a computer to perform the steps of the object recognition method of claim 10.

Citation Information

Patent Citations

  • Systems and methods for scale invariant 3D object detection leveraging processor architecture

    US9659217B2