External environment recognition device and external environment recognition method

The external environment recognition device enhances vehicle systems by aligning and matching features across multiple camera images to accurately estimate distances to unspecified targets, addressing mis-matching issues in existing technologies.

JP7897728B2Active Publication Date: 2026-07-30ASTEMO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
ASTEMO LTD
Filing Date
2022-06-27
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing technologies for external environment recognition in vehicles, such as stereo cameras, struggle with accurately measuring distances to unspecified objects and objects not directly facing the vehicle, leading to potential mis-control due to incorrect image matching.

Method used

An external environment recognition device that includes target detection units for multiple cameras, an imaging plane estimation unit, a geometric transformation estimation unit, and a feature matching unit to align and match features across images, enabling accurate estimation of distances to unspecified targets using image transformation parameters and plane equations.

Benefits of technology

Enables accurate estimation of distances to unspecified targets even when they appear differently in multiple images, improving the reliability of vehicle control systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007897728000006
    Figure 0007897728000006
  • Figure 0007897728000007
    Figure 0007897728000007
  • Figure 0007897728000008
    Figure 0007897728000008
Patent Text Reader

Abstract

To provide an external world recognition device capable of estimating the distance to an unspecified target even when the target appears in two images with different appearance.SOLUTION: An external world recognition device 100 includes: an imaging plane estimation unit 111 that estimates an imaging plane in an image 40a based on a target object detection result by a target object detection unit 101; an image transformation estimation unit 112 that estimates image transformation parameters matching the imaging plane in an image 41a based on information on the imaging plane estimated by the imaging plane estimation unit 111 and camera parameters 103 of multiple imaging units; a feature matching unit 105 that matches the features of the target object detected from the image 40a transformed using the image transformation parameters by the image transformation estimation unit 112 and the features of the target object detected from the image 41a; a position estimation unit 106 that estimates the three-dimensional position of the target object based on the results of the matching process performed by the feature matching unit 105.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an external recognition device and an external recognition method.

Background Art

[0002] In realizing automatic driving and advanced driving assistance systems, the importance of cameras for monitoring the vehicle's external environment and detecting objects necessary for vehicle driving, such as obstacles on the driving route and lane information of the driving route, is increasing. In particular, in order to improve the object detection performance, a function has been realized in which a plurality of cameras are mounted on a vehicle, information around the vehicle is acquired from the cameras, and the surrounding situation is recognized.

[0003] As a type of camera for recognizing such an external environment, for example, there is a stereo camera using a plurality of cameras. A stereo camera is one in which two cameras are arranged on a vehicle at a predetermined interval. Then, by using the parallax of the overlapping area of the images taken by the two cameras, the distance to the photographed object can be measured.

[0004] Stereo cameras include a parallel stereo camera in which the optical axes of two cameras are installed in parallel, a non-parallel stereo camera in which the optical axes are installed non-parallel, and the like. In addition, a camera system using two or more cameras may be referred to as a multi-view stereo system.

[0005] An electronic control unit (hereinafter referred to as ECU (Electronic Control Unit)) mounted on a vehicle can grasp the possibility of contact with an object in front of the vehicle by measuring the distance to the object using a plurality of images taken by a plurality of cameras. However, even if the same object is imaged by two cameras at different positions, generally the way the object appears is different between the two cameras. When collating images with different appearances, there is a possibility of collating the images at an incorrect position if there is a pattern similar to the pattern of different appearances between one image and the other image.

[0006] If images are matched at the wrong location, it can lead to errors in measuring the distance to the object based on the matching result, potentially resulting in mis-control of the vehicle. If the shape and plane model of an object in the real world are known, the way it appears in each image can be determined geometrically, and it is thought that mis-matching can be suppressed by predicting the shape and plane and matching accordingly.

[0007] Patent Document 1 describes "performing a viewpoint transformation to convert the first and second images into images from a common viewpoint by deforming at least one of the first image captured by the first camera and the second image captured by the second camera, extracting a plurality of corresponding points, and performing geometric calibration of the first and second cameras using the coordinates of the plurality of corresponding points in the first and second images before the viewpoint transformation."

[0008] Furthermore, Patent Document 2 states that, "Since the stereo camera image information captured by the imaging means is an image of the monitored surface such as the road surface or floor surface, compared to the stereo camera image information from a conventional stereo camera installed in a nearly horizontal direction, by making the image information more oriented toward the monitored surface, and furthermore, by generating 3D distance image information based on the parallelized image information obtained by parallelizing this downward-facing stereo camera image information, it becomes an overhead view image looking down on the monitored surface, so the distance to the road surface or floor surface can be determined with high accuracy." [Prior art documents] [Patent Documents]

[0009] [Patent Document 1] Japanese Patent Publication No. 2020-12735 [Patent Document 2] Japanese Patent Publication No. 2019-16308 [Overview of the project] [Problems that the invention aims to solve]

[0010] While it is appropriate to vary the method for calculating the distance from the vehicle to each object in the image, the technologies disclosed in Patent Documents 1 and 2 assume that a specific object is present in a specific location within the image. Therefore, if an unspecified object is present in a specific location within the image, the measurement of the distance to that object may be incorrect. Furthermore, while the technologies disclosed in Patent Documents 1 and 2 can calculate the distance to an object directly facing the vehicle, they may miscalculate the distance to objects not directly facing the vehicle (e.g., side walls, guardrails).

[0011] This invention was made in view of such circumstances, and aims to enable the estimation of the distance to an unspecified target even when it is captured in two images that appear differently. [Means for solving the problem]

[0012] The external environment recognition device according to the present invention includes a first target detection unit that detects a target based on a first image acquired from at least a first imaging unit among a plurality of imaging units whose imaging fields of view to the external environment overlap by at least a portion of each unit, a second target detection unit that detects a target based on a second image acquired from a second imaging unit among a plurality of imaging units, and based on the target detection result by the first target detection unit, Among the planes that constitute each target, and which are imaging planes expressed by the plane equation, The imaging plane in the first image In the image region in which the target was captured Based on the imaging plane estimation unit, the information on the imaging plane estimated by the imaging plane estimation unit, and the imaging parameters of multiple imaging units, Among the imaging planes expressed by the plane equation, The system includes an image transformation estimation unit that estimates image transformation parameters aligned with the imaging plane in the second image; a feature matching unit that compares the features of a target detected from the first image transformed using the image transformation parameters from the image transformation estimation unit with the features of a target detected from the second image; and a position estimation unit that estimates the three-dimensional position of a target based on the results of the matching process by the feature matching unit. [Effects of the Invention]

[0013] According to the present invention, even when an unspecified target is reflected in two images with different appearances, the distance to this target can be estimated. Problems, configurations, and effects other than those described above will be clarified by the description of the following embodiments.

Brief Description of the Drawings

[0014] [Figure 1] It is a diagram showing the position of a camera mounted on a vehicle according to a first embodiment of the present invention, a world coordinate system based on the vehicle, and an image coordinate system. [Figure 2] It is a diagram showing an example of an image of the foreground of a vehicle captured by a front camera and a left camera according to a first embodiment of the present invention. [Figure 3] It is a block diagram showing an example of the internal configuration of an external recognition device according to a first embodiment of the present invention. [Figure 4] It is a block diagram showing an example of the hardware configuration of a computer according to a first embodiment of the present invention. [Figure 5] It is a diagram showing an example of an image captured by a front camera according to a first embodiment of the present invention. [Figure 6] It is a diagram showing an example of the processing of a feature matching unit according to a first embodiment of the present invention. [Figure 7] It is a diagram showing the state of processing in which a feature matching unit according to a first embodiment of the present invention performs template matching using a left camera image and a front camera image. [Figure 8] It is a block diagram showing an example of the internal configuration of an external recognition device according to a second embodiment of the present invention. [Figure 9] It is a block diagram showing an example of the internal configuration of an external recognition device according to a third embodiment of the present invention. [Figure 10] It is a diagram showing an example of an area where a vehicle and a background are mixed according to a third embodiment of the present invention. [Figure 11] It is a block diagram showing an example of the internal configuration of an external recognition device according to a fourth embodiment of the present invention. [Figure 12]It is a block diagram showing an internal configuration example of a geometric transformation estimation unit and a feature matching unit according to a modified example of an external world recognition device according to a fourth embodiment of the present invention. [Figure 13] It is a block diagram showing an internal configuration example of an external world recognition device according to a fifth embodiment of the present invention.

Embodiments for Carrying Out the Invention

[0015] Hereinafter, embodiments for carrying out the present invention will be described with reference to the accompanying drawings. In this specification and the drawings, components having substantially the same function or configuration are denoted by the same reference numerals, and redundant descriptions are omitted.

[0016] [First Embodiment] FIG. 1 is a diagram showing the position of a camera mounted on a vehicle 1 according to the first embodiment, a world coordinate system based on the vehicle 1, and an image coordinate system.

[0017] FIG. 1A shows a configuration example in which a front camera 10, a left camera 20, and a right camera 30 are provided on the front, left side, and right side of the vehicle 1, respectively. In the following description, when the front camera 10, the left camera 20, and the right camera 30 are not distinguished, they are simply referred to as "cameras".

[0018] In FIG. 1A, the shooting ranges 10a of the front camera 10, the shooting range 20a of the left camera 20, and the shooting range 30a of the right camera 30 are shown separated by a one-dot chain line. The shooting ranges 10a and 20a partially overlap. Similarly, the shooting ranges 10a and 30a partially overlap, and the shooting ranges 20a and 30a partially overlap. In these overlapping ranges, the same objects are reflected in the images taken by each camera.

[0019] Based on images captured by a camera installed on vehicle 1, the ECU (external environment recognition device 100 shown in Figure 3, described later) mounted on vehicle 1 estimates and acquires object information of objects surrounding vehicle 1, such as pedestrians, other vehicles, white lines, and the road surface, as well as physical quantities such as the distance from vehicle 1 to the object, the size of the object, and the velocity of the object. Then, another ECU (an automatic driving control ECU not shown) determines the control quantities for vehicle 1 based on the acquired physical quantities and performs automatic driving or driving assistance for vehicle 1. In the following explanation, objects that may affect the driving of vehicle 1 (road surface, people, guardrails, white lines, etc.) will be referred to as "objects".

[0020] Figure 1B shows an example of a world coordinate system with the position of vehicle 1 as the origin. In this world coordinate system, the X-axis is taken in the direction of vehicle 1's movement, the Y-axis is taken in the horizontal direction perpendicular to the X-axis, and the Z-axis is taken in the vertical direction. The positions of targets other than vehicle 1 are estimated using the world coordinate system. Therefore, the distance from vehicle 1 to the targets is also estimated.

[0021] There are several methods for estimating the distance from vehicle 1 to a target. A typical method for estimating distance is to measure the distance to the target using the principle of triangulation when multiple cameras capture images of the same target. In this method, any point in world coordinates on the road is transposed in a homogeneous coordinate system as the matrix X=(X,Y,Z,1). T Let R be the external parameter matrix related to the camera's rotation angle, and T be the external parameter matrix related to the camera's mounting position. Then, let P = (R|T) be the matrix of the external parameter matrices R and T, and let K be the internal parameter matrix that manages the camera's internal state, such as its focal length and optical center.

[0022] Figure 1C shows an example of an image coordinate system for identifying an object within an image. In this image coordinate system, the upper left corner of the image is the origin, the u-axis is taken horizontally, and the v-axis is taken vertically. The position of the object within the image captured by the camera is then identified as (u,v).

[0023] Here, the transpose matrix X is the image coordinates in a homogeneous coordinate system, and the transpose matrix u=(u,v,1) is obtained. TAssuming a perspective projection model with a scale parameter of s and a lens, equation (1) holds for each camera. The scale parameter s is a scalar value. The symbolic subscript in equation (1) represents the type of camera. Here, to distinguish between the two cameras, the subscript for one camera is set to "0" and the subscript for the other camera is set to "1". Note that if there were no scale parameter s in equation (1), the right-hand side of equation (1) would become a constant multiple of the left-hand side. Therefore, the scale parameter s is provided to prevent the right-hand side of equation (1) from becoming a constant multiple of the left-hand side.

[0024]

number

[0025] Given the image coordinates of one camera, it is impossible to reconstruct the three-dimensional point of a target from these image coordinates without certain preconditions. However, if there are two cameras capturing the same position X, and the same location at position X is represented by the image coordinates (u0, u1), then the three-dimensional point X can be estimated by solving the least squares method using equation (2). However, depending on the form of the equation transformation, it is not necessary to limit the solution to using equation (2).

[0026]

number

[0027] Now, let's refer to Figure 1A to explain the meaning of equation (2). The field of view of the front camera 10 is the vertex angle of the triangle at the installation position of the front camera 10, as shown in the shooting range 10a. Similarly, the field of view of the left camera 20 is the vertex angle of the triangle at the installation position of the left camera 20, as shown in the shooting range 20a. The field of view of the right camera 30 is the vertex angle of the triangle at the installation position of the right camera 30, as shown in the shooting range 30a.

[0028] The area where the field of view of the front camera 10 and the field of view of the left camera 20 overlap is located in the front left of vehicle 1. Therefore, the distance to an object captured in the overlapping field of view of the two cameras can be estimated using the principle of triangulation. Similarly, the distance to the area where the field of view of the front camera 10 and the field of view of the right camera 30 overlap (front right of vehicle 1) can also be estimated using the principle of triangulation.

[0029] To efficiently search for an object that appears in both images at the same location in both images, the epipole constraint equation shown in equation (3) is used. When using this equation, given the image coordinates of an object that appears in one image, the candidate positions of the image coordinates of the object that appears in the other image lie on a straight line, so the search should be conducted along that straight line rather than across the entire image. The basic matrices used in the epipole constraint equation are derived from the external parameter matrices and internal parameter matrices of the two cameras shown above (for example, the front camera 10 and the left camera 20).

[0030]

number

[0031] Furthermore, by appropriately using the parameter set described above and applying a preprocessing step called parallelization, it is possible to transform the images into one where any point X in world coordinates is aligned at the same height. The search for the same location of an object can be performed with a simple process of searching horizontally in the parallelized image. This is the principle of a parallel stereo camera. In any case, it is possible to estimate the same location of an object from two images using information such as parameters specific to the camera that captured each image and installation conditions. Here, the two images are assumed to be combinations of images from front camera 10 and left camera 20, front camera 10 and right camera 30, and left camera 20 and right camera 30 in Figure 1. In the following, the process of searching two images to find the same location of an object will be simply referred to as "matching".

[0032] There are several methods for matching image information, with template matching being a typical example. Template matching involves extracting a small region, such as a rectangle, from one image around the image coordinates of a target of interest, and then searching for a region with a similar pixel distribution within that rectangle by moving through the small rectangular region from the other image using an appropriate cost function. Basic comparison methods include the sum of absolute differences and the sum of squared differences.

[0033] Another method for matching image information involves first searching for image coordinates of areas with high image features, such as the corners of a target, and identifying these as feature points. Then, the image features around these feature points are vectorized (feature-enhanced) and stored. Representative feature-enhanced techniques include SIFT (Scale Invariant Feature Transform) and ORB (Oriented FAST and Rotated BRIEF), and deep learning-based methods also exist. Deep learning-based methods match features between two images.

[0034] Here, we will explain images in which the same target is captured in the overlapping area of ​​the fields of view of two cameras. Figure 2 shows examples of images of the foreground of vehicle 1 captured by the front camera 10 and the left camera 20 shown in Figure 1. The image from the front camera 10 is called the "front camera image," and the image captured by the left camera 20 is called the "left camera image."

[0035] In front of vehicle 1 is a pedestrian crossing, and in the background is a rectangular prism-shaped landmark. Both the pedestrian crossing and the rectangular prism-shaped landmark are visible in the left camera image and the front camera image. Here, we represent areas 21 and 22 of the pedestrian crossing visible in the left camera image with dashed rectangular frames. Similarly, we represent areas 11 and 12 of the pedestrian crossing visible in the front camera image with dashed rectangular frames. The areas corresponding to the rectangular frames in the left camera image, such as the combination of areas 21 and 11 and the combination of areas 22 and 12, also exist in the front camera image. However, the appearance of the inside of the rectangular frames (parts of the pedestrian crossing) differs significantly between the left camera image and the front camera image.

[0036] Matching images of the same location in two images that appear significantly differently is difficult using the template matching method described above, and even matching using feature points or feature quantities is challenging. Normally, template matching compares pixels at the same location in an image, but matching becomes difficult when the appearance differs. Furthermore, creating feature points or features from images that appear differently also makes it difficult to match the images. As a result, shapes that are actually different may be matched, or shapes that should be considered identical may not be found in the images and therefore not matched. If the matching of the same target in different images is not performed correctly in this way, there is a possibility of inaccurate estimation of the distance to that target.

[0037] Therefore, the inventors investigated what kind of transformation should be applied to one of the images to reduce the difference in targets when matching two images that look significantly different. Below, the external environment recognition device and target matching method according to each embodiment of the present invention, which can improve the performance of matching identical objects appearing in two images, will be described with reference to the drawings from Figure 3 onward. The external environment recognition device's functions are realized by software configured in the ECU mounted on the vehicle 1.

[0038] Figure 3 is a block diagram showing an example of the internal configuration of the external environment recognition device 100 according to the first embodiment. The external environment recognition device 100 is, for example, a form of ECU mounted on a vehicle 1. In the following description, the vehicle 1 on which the external environment recognition device 100 is mounted will also be referred to as "our own vehicle 1" to distinguish it from other vehicles 2.

[0039] The external environment recognition device 100 is a device mounted on vehicle 1. The external environment refers to, for example, the area outside vehicle 1 on which the external environment recognition device 100 is mounted, and the external environment recognition device 100 recognizes various targets that exist in the external environment. The target recognition results include information such as the attributes of the target, the size of the target, and the distance from vehicle 1 to the target. The external environment recognition device 100 receives two images, each captured by two imaging units 40 and 41. The imaging unit 40 is, for example, the left camera 20 shown in Figure 1, and outputs the image 40a captured by the left camera 20 to the target detection unit 101. The imaging unit 41 is, for example, the front camera 10 shown in Figure 1, and outputs the image 41a captured by the front camera 10 to the target detection unit 102. The camera combinations shown in Figure 1 for the imaging units 40 and 41 may also be front camera 10 and right camera 30, and left camera 20 and right camera 30.

[0040] The external environment recognition device 100 includes two target detection units 101 and 102, camera parameters 103, a matching condition determination unit 104, a feature matching unit 105, and a position estimation unit 106.

[0041] The target detection unit 101 performs target detection processing on the image input from the imaging unit 40. This first target detection unit (target detection unit 101) detects a target based on a first image (image 40a) acquired from at least the first imaging unit (imaging unit 40) among a plurality of imaging units whose imaging fields of view to the outside world overlap in at least a portion. The target detection unit 102 performs target detection processing on the image input from the imaging unit 41. This second target detection unit (target detection unit 102) detects targets based on the second image (image 41a) acquired from the second imaging unit (imaging unit 41) among the multiple imaging units.

[0042] The target detection processing performed by the target detection units 101 and 102 is, for example, the process of detecting vehicle 2, white lines, pedestrians, the drivable area of ​​vehicle 1, guardrails, side walls, etc., from the input images 40a and 41a. A portion of each image shown in regions 11, 12, 21, and 22 as shown in Figure 2 is output to the subsequent functional unit as the result of the target detection processing. In addition, the target detection units 101 and 102 can also determine the attributes of each target by performing image analysis on each target through the target detection processing. The result of the target detection processing by the target detection unit 101 is output to the imaging plane estimation unit 111 of the matching condition determination unit 104. The result of the target detection processing by the target detection unit 102 is output to the feature matching unit 105. In the following description, the result of the target detection processing will also be called the "target detection result".

[0043] Furthermore, the image captured by the imaging unit 40 is input to the matching condition determination unit 104 via the target detection unit 101, and then input to the feature matching unit 105 via the matching condition determination unit 104. The image captured by the imaging unit 41 is input to the feature matching unit 105 via the target detection unit 102.

[0044] The matching condition determination unit 104 determines matching conditions for the feature matching unit 105 to match the positions of the same targets detected in images 40a and 41a. This matching condition determination unit 104 includes an imaging plane estimation unit 111 and a geometric transformation estimation unit 112.

[0045] The imaging plane estimation unit (imaging plane estimation unit 111) estimates the imaging plane in the first image (image 40a) based on the target detection result by the target detection unit (target detection unit 101). Based on the target detection result input from the target detection unit 101, the imaging plane estimation unit 111 estimates what kind of plane each target is composed of in the real world. The plane that constitutes each target is called the imaging plane and is expressed by the plane equation of the imaging plane. The imaging plane estimation unit (imaging plane estimation unit 111) estimates the imaging plane based on the size of the target and the distance to the target. The imaging plane estimation unit (imaging plane estimation unit 111) also estimates the imaging plane in the image region in which the target is captured. The plane equation can be expressed, for example, by parameters (α, β, γ, δ). The equation of the imaging plane estimated by the imaging plane estimation unit 111 is output to the geometric transformation estimation unit 112.

[0046] The image transformation estimation unit (geometric transformation estimation unit 112) estimates image transformation parameters (geometric transformation parameters) that match the imaging plane in the second image (image 41a) from among the multiple imaging units, based on the imaging plane information estimated by the imaging plane estimation unit (imaging plane estimation unit 111) and the imaging parameters (camera parameters 103) of the multiple imaging units. For example, the geometric transformation estimation unit 112 estimates the geometric transformation parameters between images at the location being captured by imaging units 40 and 41, based on the equation of the imaging plane input from the imaging plane estimation unit 111 and the camera parameters 103. Examples of camera parameters 103 include K and P included in equation (1) above. However, camera parameters 103 may be fixed values. The imaging plane estimation unit 111 and the geometric transformation estimation unit 112 may be located on the target detection unit 102 side.

[0047] The feature matching unit (feature matching unit 105) compares the features of the target detected from the first image (image 40a), which has been image-transformed using image transformation parameters (geometric transformation parameters) from the image transformation estimation unit (geometric transformation estimation unit 112), with the features of the target detected from the second image (image 41a). The feature matching unit 105 compares the features of the transformation result obtained by geometric transformation applied to the target detected by the target detection unit 101 with the features of the target detection result input from the target detection unit 102. For this reason, the feature matching unit 105 transforms the shape of the target detected from image 40a using geometric transformation parameters estimated by the geometric transformation estimation unit 112. Subsequently, the feature matching unit 105 compares the transformed shape of the target with the shape of the target detected from image 41a and calculates the distance from the vehicle 1 to the target. The distance from the vehicle 1 to the target is output to the position estimation unit 106 as the matching result from the feature matching unit 105. The matching results may be expressed, for example, as a value where a mismatch is 0% and a perfect match is 100% for each image region containing the target being matched. The feature matching unit 105 may then determine that the matching result is highly reliable if the matching result is 80% or higher, and output the highly reliable matching result to the position estimation unit 106.

[0048] The position estimation unit (position estimation unit 106) estimates the three-dimensional position of a target based on the results of the matching process performed by the feature matching unit (feature matching unit 105). For example, the position estimation unit 106 estimates the three-dimensional position of the target in the world coordinate system (see Figure 1B) (referred to as the "position estimation result") based on the matching results from the feature matching unit 105. At this time, the position estimation unit 106 can calculate the distance from the vehicle 1 to the target and include the distance to each target in the position estimation result. The position estimation result can be used for vehicle control of the vehicle 1 by other ECUs mounted on the vehicle 1 (for example, an ECU for automatic driving control) or to obtain depth information of targets in the vicinity of the vehicle 1.

[0049] Next, the hardware configuration of the computer 80 that constitutes the external environment recognition device 100 will be described. Figure 4 is a block diagram showing an example of the hardware configuration of computer 80. Computer 80 is an example of hardware used as a computer capable of operating as the external environment recognition device 100 according to this embodiment. The external environment recognition device 100 according to this embodiment realizes an image processing method performed by the cooperation of each functional block shown in Figure 3 when computer 80 (computer) executes a program.

[0050] Computer 80 comprises a CPU (Central Processing Unit) 81, ROM (Read Only Memory) 82, and RAM (Random Access Memory) 83, each connected to a bus 84. Furthermore, computer 80 includes non-volatile storage 85 and a network interface 86.

[0051] The CPU 81 reads the program code of the software that implements each function according to this embodiment from the ROM 82, loads it into the RAM 83, and executes it. Variables and parameters that occur during the calculation process of the CPU 81 are temporarily written to the RAM 83, and these variables and parameters are read out by the CPU 81 as appropriate. However, an MPU (Micro Processing Unit) may be used instead of the CPU 81. The functions of each functional part in the external environment recognition device 100 are realized by the CPU 81, ROM 82, and RAM 83.

[0052] Examples of non-volatile storage 85 include HDDs (Hard Disk Drives), SSDs (Solid State Drives), flexible disks, optical disks, magneto-optical disks, CD-ROMs, CD-Rs, magnetic tapes, or non-volatile memory. This non-volatile storage 85 stores the OS (Operating System), various parameters, and programs necessary for the computer 80 to function. ROM 82 and non-volatile storage 85 store programs and data necessary for the CPU 81 to operate, and are used as an example of a computer-readable, non-transient storage medium containing programs executed by the computer 80. Various values, such as camera parameters 103, are stored in RAM 83 or non-volatile storage 85 and read out as needed.

[0053] The network interface 86 can use, for example, a NIC (Network Interface Card), and various types of data can be sent and received with external devices via a LAN (Local Area Network), dedicated line, etc., connected to the terminals of the NIC.

[0054] The processing performed by the imaging plane estimation unit 111 will now be explained with reference to Figure 5. Figure 5 shows an example of an image captured by the front camera 10. Figure 5A is an image showing another vehicle 2, a person, a guardrail, and a side wall in the straight-ahead direction. In Figure 5A, region 51 represents the road surface detection result, region 52 represents the guardrail detection result, region 53 represents the person detection result, and region 54 represents the side wall detection result. Regions 51 to 54 are used as the target detection results.

[0055] The imaging plane estimation unit 111 estimates parameters (α, β, γ, δ) such that the imaging plane follows the plane equation αX + βY + γZ + δ = 0, based on the detection result at a certain image coordinate.

[0056] For example, suppose the target detection unit 101 detects a target from the image and the result is a road surface like the area 51. In this case, the imaging plane estimation unit 111 estimates the plane captured at those image coordinates as the vehicle contact surface on which the vehicle 1 is traveling, with Z=0, i.e., (α,β,γ,δ)=(0,0,1,0). If the target detection result is a slope, the plane equation can be expressed as, for example, hX+Z+k=0. In other words, the imaging plane estimation unit 111 can estimate it as (α,β,γ,δ)=(h,0,1,k).

[0057] Figure 5B is an image showing another vehicle 2 diagonally in front of the vehicle 1. Here, a scenario is assumed where there is a curved road in front of the vehicle 1. For example, suppose the target detection unit 101 detects another vehicle 2 and outputs the detection result as a rectangular bounding box as shown in region 55. In this case, the imaging plane estimation unit 111 estimates the plane equations of each face of the rectangular prism represented in region 55, using the front, back, and sides of the rectangular prism as the imaging planes. Note that the front of the bounding box in region 55 is the front face of vehicle 2, the back is the rear face of vehicle 2, and the sides are the left and right faces of vehicle 2.

[0058] Here, we assume that the other vehicle 2 detected by the target detection unit 101 is upright on the plane Z=0. Therefore, the imaging plane estimation unit 111 assumes that the plane equation of the face of the rectangular parallelepiped represented by region 55 is, for example, aX+bY+d=0. In other words, the imaging plane estimation unit 111 estimates the parameters of the plane equation as (α,β,γ,δ)=(a,b,0,d). In this specification, estimating the parameters of the plane equation is also referred to as "estimating the imaging plane".

[0059] However, the actual coefficient value differs depending on the face of the rectangular parallelepiped. The actual coefficient value is the estimated distance from vehicle 1 to other vehicle 2, which is output from the sensor. For example, by using a known measurement principle that allows distance to be estimated from the position of the contact surface between the road surface and other vehicle 2, as well as the installation position and angle of the monocular camera, the detection result may include distance estimation information. In this case, the imaging plane estimation unit 111 can also estimate the plane equation of the side of vehicle 1, etc., based on the distance estimation information included in the detection result. For this reason, the imaging plane estimation unit 111 can also output the estimated plane equation as the estimation result.

[0060] Furthermore, consider the case where the detection result is a pedestrian or an obstacle. In this case, if only a rectangular bounding box is output as the detection result, as shown in regions 52 and 53 of Figure 5A, the imaging plane estimation unit 111 cannot clearly estimate the imaging plane. Therefore, the imaging plane estimation unit 111 assumes, for example, that a plane perpendicular to the direction of travel of the vehicle 1 is formed. The imaging plane estimation unit 111 then expresses the plane equation as X+e=0. In other words, the imaging plane estimation unit 111 can estimate the parameters of the plane equation as (α,β,γ,δ)=(1,0,0,e). Additionally, if a rectangular bounding box is output as the detection result for region 52 where a guardrail is visible and region 54 where a side wall is visible, as shown in Figure 5A, the imaging plane estimation unit 111 can estimate the plane equation of the imaging plane corresponding to the side of the rectangular prism. The imaging plane estimation unit 111 may also estimate the parameters of the plane equation (α, β, γ, δ) for each object by assuming that the object is perpendicular to the direction of arrival of the light rays to each camera, or to the optical axis of each camera.

[0061] Incidentally, it is conceivable that there may be areas where attributes are not assigned to three-dimensional objects such as poles or trees, and the attributes of the detection result are unknown. In this case, the imaging plane estimation unit 111 assumes a three-dimensional object of a predetermined height at the position of the pole or tree, and assumes that a plane perpendicular to the direction of travel of the vehicle 1 is formed. Furthermore, the imaging plane estimation unit 111 can also estimate the parameters of the plane equation by assuming a rectangular bounding box in the area where the attributes of the detection result are unknown, as was done when the detection result was identified as a pedestrian.

[0062] As another example of how the imaging plane estimation unit 111 estimates the imaging plane, it is also possible to use quadratic or polynomial equations instead of simple linear equations for the target detection result. In that case, the imaging plane estimation unit 111 estimates the imaging plane by dividing the polynomial equation into piecewise linear forms according to the position of interest on the image. As described above, the imaging plane estimation unit 111 estimates the imaging plane of the target based on the target detection result.

[0063] Next, we will explain how the geometric transformation estimation unit 112 estimates the geometric transformation parameters. The geometric transformation estimation unit 112 estimates the geometric transformation parameters of the two images 40a and 41a captured by the two cameras (imaging units 40 and 41) based on the estimation result of the imaging plane estimated by the imaging plane estimation unit 111 and the camera parameters 103. An example of this method for estimating geometric transformation parameters will be described below.

[0064] First, the estimation result of the imaging plane estimation unit 111, which has been processed on one of the images (image 40a captured by the imaging unit 40), is used. The geometric transformation estimation unit 112, when focusing on a certain pixel coordinate in the image captured by the imaging unit 40, converts the plane equation for that pixel coordinate into a plane transformation matrix. For example, the geometric transformation estimation unit 112 rearranges the plane equation with respect to Z and introduces the plane transformation matrix C as shown in equation (4). Then, it compresses the four-dimensional world coordinate (homogeneous coordinate) X into one dimension and rearranges it as X'.

[0065]

number

[0066] The geometric transformation estimation unit 112 can transform s0u0 and s1u1 shown in equation (1) into equations (5) and (6) by using X'. The geometric transformation estimation unit 112 can derive the projection transformation matrix H (geometric transformation parameters) into equation (7) by solving equations (5) and (6) simultaneously. The camera parameters 103 input to the matching condition determination unit 104 are used in equations (5) to (7) because they contain K and P, which are included in equation (1). Equation (7) expresses the relative relationship between the two images 40a and 41a.

[0067]

number

[0068] Furthermore, depending on the parallelization process described above, the camera's installation position, and the plane equation, the geometric transformation estimation unit 112 may be able to perform the geometric transformation using affine transformations or shear transformations, which require fewer transformation parameters than projection transformations. If the geometric transformation estimation unit 112 uses fewer transformation parameters for the geometric transformation, it will lead to a reduction in the calculation time performed by the geometric transformation estimation unit 112. Therefore, the geometric transformation estimation unit 112 may switch the type of transformation according to the number of transformation parameters. In the following processes, various types of transformation processes will be collectively referred to as "geometric transformations" as an example of image transformation.

[0069] The feature matching unit 105 compares the images using the estimation results from the geometric transformation estimation unit 112 and the images captured by the imaging units 40 and 41. An example of the process by which the feature matching unit 105 compares two images will be explained with reference to Figure 6. Figure 6 shows an example of the processing of the feature matching unit 105. Images 40a and 41a are input to the feature matching unit 105 from the imaging units 40 and 41 shown in Figure 3.

[0070] The geometric transformation parameters are used as the estimation results of the geometric transformation estimation unit 112 for any point in the image 40a. The geometric transformation parameters are represented as the projection transformation matrix H in equation (7) above. The feature matching unit 105 receives the image 40a captured by the imaging unit 40 and the geometric transformation parameters estimated by the geometric transformation estimation unit 112 as input, and performs geometric transformation processing on the corresponding location in the image 40a (S1).

[0071] Furthermore, the feature matching unit 105 receives the image 41a captured by the imaging unit 41 via the target detection unit 102. The feature matching unit 105 then performs image matching (S2) between the image 40a, which underwent geometric transformation processing in step S1, and the other image 41a, which was input via the target detection unit 102. However, the area to be geometrically transformed is the region surrounding the location where the same part of the target was detected in images 40a and 41a.

[0072] In the image matching process, template matching is performed between the geometrically transformed area of ​​image 40a and the corresponding area of ​​image 41a, as well as feature point extraction and feature quantity extraction. In the case of projection transformation, since the geometric transformation itself includes a movement component, the matching area is equivalent to being estimated to some extent around the position given by the geometric transformation. Therefore, the feature matching unit 105 only needs to search around the position given by the geometric transformation. After step S2, the processing result is output to the position estimation unit 106.

[0073] Here, an example of template matching will be explained with reference to Figure 7. Figure 7 shows the process by which the feature matching unit 105 performs template matching using the left camera image and the front camera image. Here, the left camera image is image 40a shown in Figure 7, and the front camera image is image 41a shown in Figure 7.

[0074] As shown in Figure 2, the left camera image and the front camera image contain regions 21 and 22, which are characteristic parts of the intersection (where the white lines intersect). These characteristic parts also exist in the front camera image as regions 11 and 12, respectively. However, the shapes of regions 21 and 11 are different, and the shapes of regions 22 and 12 are also different, so it was not possible to simply match the shapes of each region using conventional methods.

[0075] On the other hand, by utilizing the projection transformation (geometric transformation) according to this embodiment, the targets in regions 21 and 22 included in the left camera image are transformed into the shapes of regions 71 and 72 shown at the bottom of Figure 7. At this time, the shapes included in regions 71 and 72 become similar to the shapes included in regions 11 and 12, respectively. Therefore, the feature matching unit 105 can perform the matching process more easily than searching for the shapes of regions 21 and 22 before geometric transformation from the front camera image. Thus, when the shape of the region containing the feature portion is geometrically transformed, the feature matching unit 105 can perform the matching process more easily.

[0076] Furthermore, if the plane equation is constant within a certain image region, the feature matching unit 105 may perform a projection transformation on that region before performing template matching. Alternatively, the feature matching unit 105 may perform a projection transformation for each template.

[0077] The same applies to the process of extracting feature points and feature quantities within an image (referred to as "feature point and feature quantity extraction"). The feature matching unit 105 may perform feature point and feature quantity extraction after first performing a projection transformation on the image. Alternatively, the feature matching unit 105 may incorporate projection transformation into the feature point and feature quantity extraction process and then perform feature point and feature quantity extraction.

[0078] The position estimation unit 106 calculates the coordinate points of the target in the world coordinate system using equation (2) above, based on the corresponding position obtained as a result of the matching by the feature matching unit 105. The external environment recognition device 100, with the configuration described above, can respond to changes in the environment while vehicle 1 is in motion, such as vehicle 1, pedestrians, and roads, while improving the accuracy of matching targets in the two images and the accuracy of estimating the distance from vehicle 1 to the target.

[0079] In the external environment recognition device 100 according to the first embodiment described above, in order to estimate the distance from the vehicle 1 to the target, the device estimates a plane equation of the imaging plane corresponding to the position in the image 40a for each target, in accordance with the driving environment which changes dynamically as the vehicle moves. The external environment recognition device 100 then performs a geometric transformation for each target and compares the features of the target detected from the image 41a. Therefore, the external environment recognition device 100 can estimate the distance to the target even if an unspecified target is captured in two images that appear differently. Furthermore, since the position estimation unit 106 estimates the position of the target whose features have been matched, it is possible to reduce the error in the distance to the target.

[0080] [Second Embodiment] Next, an example of the configuration and processing of the external environment recognition device 100A according to a second embodiment of the present invention will be described with reference to Figure 8. Figure 8 is a block diagram showing an example of the internal configuration of the external environment recognition device 100A according to the second embodiment.

[0081] The external environment recognition device 100A performs imaging plane estimation processing and geometric transformation parameter estimation for two images input from imaging units 40 and 41, respectively. Therefore, the external environment recognition device 100A includes multiple imaging plane estimation units (imaging plane estimation units 111, 113) and multiple image transformation estimation units (geometric transformation estimation units 112, 114), respectively, for the target detected from the first image (image 40a) and the target detected from the second image (image 41a). As shown in Figure 8, the matching condition determination unit 104A of the external environment recognition device 100A includes imaging plane estimation units 111, 113 and geometric transformation estimation units 112, 114. The imaging plane estimation units 111, 113 have the same function, and the geometric transformation estimation units 112, 114 have the same function.

[0082] The target detection unit 101 detects targets from images captured by the imaging unit 40. Based on this detection result, the imaging plane estimation unit 111 estimates the imaging plane, and the geometric transformation estimation unit 112 estimates the geometric transformation parameters. Similarly, the target detection unit 102 detects targets from the image captured by the imaging unit 41. Based on the target detection result, the imaging plane estimation unit 113 estimates the imaging plane, and the geometric transformation estimation unit 114 estimates the geometric transformation parameters.

[0083] The feature matching unit 105A receives two results through the geometric transformation estimation units 112 and 114. Therefore, the function of the feature matching unit 105A is different from the function of the feature matching unit 105 in the first embodiment shown in Figure 3.

[0084] Several methods are conceivable for implementing the functionality of the feature matching unit 105A. In one method, the feature matching unit 105A obtains a comparison result between the geometric transformation of image 40a using geometric transformation parameters estimated by one of the geometric transformation estimation units 112 and image 41a. The feature matching unit 105A also obtains a comparison result between the geometric transformation of image 41a using geometric transformation parameters estimated by the other geometric transformation estimation unit 114 and image 40a. Then, the feature matching unit 105A compares the two comparison results obtained. If it is determined from the two comparison results that the same target was matched at the same location, both comparison results can be considered to have a high degree of confidence.

[0085] Therefore, the feature matching unit (feature matching unit 105A) obtains the result of image transformation (geometric transformation) of the first image (image 40a) using the image transformation parameters (geometric transformation parameters) estimated by one image transformation estimation unit (geometric transformation estimation unit 112), and the result of matching the second image (image 41a) with the second image (image 41a), and obtains the result of geometric transformation of the second image (image 41a) using the image transformation parameters (geometric transformation parameters) estimated by the other image transformation estimation unit (geometric transformation estimation unit 114), and compares the respective matching results to assign a high level of confidence to the matching result when the target detected in the first image (image 40a) and the target detected in the second image (image 41a) are matched at the same three-dimensional position. The level of confidence assigned to the matching result here may be expressed as a value where, for each target detected in each image, the mismatch is 0% if they are not matched at the same three-dimensional position, and 100% if they are matched at the same three-dimensional position. Furthermore, the feature matching unit 105 may output a highly reliable matching result to the position estimation unit 106 if the confidence level is 80% or higher.

[0086] Furthermore, the feature matching unit 105A may also match the estimated imaging plane. In this case, the feature matching unit 105A compares the imaging plane estimated by the imaging plane estimation unit 111 with the imaging plane estimated by the imaging plane estimation unit 113. At this time, the feature matching unit 105A uses the imaging plane estimated by the imaging plane estimation unit 113 as a reference, and if the imaging plane estimated by the imaging plane estimation unit 111 is different, the reliability of the matching result is low, and the unit determines it to be a matching error. The feature matching unit 105A can then choose not to output the feature matching result to the position estimation unit 106 for targets that are the target of the estimated imaging plane. In this way, the feature matching unit 105A outputs the matching result to the position estimation unit 106, retaining only the matching results with high reliability, or it may judge matching results with low reliability as noise and not output the matching result. As a result, the position estimation unit 106 can estimate the position and distance of targets limited to those targets with high reliability in the matching result.

[0087] In the external environment recognition device 100A according to the second embodiment described above, the matching condition determination unit 104A is configured to include imaging plane estimation units 111, 113 and geometric transformation estimation units 112, 114. The feature matching unit 105A can determine the reliability of the matching result by comparing the estimation results using one of the estimation results output from the geometric transformation estimation units 112, 114 as a reference. Therefore, it becomes easier to determine whether the target has been correctly matched based on the reliability of the matching result by the feature matching unit 105A, and the position estimation unit 106 can also accurately estimate the position of the target.

[0088] [Third Embodiment] Next, an example of the configuration and processing of the external environment recognition device 100B according to the third embodiment of the present invention will be described with reference to Figures 9 and 10.

[0089] When multiple targets with different attributes are mixed within the image region specified by the target detection unit 101, several methods can be considered for the matching condition determination unit 104B shown in Figure 9 to accurately detect targets of each attribute. For example, when the matching condition determination unit 104B receives the detection result of the target detected by the target detection unit 101, it can prioritize measuring targets that are in the foreground of the target to be matched, based on the attributes of the target detected by the target detection unit 101 and the plane equation estimated by the imaging plane estimation unit 111.

[0090] Furthermore, if the target detection unit 101 detects multiple attributes for each target, the matching condition determination unit 104B may select the attributes of the dominant target. A dominant target is, for example, the target that is in the foreground when multiple targets are visible in the image area. In this case, the feature matching unit 105 matches the features of the targets according to the attributes of the dominant target that is in the foreground relative to the controlled targets, and estimates the position of the targets. Here, we will describe the matching condition determination unit 104B of the external environment recognition device 100B, which is capable of selecting a dominant target from among multiple targets detected from images 40a and 41a.

[0091] Figure 9 is a block diagram showing an example of the internal configuration of the external environment recognition device 100B according to the third embodiment. In the external environment recognition device 100B according to the third embodiment, when multiple targets with different attributes are mixed within an image area specified by the target detection unit 101, a process is performed to select a specific target as the target for template matching by the feature matching unit 105.

[0092] The matching condition determination unit 104B of the external environment recognition device 100B includes an imaging plane estimation unit 111, a geometric transformation estimation unit 112, and a transformation selection unit 115. The transformation selection unit (transformation selection unit 115) selects the closest target from among multiple targets detected from the first image (image 40a) and the second image (image 41a), and selects the image transformation parameters (geometric transformation parameters) estimated by the image transformation estimation unit (geometric transformation estimation unit 112) for the selected target. For example, the transformation selection unit 115 selects the geometric transformation parameters estimated for the target with dominant attributes from the geometric transformation parameters for multiple targets estimated by the geometric transformation estimation unit 112. The geometric transformation parameters selected by the transformation selection unit 115 are then output to the feature matching unit 105.

[0093] The feature matching unit (feature matching unit 105) performs image transformation on the detected target from the first image (image 40a) using the image transformation parameters (geometric transformation parameters) selected by the transformation selection unit (transformation selection unit 115). The feature matching unit 105 then compares the features of the geometrically transformed target with the features of the target detected by the target detection unit 102 and outputs the comparison result to the position estimation unit 106.

[0094] Here, a specific example of the processing performed by the matching condition determination unit 104B shown in Figure 9 will be explained with reference to Figure 10. Here, as an example, the matching condition determination unit 104B uses a plane equation to select the attributes of the target that is close to the vehicle 1, i.e., closer to the vehicle.

[0095] Figure 10 shows examples of areas 56 and 57 where vehicle 2 and the background are mixed. Assume that vehicle 2 is traveling toward vehicle 1.

[0096] In region 56, a portion of the left side of vehicle 2 and the road surface in the background are mixed together. Similarly, in region 57, a portion of the lower side of vehicle 2 and the road surface in the background are mixed together. Typically, when the road surface is included in the upper part of an image, the three-dimensional object is in front of the road surface. When the frame targeted for template matching by the feature matching unit 105 is region 56, the shape of vehicle 2, which is in front of the background, is given priority as the target for matching.

[0097] Conversely, if the image includes both a three-dimensional object and a road surface, the road surface at the bottom of the image is in front of the three-dimensional object. When the frame targeted for template matching by the feature matching unit 105 is region 57, the point where the lower part of vehicle 2 meets the road surface is closer to the vehicle itself, so the shape of the road surface in front of vehicle 1 is given priority as the target for matching.

[0098] As another example, from the perspective of vehicle control, priority may be given to objects (e.g., vehicle 2) that would have a significant impact on the vehicle's movement if it were to come into contact with it, and these could be determined as matching targets simply based on the attribute information of each vehicle. In this case, the conversion selection unit (conversion selection unit 115) prioritizes selecting targets that have attributes that have a significant impact on the vehicle's movement. As a result, the position estimation unit 106 prioritizes estimating the distance to targets with attributes that have a significant impact on the vehicle's movement, allowing the vehicle to take control such as avoiding targets whose distance has been estimated.

[0099] In the external environment recognition device 100B according to the third embodiment described above, when multiple targets with different attributes are detected in a single image region of image 40a, a geometric transformation parameter can be selected for one of the targets, the geometric transformation of this target can be performed, and it can be compared with the target in the corresponding image region of image 41a. As a result, the feature matching unit 105 can easily match the targets to be matched in images 40a and 41a, and the position estimation unit 106 can accurately estimate the position of the matched targets.

[0100] [Fourth Embodiment] Next, an example of the configuration and processing of the external environment recognition device 100C according to the fourth embodiment of the present invention will be described with reference to Figures 11 and 12.

[0101] The road surface model (plane equation) estimated by the imaging plane estimation unit 111 is expected to deviate from the actual plane model as the distance of the target from the vehicle 1 increases. In other words, the plane equation estimated by the imaging plane estimation unit 111 based on the detection results of the target detection unit 101 may contain errors in the plane equation itself. Therefore, it is conceivable that errors in the plane equation may affect the accuracy of feature matching by the feature matching unit 105B. Here, we will describe an example configuration of the external environment recognition device 100C in which errors in the plane equation do not affect the accuracy of feature matching by the feature matching unit 105B.

[0102] Figure 11 is a block diagram showing an example of the internal configuration of the external environment recognition device 100C according to the fourth embodiment. In the external environment recognition device 100C according to the fourth embodiment, when the plane equation contains errors, a process is performed to match the features of the target.

[0103] The external environment recognition device 100C has the same configuration as the external environment recognition device 100 shown in Figure 3, but differs in that it includes an error estimation unit 107 connected to the geometric transformation estimation unit 112A of the matching condition determination unit 104C. The error estimation unit (error estimation unit 107) estimates the error of the imaging plane estimated by the imaging plane estimation unit (imaging plane estimation unit 111). For example, the error estimation unit 107 estimates the error of the plane equation estimated by the imaging plane estimation unit 111 and outputs the estimation result to the geometric transformation estimation unit 112A. The error of the plane equation is a value determined, for example, in experiments or during the design phase. As for the error estimation result by the error estimation unit 107, a method can be considered in which the output error increases the further away the target is. Therefore, the error estimation unit 107 estimates a large error for objects farther away from the vehicle 1 and a small error for objects closer to the vehicle 1.

[0104] One way to use the errors output from the error estimation unit 107 is, for example, to prevent the geometric transformation estimation unit 112A from outputting geometric transformation parameters for plane equations where the error estimation result by the error estimation unit 107 exceeds a certain level. As a result, the feature matching unit 105B does not need to use the geometric transformation parameters estimated by the geometric transformation estimation unit 112A. This processing is performed because it is expected that the geometric transformation of the imaging plane by the feature matching unit 105B may have a negative impact on the feature matching process. If the feature matching unit 105B does not use geometric transformation parameters, it does not perform a geometric transformation of the imaging plane and matches the targets detected from images 40a and 41a as they are.

[0105] The image transformation estimation unit (geometric transformation estimation unit 112A) determines the range of error in the imaging plane to be image transformed based on the error in the imaging plane. Then, the geometric transformation estimation unit 112A uses the error estimation result of the plane equation input from the error estimation unit 107 to determine whether or not to output the estimated geometric transformation parameters for the imaging plane estimated by the imaging plane estimation unit 111. Since the range of error in the imaging plane is determined in this way, geometric transformation parameters estimated for imaging planes with large errors will not be output. If no geometric transformation parameters are output from the geometric transformation estimation unit 112A, the feature matching unit 105B matches the areas where targets are detected in images 40a and 41a without applying geometric transformation, and outputs the matching result to the position estimation unit 106.

[0106] In the external environment recognition device 100C according to the fourth embodiment described above, the geometric transformation estimation unit 112A uses the estimation result of the error of the plane equation to determine whether or not to output geometric transformation parameters. If geometric transformation parameters are not output, the feature matching unit 105B matches the area where the target was detected without performing geometric transformation. This improves the accuracy of matching the features of the target compared to matching an area geometrically transformed from an imaging plane using a plane equation with a large error.

[0107] [Modified version of the fourth embodiment] Here, we will describe another method that uses the error estimation result output by the error estimation unit 107. Figure 12 is a block diagram showing an example of the internal configuration of the geometric transformation estimation unit 112A and the feature matching unit 105B in a modified version of the external environment recognition device 100C. Here, we will focus on explaining the configuration example of the geometric transformation estimation unit 112A and the feature matching unit 105B of the external environment recognition device 100C shown in Figure 11.

[0108] It is assumed that the parameters of the plane equation estimated by the imaging plane estimation unit 111 from the imaging plane are empirically deviated from the original parameters (α, β, γ, δ) by an error (±εα, ±εβ, ±εγ, ±εδ). In this case, the geometric transformation estimation unit 112A duplicates the estimation result of the imaging plane N times for each imaging plane within the range of this error. There are various methods for duplication, but Figure 12 shows an example of the configuration and processing of the geometric transformation estimation unit 112A capable of executing one of these methods.

[0109] The image transformation estimation unit (geometric transformation estimation unit 112A) shown in Figure 12 generates multiple imaging planes based on the error of the imaging plane estimated by the error estimation unit (error estimation unit 107), and estimates multiple image transformation parameters (geometric transformation parameters) for each of the multiple imaging planes. This geometric transformation estimation unit 112A comprises an imaging plane N generation unit 112A-1 and a geometric transformation N estimation unit 112A-2.

[0110] The imaging plane N generation unit 112A-1 randomly adjusts the parameters so that they fall within the above error range, and adds them to the parameters of the plane equation estimated by the imaging plane estimation unit 111 to generate multiple imaging planes of N patterns.

[0111] The geometric transformation N estimation unit 112A-2 estimates multiple geometric transformation parameters for each of the generated N patterns of imaging planes. Therefore, the geometric transformation N estimation unit 112A-2 is capable of estimating the geometric transformation parameters for N patterns. The geometric transformation N estimation unit 112A-2 estimates the parameters, which are then output to the feature matching unit 105B.

[0112] The feature matching unit (feature matching unit 105B) performs a matching process multiple times between the features of an object detected from the first image (image 40a), which has been image-transformed using multiple image transformation parameters (geometric transformation parameters), and the features of an object detected from the second image (image 41a). The result of the matching process with the highest evaluation is output to the position estimation unit (position estimation unit 106). This feature matching unit 105B includes a feature N matching unit 105B-1 and a feature matching result selection unit 105B-2.

[0113] The Feature N Matching Unit 105B-1 receives the estimated results of N patterns of geometric transformation parameters from the Geometric Transformation N Estimation Unit 112A-2 of the Geometric Transformation Estimation Unit 112A. Based on the estimated results of the N patterns of geometric transformation parameters received, the Feature N Matching Unit 105B-1 performs N different matching processes on the target features. The results of these N matching processes are output to the Feature Matching Result Selection Unit 105B-2.

[0114] The feature matching result selection unit 105B-2 selects the matching process result with the highest evaluation during matching from the results of the N matching processes input from the feature N matching unit 105B-1, and outputs the selected matching process result to the position estimation unit 106. The position estimation unit 106 estimates the position of the target based on the results of the matching process input from the feature matching unit 105B.

[0115] The external environment recognition device 100C, a modified version of the fourth embodiment described above, includes a geometric transformation estimation unit 112A and a feature matching unit 105B, which allows the device to obtain estimation results for N patterns of geometric transformation parameters from the estimation results for N patterns of the imaging plane. Then, after N matching processes are performed based on the estimation results of the N patterns of geometric transformation parameters, the result of the matching process with the highest evaluation during the matching is selected. Therefore, even if the parameters of the plane equation estimated by the imaging plane estimation unit 111 contain errors compared to the original parameters, the position estimation unit 106 can accurately estimate the position of the object shown in the images 40a and 41a.

[0116] [Fifth Embodiment] Next, an example of the configuration and processing of the external environment recognition device 100D according to the fifth embodiment of the present invention will be described with reference to Figure 13. Figure 13 is a block diagram showing an example of the internal configuration of the external environment recognition device 100D.

[0117] The external environment recognition device 100D includes, in addition to the target detection units 101 and 102, the matching condition determination unit 104, and the position estimation unit 106, a parallax calculation unit 108 and a parallax validity verification unit 109.

[0118] The parallax calculation unit (parallax calculation unit 108) calculates the parallax from the first image (image 40a) input from the imaging unit 40 and the second image (image 41a) input from the imaging unit 41, and compares the positions of the targets reflected in the first image (image 40a) and the second image (image 41a). That is, the parallax calculation unit 108 calculates the parallax of images 40a and 41a without using the plane equation estimated by the imaging plane estimation unit 111 based on the target detection result of the target detection unit 101. The parallax calculation unit 108 may use a known method for measuring parallax as the method for calculating parallax. By calculating the parallax, the parallax calculation unit 108 outputs the result of comparing the positions of the same targets reflected in images 40a and 41a to the parallax validity verification unit 109 as a second comparison result.

[0119] The parallax validity verification unit 109 is a modified version of the feature matching unit 105 according to the first embodiment. The feature matching unit (parallax validity verification unit 109) compares a first matching result obtained by comparing the features of a target detected from a first image (image 40a) that has been image-transformed using image transformation parameters (geometric transformation parameters) with the features of a target detected from a second image (image 41a), and compares this result with a second matching result obtained by parallax. Based on the validity of the comparison result, it outputs either the first matching result or the second matching result to the position estimation unit (position estimation unit 106). That is, the parallax validity verification unit 109 compares a target in an image region that has been geometrically transformed using geometric transformation parameters estimated by the geometric transformation estimation unit 112 using the plane equation estimated by the imaging plane estimation unit 111, with a target in the image region of image 41a, and obtains a first matching result. The parallax validity verification unit 109 then compares the first matching result with the second matching result and verifies the validity of the second matching result calculated by the parallax calculation unit 108. The validity of the matching result by the parallax validity verification unit 109 is expected to be determined by a method such as the degree of matching in template matching.

[0120] Furthermore, within the range where the plane equation is estimated by the imaging plane estimation unit 111, it is assumed that the accuracy of matching the geometrically transformed target based on the geometric transformation parameters estimated by the geometric transformation estimation unit 112 will be high. However, in the method of determining the corresponding position of the target using the parallax calculated by the disparity calculation unit 108 from images 40a and 41a without the estimation of the plane equation by the imaging plane estimation unit 111, conditional branching due to the plane equation does not occur. For this reason, there is an advantage in that the disparity calculation process can be sped up by implementing the disparity calculation unit 108 in hardware.

[0121] Therefore, the parallax validity verification unit 109 limits the first matching result obtained using the plane equation and geometric transformation parameters, and uses this limited first matching result in the verification process to determine whether the second matching result obtained by the parallax calculation unit 108 is correct or not. This processing is performed because if the parallax validity verification unit 109 were to obtain the first matching result obtained using the plane equation and geometric transformation parameters for all targets commonly included in images 40a and 41a, the processing load on the parallax validity verification unit 109 would be high.

[0122] The parallax validity verification unit 109 compares the limited first matching result with the second matching result. If the validity of the first matching result using geometric transformation is high, it adopts the first matching result and outputs the first matching result to the position estimation unit 106. For example, if the target is detected at a position close to the vehicle 1, the parallax will be large, and the validity of the second matching result, which is matched by parallax, will be low. On the other hand, if the target is close to the vehicle 1, the imaging plane is accurately estimated, and the validity of the first matching result using geometric transformation will be high. Therefore, the position estimation unit 106 estimates the position of the target using the first matching result, similar to the external environment recognition device 100 according to the first embodiment described above.

[0123] The parallax validity verification unit 109 compares the first matching result with the second matching result, and if the validity of the second matching result is low, it calculates the parallax of images 40a and 41a and outputs the second matching result, which has been matched, to the position estimation unit 106. For example, if the target is detected at a position far from the vehicle 1, the parallax will be small, and the validity of the second matching result, which has been matched by parallax, will be high. On the other hand, because this target is at a position far from the vehicle 1, the imaging plane is estimated inaccurately, so the validity of the first matching result by geometric transformation is low. For this reason, the position estimation unit 106 estimates the position of the target using the second matching result. If the validity of the second matching result is low, there may be targets where the difference between the first and second matching results is large. In this case, the parallax validity verification unit 109 may compare the first and second matching results again for the image region in which this target is shown.

[0124] In the external environment recognition device 100D according to the fifth embodiment described above, the validity of the matching result is verified by comparing the image region of image 40a and the image region of image 41a, which have been geometrically transformed using geometric transformation parameters output from the matching condition determination unit 104, with the second matching result of images 40a and 41a using the parallax calculated by the parallax calculation unit 108. The parallax validity verification unit 109 then determines whether to output the result to the position estimation unit 106 as either the first matching result or the second matching result, depending on the validity of the matching result. If the validity obtained by comparing the first and second matching results is low, the position estimation unit 106 can estimate the position of the target using the second matching result. Therefore, depending on whether the target is close to or far from the vehicle 1, either the first matching result or the second matching result will be output to the position estimation unit 106. As a result, the position estimation unit 106 can quickly estimate the position of targets at various locations.

[0125] [Differentiation] In the embodiments described above, the present invention was explained as an external environment recognition device mounted on a vehicle 1. However, this external environment recognition device may also be used, for example, in a self-propelled robot equipped with multiple cameras, or in an infrastructure monitoring system that monitors the premises using images captured by multiple cameras. In an infrastructure monitoring system, for example, a configuration in which multiple cameras are installed in one location is envisioned.

[0126] In addition to the embodiments described above, it is also possible to configure an external environment recognition device that has the function of directly calculating distance from the results of machine learning. In this case, if the external environment recognition device can calculate the distance to a target using machine learning, the distance can be calculated at the time the target is detected from the image, thereby reducing the computational load on the external environment recognition device.

[0127] In addition to the embodiments described above, the external environment recognition device may also perform parallax matching, as used in conventional stereo cameras. In this case, the external environment recognition device can refer to the result that is estimated to have the best matching result from the results of matching multiple images, and can also estimate the positional relationship of each target using pre-prepared map information.

[0128] It should be noted that the present invention is not limited to the embodiments described above, and various other applications and modifications can be taken as long as they do not deviate from the gist of the present invention as described in the claims. For example, the embodiments described above are detailed and specific explanations of the configuration of the apparatus and system in order to clearly illustrate the present invention, and are not necessarily limited to having all the configurations described. Furthermore, it is possible to replace some of the configurations of the embodiments described here with the configurations of other embodiments, and it is also possible to add the configurations of other embodiments to the configuration of one embodiment. In addition, it is possible to add, delete, or replace some of the configurations of each embodiment with other configurations. Furthermore, the control lines and information lines shown are those deemed necessary for explanatory purposes, and not all control lines and information lines are necessarily shown in the actual product. In reality, it is safe to assume that almost all components are interconnected. [Explanation of Symbols]

[0129] 1,2...Vehicle, 40,41...Imaging unit, 40a,41a...Image, 100...External environment recognition device, 101,102...Target detection unit, 103...Camera parameters, 104...Matching condition determination unit, 105...Feature matching unit, 106...Position estimation unit, 111...Imaging plane estimation unit, 112...Geometric transformation estimation unit, 113...Imaging plane estimation unit, 114...Geometric transformation estimation unit

Claims

1. A first target detection unit detects a target based on a first image acquired from at least one of several imaging units, the imaging field of view of which overlaps with at least a portion of the external environment. A second target detection unit detects a target based on a second image acquired from a second imaging unit among the multiple imaging units, Based on the detection result of the first target detection unit, an imaging plane estimation unit estimates the imaging plane in the first image, which is one of the imaging planes that constitute each target and is expressed by a plane equation, within the image region in which the target was captured. An image transformation estimation unit estimates image transformation parameters that match the imaging plane in the second image from among the imaging planes represented by the plane equation, based on the information of the imaging plane estimated by the imaging plane estimation unit and the imaging parameters of the plurality of imaging units. A feature matching unit compares the features of the target detected from the first image, which has been image-converted using the image conversion parameters of the image conversion estimation unit, with the features of the target detected from the second image. The system includes a position estimation unit that estimates the three-dimensional position of the target based on the matching results of the feature matching unit. External world recognition device.

2. The system comprises a plurality of imaging plane estimation units and a plurality of image conversion estimation units, each provided for the target detected from the first image and the target detected from the second image, The feature matching unit obtains a comparison result between the first image converted using the image conversion parameters estimated by one of the image conversion estimation units and the second image, and obtains a comparison result between the second image converted using the image conversion parameters estimated by the other image conversion estimation unit and the first image, and compares each of the comparison results to assign a high degree of confidence to the comparison result when the target detected in the first image and the target detected in the second image are matched at the same three-dimensional position. The external environment recognition device according to claim 1.

3. The system includes a conversion selection unit that selects from among a plurality of targets detected from the first and second images that are close to the vehicle, and selects the image conversion parameters estimated by the image conversion estimation unit for the selected target. The feature matching unit performs image conversion on the target detected from the first image using the image conversion parameters selected by the conversion selection unit. The external environment recognition device according to claim 2.

4. The conversion selection unit prioritizes selecting targets that have attributes that significantly affect the vehicle's movement. The external environment recognition device according to claim 3.

5. The system includes an error estimation unit that estimates the error of the imaging plane estimated by the imaging plane estimation unit, The image conversion estimation unit determines the range of error in the imaging plane to be converted based on the error in the imaging plane. The external environment recognition device according to claim 1.

6. The image conversion estimation unit generates a plurality of imaging planes based on the error of the imaging plane estimated by the error estimation unit, and estimates a plurality of image conversion parameters for each of the plurality of imaging planes. The feature matching unit performs a matching process multiple times between the features detected from the first image, which has been image-transformed using a plurality of image transformation parameters, and the features detected from the second image, and outputs the result of the matching process with the highest evaluation to the position estimation unit. The external environment recognition device according to claim 5.

7. The system includes a parallax calculation unit that calculates parallax from the first image and the second image and compares the positions of targets reflected in the first image and the second image. The feature matching unit compares a first matching result obtained by matching the features of the target detected from the first image, which has been image-converted using the image conversion parameters, with the features of the target detected from the second image, with a second matching result obtained by matching using the parallax, and outputs either the first matching result or the second matching result to the position estimation unit based on the validity of the comparison result. The external environment recognition device according to claim 1.

8. A process in which a first target detection unit detects a target based on a first image acquired from at least one of several imaging units whose imaging fields of view to the outside world overlap in at least a portion thereof, Based on the second image acquired from the second imaging unit among the multiple imaging units, a second target detection unit detects a target. The imaging plane estimation unit performs a process of estimating the imaging plane in the first image, based on the target detection result by the first target detection unit, among the imaging planes that constitute each target and are expressed by a plane equation, in the image region in which the target was captured. Based on the information of the imaging plane estimated by the imaging plane estimation unit and the imaging parameters of the plurality of imaging units, the image transformation estimation unit performs a process to estimate image transformation parameters that match the imaging plane in the second image from among the imaging planes represented by the plane equation, A process of comparing the features of the target detected from the first image, which has been image-converted using the image conversion parameters of the image conversion estimation unit, with the features of the target detected from the second image, The process includes estimating the three-dimensional position of the target based on the results of the matching process. How to recognize the external world.