Method for measuring environmental terrain

By using a monocular camera and polar line correction technology with wide field of view lens on vehicles, the problem of difficulty in obtaining the environmental terrain depth of a monocular camera is solved, real-time dense 3D point cloud reconstruction is realized, and accurate information acquisition of the autonomous driving system is supported.

CN115023736BActive Publication Date: 2025-07-18CONNAUGHT ELECTRONICS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080095163.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-13
Filing Date
2020-12-08
Publication Date
2025-07-18
Estimated Expiration
2040-12-08

AI Technical Summary

Technical Problem

The prior art is difficult to obtain efficient and accurate environmental terrain depth information from monocular cameras, especially when used on vehicles, and it is impossible to form dense 3D point clouds in real time to assist in autonomous driving.

Method used

By using a monocular camera with a wide field of view lens, combined with a motion stereo module and polar line correction technology, the camera's motion and vehicle odometer information are used to map images to a non-planar surface for depth calculations, forming a dense 3D point cloud.

Benefits of technology

It realizes the acquisition of dense and accurate environmental terrain depth information from a monocular camera, can reconstruct the local terrain around the vehicle in real time, and supports the acquisition of accurate information from the autonomous driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115023736B_ABST
    Figure CN115023736B_ABST
Patent Text Reader

Abstract

A method for forming a point cloud corresponding to a terrain of an imaging environment, comprising: acquiring a first image of the environment with a camera having a WFOV lens mounted on a vehicle; changing the camera pose by an adjustment greater than a threshold; and acquiring a second image of the environment with the camera in the changed pose. The images are mapped onto corresponding surfaces to form corresponding mapped images defined by the same non-planar geometry. One of the first or second mapped images is divided into pixel blocks; and for each pixel block, a depth map is formed by performing a search on the other mapped image to evaluate a position difference of the position of the pixel block in each mapped image. When the vehicle moves in the environment, the depth map is converted into a part of a point cloud corresponding to a local terrain of the vehicle's surrounding environment; and the point cloud is scaled according to the adjustment of the camera pose.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a method for measuring the topography of an environment imaged by a camera. More specifically, the present invention relates to a method for measuring the topography of an environment that is imaged using dense depth measurements evaluated from motion stereo images of the environment. Background Art

[0002] A camera is a device that produces an image of a scene or environment. When two cameras produce images of the same scene from different positions, the different images can be compared to determine the depth of various parts of the scene, which is a measure of the relative distance to a plane defined by the two camera positions. Under certain assumptions and / or certain information, the relative depth can be calibrated to an absolute distance measurement. This is the depth principle of parallax imaging. Depth measurements can be used to approximate the topography of the imaged environment.

[0003] Generally, depth from parallax imaging requires an N-eye system, where N > 1. Typically, a binocular system has two cameras that produce a synchronized pair of images of the scene. Features in one image of the pair can be matched to corresponding features in the other image. Features can include different imaging elements, such as corners or regions of similar pixels, blobs, but features can also include any given pixel of the image. Then, the positional differences of the matched features between the images can be used to calculate the parallax. Based on the difference in features and the known separation of the cameras in the binocular system, the depth of the features can be evaluated. Generally, the images acquired by the binocular system are mapped onto a surface to assist subsequent image processing or to make the acquired images more suitable for viewing.

[0004] Gehrig's "Large Field of View Stereo for Automotive Applications" in OmniVis, Volume 1, 2005, relates to cameras placed on the left and right of an automotive rearview mirror and describes options for analyzing stereo vision with a large field of view and performing object detection.

[0005] The "Three Dimensional Measurement Using Fisheye Stereo Vision, Advances in Theory and Applications of Stereo Vision" in Chapter 8 of the book "Three Dimensional Measurement Using Fisheye Stereo Vision, Advances in Theory and Applications of Stereo Vision" published by Yamaguchi in 2011 discloses mapping fisheye images onto a plane and matching features, and concludes that fisheye stereo vision allows for the measurement of 3D objects in a relatively large space.

[0006] "Omnidirectional Stereo Vision" by Zhu in IEEE ICAR, 2001 is concerned with the configuration of omnidirectional stereo imaging and conducts numerical analysis on omnidirectional representation, epipolar geometry, and depth error characteristics.

[0007] "Direct Fisheye Stereo Correspondence Using Enhanced Unified Camera Model and Semi-Global Matching Algorithm" by Bogdan et al. in ICARCV 2016 proposes a fisheye camera model that projects lines onto conic curves and describes a matching algorithm for a fisheye stereo system to compute dense direct stereo correspondences without rectifying fisheye images.

[0008] "Binocular Spherical Stereo" by Li in the IEEE journal on Intelligent Transportation Systems 9, 589, 2008 is concerned with binocular fisheye stereo images and describes the conversion into spherical images and the use of latitude-longitude representation to accelerate feature point matching.

[0009] "Fish-Eye-Stereo Calibration and Epipolar Rectification" by Abraham et al. in the journal Photogrammetry and Remote Sensing 59, 278, 2005 is concerned with the calibration and epipolar rectification of fisheye stereo images and discusses the generation of epipolar images.

[0010] "On the Accuracy of Dense Fisheye Stereo" by Schnedier et al. in IEEE Robotics and Automation, 1, 227, 2016 analyzed the epipolar rectification model of a fisheye stereo camera and discussed the related accuracy.

[0011] "Omnidirectional stereo vision using fisheye lenses" proposed by Drulea et al. in ICCP IEEE in 2014 involves an omnidirectional stereo system and the segmentation of fisheye lens images into rectified images. Stereo matching algorithms are applied to each pair of rectified images to form a point cloud.

[0012] The object of the present invention is to overcome at least some of the limitations of this related work. Summary of the Invention

[0013] The present invention is defined by the independent claims.

[0014] Embodiments of the present invention provide a method for recovering dense and accurate depth information from images acquired by a camera with a wide field of view lens. This enables images from a monocular camera on a vehicle to form part of a point cloud corresponding to the local topography of the environment around the vehicle as the vehicle moves in the environment.

[0015] The dependent claims provide further optional features. Brief Description of the Drawings

[0016] Embodiments of the present invention will now be described by way of example with reference to the accompanying drawings, in which:

[0017] Figure 1 A vehicle equipped with cameras is schematically shown, each camera being capable of operating according to the present invention;

[0018] Figure 2 A bowl surface is shown onto which the stitched images from multiple cameras can be mapped;

[0019] Figure 3 An unrectified stereo camera setup is shown;

[0020] Figure 4 The disparity in the rectified stereo camera setup is shown;

[0021] Figure 5a and 5b A spherical epipolar rectification surface for a moving camera is shown, Figure 5a The situation when the camera observes along the motion axis is shown, Figure 5bShows the situation when the camera observes perpendicular to the axis of motion;

[0022] Figure 6a and 6b Shows an upright cylindrical epipolar correction surface for moving the camera, Figure 6a Shows the situation when the camera observes along the axis of motion, Figure 6b Shows the situation when the camera observes perpendicular to the axis of motion;

[0023] Figure 7 Shows a cylindrical correction surface for moving the camera when the camera observes along the axis of motion and the cylinder is concentric with the axis of motion;

[0024] Figure 8a and 8b Shows spherical and cylindrical correction surfaces with baseline alignment respectively;

[0025] Figure 9 Shows a spherical coordinate system for epipolar correction;

[0026] Figure 10 Shows a conical correction surface;

[0027] Figure 11 Shows a correction surface including multiple planes;

[0028] Figure 12a and 12b Shows the corresponding multi-part correction surfaces according to embodiments of the present invention respectively;

[0029] Figure 13a and 13b Shows the geometric relationship for calculating depth using a planar correction surface; and

[0030] Figure 14 Shows the geometric relationship for calculating depth using a spherical correction surface. Detailed Description

[0031] For many tasks involving driving a vehicle, obtaining information about the local environment is important for safely completing the task. For example, when parking, it is beneficial to display a real-time image of the vehicle's surroundings to the driver.

[0032] The driver of the vehicle does not need to be human, as the vehicle can be self-driving, i.e., an autonomous vehicle. In this case, the accuracy of the information obtained is particularly important for identifying objects and avoiding misleading the vehicle's driving system with the information obtained. The driver can also be a combination of a human and one or more automated systems for assisted driving.

[0033] The sensitivity of the camera used in the present invention need not be limited to any particular wavelength range, but most commonly, it will be used in conjunction with a camera sensitive to visible light. The camera is typically in the form of a camera module, including a housing for the lens and sensor, the lens for focusing light onto the sensor. The camera module may also have electronics for powering the sensor and enabling communication therewith, as well as processing electronics that may process the image. Such processing may be low-level image signal processing, such as gain control, exposure control, white balance, denoising, etc., and / or it may include more powerful processing, such as for computer vision.

[0034] When imaging the environment around a vehicle, a single camera typically does not have a sufficient field of view to acquire all the data needed. One way to address this issue is to use multiple cameras. In Figure 1 FIG., a vehicle 100 is shown having four cameras 101, 102, 103, 104 located around the periphery of the vehicle. One edge of each camera's field of view is marked with a dashed line 101a, 102a, 103a, 104a. This configuration of cameras results in the fields of view overlapping in regions 101b, 102b, 103b, 104b. The illustrated configuration is merely exemplary. The teachings disclosed are equally applicable to other camera configurations.

[0035] The illustrated field of view is approximately 180 degrees. A wide field of view is typically achieved by a camera with a wide-field-of-view lens, such as a fish-eye lens. Fish-eye lenses are preferred because they are typically cylindrically symmetric. In other applications of the present invention, the field of view may be less than or greater than 180 degrees. Although fish-eye lenses are preferred, any other lens that provides a wide field of view may be used. In this document, a wide field of view is a lens with a field of view of more than 100 degrees, preferably more than 150 degrees, and more preferably more than 170 degrees. Generally, cameras with such a wide field of view result in imaging artifacts and distortions in the acquired images.

[0036] The lens focuses light onto a sensor, which is typically rectangular. Thus, the acquired data is affected by a combination of the lens's artifacts and distortions and the limited sensitive surface effects of the sensor. As a result, the acquired image is a distorted representation of the imaging scene. The acquired distorted image can be at least partially corrected by a process that includes mapping the acquired data onto another surface. Mapping onto certain surfaces makes subsequent processing techniques more accurate or simpler. Several particularly advantageous surfaces will be described in more detail later.

[0037] In as Figure 1In the case of a vehicle with multiple cameras as shown, it is preferable to display a single image of the local environment rather than multiple images from multiple cameras. Therefore, the images from the cameras are combined. There are various known methods to perform such combination. For example, the images can be stitched together by identifying features in the overlapping regions. The positions of the identified features in the images can then be used to map one image to another. Many other methods for stitching images with overlapping regions are known to those skilled in the art. For example, German Patent Application No. 102019131971.4 (Reference No.: 2018PF02667) titled "An image processing module" filed on November 26, 2019, and German Patent Application No. DE102019126814.1 (Reference No.: 2019PF00721) titled "An electronic control unit" filed on October 7, 2019, give examples of electronic control units that provide a panoramic view of the vehicle to assist the driver, especially when parking.

[0038] As Figure 1 shown, if the camera configuration can provide images from all directions around the vehicle, the stitched image provides a sufficient view of information to generate views in other directions.

[0039] Since the stitched image consists of planar images stitched together, the stitched image itself appears planar. By mapping the stitched image onto a non-flat surface, a better display of the stitched image can be achieved. For example, the Figure 2 shown bowl-shaped surface 300 can be used. For this surface, nearby objects are mapped onto the flat floor, while distant objects are mapped onto the sidewall of the bowl. Since vehicles often travel on flat land with objects such as distant trees and buildings, such mapping usually produces a better and more accurate display of the stitched image than a flat mapping.

[0040] The optimal surface would be the surface corresponding to the terrain of the scene imaged by the cameras. The present invention provides a method for generating an approximation of the surrounding terrain by calculating the depth of partial images. This is implemented in real time without having to incur the cost of binocular stereo cameras and / or cameras using non-standard lenses.

[0041] Motion Stereo Module

[0042] The present invention relates to a method for processing motion stereo images from a monocular camera using a motion stereo module. The motion stereo module recovers dense and accurate depth measurements from a pair of images of a scene, thereby allowing the terrain of the imaged scene to be reconstructed. By using the method described below, the motion stereo module can operate fast enough to be completed in real time. In other words, the processing is fast enough such that the image display based on the live camera feed is not adversely affected.

[0043] Depth from disparity imaging allows depth measurements to be extracted from a pair of images of a scene. Typically, depth from disparity imaging uses a pair of images obtained from a stereo camera, which includes two cameras placed close to each other or integrated into the same device, such that a synchronized pair of images can be obtained directly. However, the present invention is based on a moving monocular camera. The pair of images from the monocular camera is obtained by capturing one image with the monocular camera and then adjusting the pose of the monocular camera, i.e., moving the monocular camera, and acquiring another image.

[0044] The resulting depth measurements can be formed as part of a 3D point cloud or a similar 3D reconstruction that approximates the terrain of the environment imaged by the camera. The 3D reconstruction enables better assessment and measurement of static features. If the camera is mounted on a vehicle, examples of common static features include curbs, ramps, surface irregularities, or larger objects such as utility poles, trees, obstacles, walls, and parked vehicles, thereby providing valuable information. Thus, the present invention makes it easier to detect such objects.

[0045] Preferably, there is a dense number of depth measurements, as this allows for a higher resolution 3D reconstruction. Known dense reconstruction techniques require a stereo camera. Relative to a monocular camera, a stereo camera requires complex hardware and frame synchronization, but due to the fixed separation between the two cameras, stereo cameras have the advantage that they provide sufficient binocular disparity for depth estimation and 3D reconstruction, independent of camera motion.

[0046] Stereo and Structure from Motion

[0047] Forming a 3D reconstruction using mobile monocular camera technology is typically attempted using classical structure from motion techniques. Such techniques only produce a sparse set of depth measurements. This sparse set results in a limited set of points in the 3D reconstruction, making it less representative of the local terrain.

[0048] For this method, image pairs are generated using knowledge of the camera motion between the captured images. Before the depth of the disparity processing begins, correction of the image pairs is usually also required. Many surfaces can be corrected. Epipolar correction on spherical or cylindrical surfaces has particular advantages because the available field of view is improved while the source image resolution is distributed advantageously within the corrected image. This is particularly important for cameras that produce highly distorted shapes, such as cameras with fish-eye lenses.

[0049] Typically, 3D reconstruction methods that utilize binocular disparity and epipolar geometry principles take as input at least two images of the same scene captured from different camera poses. The precise movement of the camera (changes in position and orientation) can be determined dynamically using computer vision techniques or inertial sensors. When such a camera is mounted on a vehicle, the motion can be evaluated at least in part based on on-vehicle odometer information. This information is typically available on the vehicle CAN or FlexRay bus of modern vehicles.

[0050] Thus, to obtain depth measurements, first an image of the scene is acquired with the camera; the camera is moved and another image of the scene is acquired. The resulting images are corrected by mapping them onto a common plane or a suitable surface defined by a specific epipolar geometry. The corrected images are processed to evaluate the depth from the disparity. For example, known matching algorithms calculate the disparity for all pixels between the corrected images. The disparity information is then converted into depth or directly into a 3D point cloud. In some embodiments, instead of matching each pixel, pixel blocks are matched. The pixel blocks can overlap such that a pixel is included in multiple pixel blocks. These blocks do not have to be rectangular and can be of any size that allows for matching.

[0051] When a monocular camera is attached to a vehicle, the camera motion can be calculated with reasonable accuracy at low accelerations from the on-vehicle odometer sensor in two degrees of freedom (two parameters). Such an odometer sensor provides the longitudinal speed and the yaw angular velocity or the steering angle of the vehicle. If the vehicle motion is truly planar, this information is sufficient. However, in reality, due to the dynamic response of the suspension to road unevenness, acceleration, deceleration, and turning, the vehicle motion is more complex. This complexity results in instantaneous changes in pitch, roll, and height. In addition, mechanical measurements and their transmission on the system bus are affected by delays and are not synchronized with the camera frames by default. German Patent Application No. 102019114404.3, filed on May 29, 2019, titled "Image acquisition system" (reference number: 2018PF02113(SIE0883)) discloses techniques for handling these variations in the attitude of the vehicle relative to the road surface.

[0052] However, the dynamic motion of a vehicle and a camera can be fully characterized by six degrees of freedom (6-DoF) with at least three position parameters (X, Y, Z) and three rotation parameters (yaw, pitch, and roll). When the relative camera motion between two images can be estimated in 6-DoF, the motion stereo module produces the most reliable and accurate results. The 6-DoF also provides scaling information, i.e., the length of the translation vector used to correctly scale the 3D reconstruction (e.g., the point cloud). The rectified images and the disparity map are invariant to the estimated scaling. Known related relative pose estimation problems are solved using known techniques.

[0053] Epipolar Rectification

[0054] By observing Figure 3 , and considering the cameras C1 and C2 as pinhole cameras and the virtual image planes in front of the optical centers of the cameras C1 and C2, the following useful terms can be defined. The epipole is the intersection of the baseline extending between the optical centers and the image plane. The epipole can be considered as the image of the optical center of the other camera in one camera. The epipolar plane is the plane defined by the 3D point Z1 and the optical center. The epipolar line is the line where the epipolar plane intersects the image plane. It is the image of a bundle of light rays passing through the optical center in one camera and is an image point in the other camera. All epipolar lines intersect at the epipole.

[0055] For motion stereo, a known stereo matching module (e.g., the Renesas STV hardware accelerator) can be used to perform depth analysis, which takes the rectified images as input, meaning two images where the epipolar lines have been mapped to horizontal scan lines and have the same vertical offset in both images. In principle, the geometry of any scene can be reconstructed from two or more images captured from different camera poses by knowing the feature point correspondences between these images.

[0056] Given such correspondences, e.g., in the form of a dense optical flow field, the 3D point cloud or 3D reconstruction can be calculated through a mathematical process called triangulation, where the light rays are back-projected from the camera viewpoints through their respective image points and intersect in 3D space by minimizing an error metric. However, unlike optical flow, for epipolar (i.e., stereo) rectification, the correspondence problem is reduced to a 1D search along the conjugate epipolar lines, and triangulation is reduced to a simple formula for solving the ratio between similar triangles. The 1D search is efficiently performed by a stereo matching algorithm, which typically applies techniques to aggregate information from multiple searches and produces a robust 1D correspondence in the form of a disparity map. In this way, most of the computational burden of triangulation is transferred to epipolar rectification. An example of this effect can be seen in Figure 4 where the rectification of the left image produces horizontal epipolar lines, in contrast to Figure 3 , Figure 3The epipolar line in it is the diagonal of the image plane.

[0057] In the case where one image provides a reference pose, knowledge of the intrinsic calibration parameters of each captured image and the relative pose of the camera allows calibration of the image. The reference pose can be given relative to an external reference frame or can be arbitrarily set to zero, i.e., the camera origin is at (0,0,0) and the axes are defined by the standard basis vectors (identity rotation matrix). The intrinsic parameters can always be assumed to be known and constant. However, this may not always be a safe assumption, e.g., due to variations caused by thermal expansion and contraction of the materials including the camera. An alternative is to compensate for any variations by characterizing and taking into account such variations, or by using an online intrinsic calibration method to periodically update the intrinsic calibration information stored in the system.

[0058] In motion stereo, the relative pose varies according to the motion of the vehicle and can be dynamically estimated for each stereo image pair. The relative pose is completely determined by at least 6 parameters (3 position parameters and 3 rotation parameters), i.e., with 6 degrees of freedom, or can be "scaled" by at least 5 parameters (3 rotation parameters and 2 position parameters), i.e., with 5 degrees of freedom missing the 6th degree of freedom (scale). In the latter case, the two position parameters represent the direction of the translation vector, e.g., in projected coordinates or spherical coordinates of unit length. The missing "scale" or so-called "scale ambiguity" is a typical obstacle in monocular computer vision, which arises from the simple fact that the scene geometry and the camera translation vector can be scaled together without affecting the positions of the feature points and their correspondences in the captured image, and thus, conversely, the scale usually cannot be recovered solely from such correspondences. Note that estimating the scale or absolute scale is not necessary for epipolar rectification. In other words, since 5 degrees of freedom (3 rotation parameters and 2 position parameters) provide sufficient information, the epipolar rectified images are invariant to the estimated scale. However, scaling allows obtaining correct depth measurements, enabling a more realistic 3D reconstruction (i.e., correct 3D point cloud coordinates).

[0059] Epipolar rectification can be performed by directly mapping the images to a suitable plane or surface and resampling them in a way that satisfies two simple geometric constraints - in the two rectified images, conjugate epipolar lines or curves are mapped along the horizontal scan lines and with the same vertical offset. The simplest form of epipolar rectification uses two coplanar surfaces (image planes) and a Cartesian sampling grid oriented parallel to the baseline (camera translation vector). For cameras using fish-eye lenses, due to the mathematical limitations of perspective projection, this method severely limits the obtained field of view and causes quality loss in the rectified images due to wide-angle perspective effects (such as pixel "stretching").

[0060] Even when multiple planes are used to increase the reconstructed field of view, these still cannot reach regions close to the extended focus because this would require an image plane of near-infinite size. This is a particular problem for forward and backward cameras on a vehicle, where the extended focus is approximately near the center of the fisheye image when the vehicle is moving on a straight path.

[0061] The extended focus in side-view or rear-view mirror cameras is typically located in an image region of less interest. However, even for these cameras, the planar surface imposes a limitation on the reconstructed field of view.

[0062] Generalized Epipolar Rectification

[0063] To mitigate the above problems and enhance the reconstructed field of view in the horizontal direction (HFOV), vertical direction (VFOV), or both, non-planar mapping surfaces such as spheres, cylinders, or polynomial surfaces can be effectively used for epipolar rectification.

[0064] As an example, consider Figure 5a , where two points C1 and C2 representing the camera viewpoints define a translation vector, i.e., the "baseline". The intersections of the baseline with each mapping surface are the poles E1 and E2, and any object point P in the scene, along with the two points C1 and C2, define an epipolar plane. There is an infinite number of epipolar planes, rotating around the baseline. This family of planes is called an epipolar pencil. In this case, the intersection of any one of these planes with the mapping surface defines an "epipolar line" or curve, which is circular for a spherical mapping surface and elliptical for an upright cylindrical mapping surface. Any object point P is mapped to points Q1 and Q2, which belong to the same epipolar plane and conjugate epipolar curve, and they are mapped to a horizontal scan line and have the same vertical offset in the two rectified images.

[0065] The mapping of fisheye image pixels along the epipolar line or curve to pixels on the horizontal scan line in the rectified image can be achieved by "sampling" rays along each epipolar line or curve through their respective viewpoints and then using the intrinsic calibration and relative pose information to trace each ray back to the fisheye image. In this way, each pixel in the rectified image can be traced back to an image point in the fisheye image. The intensity value of the nearest source pixel can be obtained directly, or by using a reconstruction and / or anti-aliasing filter that takes into account adjacent pixel values, such as a bilinear filter (i.e., bilinear interpolation can be used).

[0066] This process can be performed with very high computational efficiency by constructing a sparse lookup table for a subset of pixels in the rectified image (e.g., every 16 pixels in the horizontal and vertical directions), which stores the corresponding fisheye image coordinates of the pixel with fractional precision, e.g., having 12 integer bits and 4 fractional bits, i.e., 16 bits per coordinate or 32 bits per pixel, to save memory bandwidth and improve runtime performance. Then a software- or hardware-accelerated "renderer" can be used to very quickly "render" the rectified image using these lookup tables by interpolating the coordinates of the missing pixels. This reduces the number of rays that need to be computed for each rectified image, e.g., by a factor of 256 when using 1:16 subsampling in both directions.

[0067] To save memory bandwidth, the lookup table can also be compressed by storing deltas instead of absolute image coordinates (e.g., 4-bit integer deltas with 4 fractional bits, i.e., 8 bits per coordinate or 16 bits per pixel). In this case, the initial absolute coordinates are stored as a seed for the whole table or for each row so that the absolute coordinates can be incrementally recovered from the stored deltas during rendering. Additionally, the lookup table can be subdivided into smaller regions, where each region is assigned an offset value that will be applied to all deltas within that region during rendering.

[0068] For a spherical surface, rays can be sampled at angular intervals along each meridian curve, which for a spherical surface is a circle. The angular intervals that can be rotated around a baseline define meridian planes such that each plane is mapped to a discrete horizontal scan line in the rectified image. The mathematical functions that map the horizontal pixel coordinate x of the rectified image to the polar angle θ and the vertical pixel coordinate y (number of scan lines) to the azimuth can be summarized as: θ = f(x) and or can be linear in the simplest case: θ = s x x and where s x and s y are constant scale factors. The polar angle and the azimuth define the ray directions for sampling the fisheye image pixels. More complex functions can be used to distribute the source image resolution over the rectified image in a beneficial way.

[0069] Note that the concept of mapping a surface, such as a sphere ( Figure 5a and 5b ) or a cylinder ( Figure 6a and 6bIt is only used as a modeling aid because the process can be abstracted mathematically and procedurally. The intersection of the epipolar plane with the mapping surface simply defines a path along which rays are sampled to achieve the desired alignment of epipolar lines or curves in the rectified image. The rays along this path can also be sampled at intervals of "arc length" rather than angular intervals. By modifying the shape of the 3D surface (e.g., modifying it to an oblate or prolate ellipsoid) or the mapping function or both, the two representations can be made equivalent.

[0070] As Figure 6a and 6b shown, for an upright cylindrical surface, rays can be sampled at angular intervals along each epipolar curve, in which case the epipolar curve is an ellipse, but since the vertical dimension is linear, the epipolar plane is defined in Cartesian coordinate intervals similar to a planar surface. Similarly, for a horizontal cylindrical surface, since the horizontal dimension is linear and the epipolar plane is defined in angular intervals rotating around the baseline, rays can be sampled in Cartesian coordinate intervals along each epipolar line.

[0071] Turning to Figure 5a , 5b , 6a, 6b and 7; the planes represented by rectangles demonstrate how a planar surface with a finite field of view can be used to replace a spherical or cylindrical surface for epipolar rectification. For vehicle-mounted cameras facing forward and backward, these planes can be oriented perpendicular to the baseline, which is also the direction of camera motion. In this case, the epipolar lines radiate radially from the poles E1 and E2, and a polar coordinate sampling grid can be used to map them along the horizontal scan lines in the rectified image. This configuration allows the central part of the image to be rectified in a manner similar to spherical modeling. When the image plane approaches infinity, the field of view is limited to a theoretical maximum of 180 degrees.

[0072] In general, the poles can also be in the image (especially for forward motion when acquiring images from the forward or backward direction) or near the image. In this case, a linear transformation cannot be used. A radial transformation can be used here. However, if the epipolar lines are outside the image when acquiring images from the left or right direction, as in the case of forward motion, the transformation is similar to a linear transformation.

[0073] Fisheye Correction

[0074] Calculating the grid to correct for fisheye distortion. For each grid point x, y in the undistorted image, the grid point x′, y′ in the distorted space can be defined as follows:

[0075] a = (x - c x )

[0076] b = (y - c y )

[0077]

[0078]

[0079] r′ = k4θ 4 + k3θ 3 + k2θ 2 + k1θ

[0080]

[0081]

[0082] The value of the focal length f should be determined as follows. At the distortion center c x , c y the distortion is minimal. The pixels in the output image should not be compressed (or magnified). Near the distortion center, the k1 parameter dominates. If f = k1, then for a small angle θ around the distortion center, the pixels in the undistorted image will be in the same proportion because r = f·tan(θ) ≈ fθ and r′ ≈ k1θ.

[0083] However, the camera may not allow access to its distortion parameters, which in this case are k1…k4. In such a case, the focal length can be determined by using a small-angle artificial light with minimal distortion. The resulting pixel positions can be used to estimate the focal length. When calculating the grid for each grid point (x, y) in the direction of the selected virtual image plane, the following ray is created: v = (x, -y, -f·s), where s is the scale factor.

[0084] This ray is rotated by the rotation matrix of the image plane relative to the original camera image plane. The resulting vector can be used to return the pixel coordinates in the original camera image.

[0085] General Single-Step Image Mapping

[0086] Planar or radial plane mapping can operate on the undistorted (fisheye corrected) image. The steps of correcting distortion and rectification can also be combined into a single image mapping process. This can be more efficient and save memory.

[0087] As mentioned above, not only planes, but also other shapes, such as spherical or cylindrical, can be used for correction. In terms of the trade-off between image distortion and the feasibility of depth estimation, these other shapes may be beneficial. In general, the epipolar line is an epipolar curve. Figure 8a and 8b show spherical, planar, and cylindrical mappings with baseline alignment. The pole is determined by the movement of the camera.

[0088] Based on spherical or planar mapping, other mappings can be achieved through coordinate warping.

[0089] Determine 3D Pole

[0090] In an embodiment with a camera mounted on a vehicle, the odometer from the vehicle system bus can provide the rotation and movement of the vehicle. To calculate the pole position of the current camera, the position of the previous camera is calculated in the current vehicle coordinate system. To calculate the pole of the previous camera, the position of the current camera in the previous vehicle camera system is calculated. The vector from the current camera to the previous camera (and vice versa) points to the pole, and this vector is called the baseline. The following formula is based on the measured mechanical odometer:

[0091]

[0092]

[0093] where, is the current vehicle-to-world rotation, and is the previous world-to-vehicle rotation. These two matrices can be combined into an incremental rotation matrix. c v is the position of the camera in vehicle coordinates (the extrinsic calibration of the camera). δ is the incremental translation in world coordinates. Visual odometry can provide the incremental translation in vehicle coordinates, in which case the equation is further simplified.

[0094] Baseline Alignment

[0095] To make the mapping calculation easier, the resulting image can be baseline-aligned. Geometrically, this means that the virtual camera rotates in such a way as to compensate for the camera's movement so that the virtual camera is perpendicular to the baseline or collinear with the baseline. Thus, the vehicle coordinate system is rotated to the baseline. This rotation can be determined by rotation about an axis and an angle. The angle between the epipolar vector and the vehicle coordinate axes is determined as follows:

[0096] c(θ) = -e c ·(1, 0, 0) or cox(θ) = -e c ·(-1, 0, 0)

[0097] The rotation axis is the cross product of two vectors:

[0098] u = e c ×(1, 0, 0) or u = e c ×(-1, 0, 0)

[0099] Then the rotation matrix can be determined, for example, using the standard axis-angle formula:

[0100]

[0101] Using this rotation (referred to as ), any ray in the polar coordinate system can be transformed to the vehicle coordinate system, and vice versa.

[0102] Polar Coordinate System

[0103] To define the mapping for correction, a spherical coordinate system is defined with the pole as the pole and using the latitude and longitude angles, as Figure 9 shown. In the absence of any rotation (pure forward or backward movement of the vehicle), the pole points in the direction of the X-axis of the vehicle coordinate system. In this case, ±90° latitude points to the X-axis and 0° longitude points to a point above the vehicle (0, 0, 1). The circles around the sphere that define the plane parallel to the X-axis show the circle of polar lines on the sphere. Any mapping should map these circles to horizontal lines in the corrected image.

[0104] Spherical Mapping

[0105] When considering the Figure 9 geometry shown, in the spherical mapping, the polar line coordinates are directly mapped to the x and y coordinates in the output image. Latitude is mapped to the y coordinate and longitude is mapped to the x coordinate. The pixel density of the output image defines the density of the angles. For each view, the ranges of longitude and latitude can be defined. For example, in the Figure 1 vehicle environment shown:

[0106] · Rear camera: longitude [0°, 360°], latitude [-10°, -90°]

[0107] · Front camera: longitude [0°, 360°], latitude [10°, 90°]

[0108] · Left mirror camera: longitude [70°, 180°], latitude [-80°, 80°]

[0109] · Right mirror camera: longitude [-70°, -180°], latitude [-80°, 80°]

[0110] Now, these are mapped to the camera image coordinates in the following order:

[0111] · Convert the polar line coordinates to a ray vector of (x, y, z) coordinates.

[0112] · Rotate the ray from the polar line coordinates to the vehicle coordinates.

[0113] · Rotate the ray from the vehicle to the camera coordinate system.

[0114] · Use camera projection to determine the original source image coordinates.

[0115] The following formula converts polar coordinates to a ray vector with (x, y, z) coordinates:

[0116]

[0117] where, is the longitude and γ is the latitude. The ray is converted to a camera ray as follows:

[0118]

[0119] R VC is the camera rotation matrix from the vehicle to the camera coordinate system. Then built-in functions are applied to retrieve the pixel coordinates in the source image to obtain the epipolar coordinates

[0120] Baseline Alignment Plane Mapping

[0121] For planar mapping, the same mechanism as for spherical mapping is used. First, the planar coordinates (x, y) are converted to epipolar coordinates, but the calculation steps are the same. The conversion is as follows:

[0122]

[0123] For rear and front cameras on the vehicle, since the latitude values converge to singularities that require radial mapping, the conversion is as follows:

[0124]

[0125] The viewport can be defined as the angular extent, just like in spherical mapping. These can be converted to ranges in planar coordinates, which can be converted to pixel ranges at a given pixel density.

[0126] Cylindrical Mapping

[0127] Cylindrical mapping is a hybrid of spherical and planar. They are especially useful for mirror cameras on the vehicle. The mappings for a vertical or horizontal cylinder are respectively:

[0128] or

[0129] Conical Mapping

[0130] Conical mapping and spherical mapping have the property of being closer to an extended focus than cylindrical mapping because less stretching is required due to the shape of the mapping surface. The drawback is that these mappings do not preserve the shape of an object when it moves from one camera image to the next. However, the range on the conical view is much better, and the detection quality, especially for nearby detections, is better than in cylindrical mapping.

[0131] On some vehicles, the conical viewport may reduce the quality of ground and curb detection, but it is computationally cheaper than the spherical viewport. Figure 10 Shows an example of a conical mapping surface. When considering embodiments with cameras mounted on a vehicle, the conical view may only be used for front and rear cameras, not side cameras.

[0132] General Mapping

[0133] In addition to spheres, planes, or cylinders, other mappings are possible. They can be defined by the following three functions: f(x), g(y), and h(x). Assuming Figure 1 the camera configuration in

[0134] For the mirror (left and right views):

[0135] For the rear and front views:

[0136] For the plane case: f(x) = h(x) = atan(x), g(y) = atan(y)

[0137] For the spherical case: f(x) = h(x) = x, g(y) = y

[0138] Custom-designed functions may be beneficial for the correct balance between the field of view and spatial resolution in the depth map or the distortion of the stereo matcher. Ideally, the functions f(x), g(y), and h(x) should achieve the following goals:

[0139] · atan(x) ≤ f(x) ≤ x and atan(y) ≤ g(y) ≤ y to achieve a good balance between spatial resolution, field of view, and distortion;

[0140] · x ≤ h(x) ≤ c·x, requiring the function to extend near the poles to have a greater parallax resolution there, but to return to linear behavior for larger x; and

[0141] · The formula is simple and the computational amount is small.

[0142] For example:

[0143] Map to Grid

[0144] Computing the mapping from epipolar coordinates to source pixel coordinates for each destination pixel is time consuming. This can be accelerated by computing these mappings only for a grid of destination coordinates. This mapping can be used using automatic destination coordinate generation using a regular grid in the destination coordinates. For each node on the grid, a mapping to the original source pixel coordinates is computed. Using bilinear interpolation, the pixels in the destination image are mapped according to the provided grid. The grid cell size can be defined in such a way that the runtime of the grid creation is feasible, but on the other hand, the distortion due to the grid is small enough.

[0145] Surface Combination

[0146] It is also possible to combine surfaces for epipolar mapping, including using multiple planes for epipolar correction. This is discussed in "Omnidirectional stereo vision using fisheye lenses" by Drulea et al. cited above. For example, Figure 11 A mapping surface comprising two planes is shown in . If this is considered in the case of a vehicle-mounted camera, a first horizontal plane H will be used, for example, to reconstruct the ground; a second plane V perpendicular to the first plane will be used for other objects.

[0147] Multiple additional surfaces can be used. Since cylindrical (and planar) viewports perform well near the car when detecting the ground, curbs, and objects, it is advantageous to continue to benefit from their properties. In order to also detect near the expanded focus, another surface can be added to cover that area. Ideally, the mapping surface should balance the stretching of the image. Since the expanded focus is a small part of the image, the viewport should also be small; and it should cover the range starting from the detection end of the cylindrical view and close to the expanded focus.

[0148] Figure 12a A combination of a cylindrical viewport 120 and a frustoconical viewport 122 (based on a frustum of a cone) is shown. Such modified surfaces are advantageous for front and rear vehicle cameras where the cameras move in / out of the scene. Rearview mirror cameras typically only use cylindrical surfaces because they move laterally to the scene.

[0149] Figure 12b Another modified surface is shown, in which a hemispherical surface 124 is added to the cylindrical surface 120. Note that for both modified surfaces, mapping the area immediately adjacent to the extended focus is not useful, since no valid disparity and depth can be extracted in this area. Therefore, the surface under consideration may have a portion of the surface 126 missing at the extended focus, such as Figure 12a and 12b shown.

[0150] Depth Calculation and Point Cloud Generation

[0151] After running the stereo matcher, the resulting disparity map is an image that contains the horizontal movement from the previous image to the current image. These movements can be converted into depth information or directly into a point cloud of 3D points. Depending on the epipolar rectification method used, different formulas can be used to calculate this conversion.

[0152] Plane Depth Calculation

[0153] Baseline-aligned planar epipolar rectification is a special case that allows the determination of depth through a simple calculation procedure. Turning Figure 13a and considering a camera mounted on the side of a vehicle. The depth d can be calculated using similar triangles. Relative to the focal length f with depth d, the disparity q is similar to the baseline b = |δ|:

[0154]

[0155] where f and q can be given in pixels.

[0156] Correspondingly, turning Figure 13b and considering a camera mounted on the front or rear of a vehicle, the following relationship can be derived:

[0157]

[0158]

[0159] This can be simplified to:

[0160]

[0161] It can be seen that there is not only a dependence on the disparity q2 - q1, but also a dependence on q1. This means that the depth calculation in the disparity map depends on your position in the map. Similarly, this means that at the same distance, an object with a small q1 has a smaller disparity than an object with a large q1. This reduces the spatial resolution in these cases because the feature matcher operates at the pixel level. Naturally, when approaching the poles, the spatial resolution approaches zero. However, different mapping functions can reduce this effect by balancing the relationship between depth and disparity to some extent.

[0162] Spherical Depth Calculation

[0163] In Figure 14 the spherical case shown, depth calculation is more difficult, but by considering a single epipolar circle in the spherical mapping, the following triangular relationship can be derived:

[0164]

[0165]

[0166] The angles α and β (latitude) correspond to the horizontal pixel positions of the matched pixels in the spherically rectified image and are easily retrievable. In addition, longitude can be considered. The depth d calculated above is the distance from the baseline, so that:

[0167]

[0168] is the depth from the virtual camera plane. Although spherical mapping provides a large field of view, the relationship between parallax and distance is not very favorable. For small angles α, β, an approximate variant of the above equation is:

[0169]

[0170]

[0171] The relationship between parallax and depth is not constant and, in the case of depth, is even quadratic with the angle of incidence. This means that for a mirror view camera, at the edges of the rectified image, the spatial resolution is very low. Therefore, different mappings that balance the field of view and spatial resolution may be advantageous.

[0172] Cylindrical Depth Calculation

[0173] The cylindrical case is again a mixture of planar and spherical. Depending on whether it is a horizontal or vertical cylinder, the depth is calculated using the above planar or spherical method. In the case of planar depth, the execution of the longitude adjustment is equivalent to the spherical longitude adjustment.

[0174] General Depth Calculation

[0175] For spherical mapping, the following equation is used:

[0176]

[0177]

[0178] According to the above definitions (see "General Mapping" section), the following can be used:

[0179] α = f(x1), β = f(x2) or α = h(x1), β = h(x2)

[0180] Taking longitude into account, it follows that:

[0181] d′ = sin(g(y))d

[0182] Point Cloud from Depth

[0183] If the above method is used to calculate the distance of each disparity image pixel relative to the virtual camera plane, the resulting coordinates can be converted into vehicle coordinates to form a point cloud. Assume that the Euclidean coordinates in the epipolar reference system are relative to the selected camera (current) matrix can be used to rotate into the vehicle coordinate system and add the external position of the camera to obtain vehicle coordinates. Computationally, this method is very effective for baseline-aligned plane mapping because for each pixel, only very few and very simple operations are involved.

[0184] Triangulation Point Cloud

[0185] As an alternative to forming a point cloud from depth, triangulation can be used to generate a point cloud from a disparity map. This takes as input the following: two rays of the vehicle and the motion vector in the coordinate system under consideration. Through some basic and known operations (multiplication, addition, subtraction, and division), the output of the 3D position of the triangulated rays can be obtained.

[0186] To determine the rays from the disparity map, two methods can be chosen: using the depth calculation mentioned above, or using a pre-computed ray grid.

[0187] The former method is less effective than directly generating a point cloud from depth values. The latter method is effective. When generating the grid for the epipolar mapping mentioned above, the required rays are also calculated. Intermediate result:

[0188]

[0189] can be used for each node in the grid and can be stored for later use. r is the ray after rotating from the epipolar coordinate system to the vehicle coordinate system. Now, for each pixel in the disparity map, the corresponding ray can be bilinearly interpolated (or using a more advanced method, such as spline interpolation) from the rays stored in the nodes around the pixel. When choosing the grid cell size for the epipolar mapping, the method described in this section can be considered so that the generated point cloud is accurate enough.

[0190] Triangulation Method

[0191] Since the method of triangulation is known, for the sake of brevity, only the final equations are given. Given two vehicle rays r c , r p and the incremental motion δ of the vehicle, the following six quantities can be defined:

[0192] a = r c · r c = |r c | 2

[0193] b = rp ·r c

[0194] c = r p ·r p = |r p |

[0195] d = r c ·(-δ)

[0196] e = r p ·(-δ)

[0197] p = ac - b 2

[0198] If both rays are unit vectors, the values of a and c will be 1, and the following formula can be further simplified:

[0199]

[0200]

[0201] To provide two intersection points of the rays. Usually, two rays in 3D space do not necessarily intersect. The average of two 3D points will be the closest point of the two rays. However, in our example, the two rays intersect the polar curve, so there will be intersection points. This means that only one of the above equations is sufficient. However, by calculating two 3D points and using the average result as the final 3D point output, inaccuracies in ray calculations caused by interpolation can be at least partially resolved.

[0202] Application

[0203] By using the above techniques, images from a moving monocular camera on a vehicle with a wide - field - of - view lens can be used to form part of a point cloud that approximates a portion of the local terrain of the vehicle's surroundings as the vehicle moves through the environment. The images formed by stitching together the images from the moving monocular camera can then be mapped onto the surface defined by the point cloud. Images of the resulting virtual scene can be evaluated from any virtual camera pose. Since the virtual images have a similar topography to the imaging scene regardless of the virtual camera pose, the images provided will be realistic.

Claims

1. A method for forming a point cloud corresponding to the terrain of an imaging environment, comprising: Obtaining a first image of the environment using a camera (C1) with a wide field of view lens mounted on a vehicle; Changing the camera attitude by an adjustment greater than a threshold; Obtaining a second image of the environment using the camera in the changed attitude (C2); Mapping the first image onto a first surface to form a first mapped image; Mapping the second image onto a second surface to form a second mapped image, wherein the first surface and the second surface are defined by the same non-planar geometry (120, 122, 124); wherein the non-planar geometry is axisymmetric about the optical axis of the camera; wherein the non-planar geometry includes: A first non-planar surface (120) defining a cylindrical surface; and A second non-planar surface (122, 124), the first non-planar surface extending from the camera to the second non-planar surface, and the second non-planar surface converging from the first non-planar surface towards the optical axis; Dividing one of the first mapped image or the second mapped image into pixel blocks; For each pixel block, forming a depth map by performing a search on the other of the first mapped image or the second mapped image to evaluate the position difference of the pixel block in each mapped image; When the vehicle moves in the environment, converting the depth map into a part of a point cloud corresponding to the local terrain of the vehicle's surrounding environment; And scaling the point cloud according to the adjustment of the camera attitude.

2. The method according to claim 1, wherein, Performing the search includes performing a one-dimensional search on the image data along the rectified epipolar line (b).

3. The method according to claim 1 or 2, wherein The method further includes evaluating the change in the camera attitude using the odometry information provided by the vehicle odometry sensor.

4. The method according to any one of the preceding claims, wherein, The wide field of view lens is a fish-eye lens.

5. The method according to claim 1, wherein, The second non-planar surface defines a frustum of a cone surface (122).

6. The method according to claim 1, wherein The second non-planar surface defines a hemispherical surface (124).

7. The method according to claim 6, wherein Performing the mapping by constructing a sparse look-up table for a subset of pixels in the mapped image and storing the corresponding camera image coordinates of the pixels with decimal precision.

8. A vehicle (100) for moving in a forward direction, the vehicle comprising: A camera (101) whose optical axis is aligned in the forward direction; And A camera (104) whose optical axis is aligned in the backward direction, the backward direction being opposite to the forward direction; wherein each camera operates according to the method of any one of claims 1 to 7.

9. The vehicle according to claim 8, wherein, The vehicle further includes: Two cameras (102, 103) aligned in opposite directions perpendicular to the forward direction, wherein the two cameras operate according to the method of any one of claims 1 to 4.

10. The vehicle according to claim 8 or 9, wherein, The cameras provide information for the vehicle's automatic parking or autonomous driving system.

11. A method for displaying a scene around a vehicle, comprising: Obtaining images from a plurality of cameras located around the vehicle with overlapping fields of view; Stitching the images into a stitched image; Mapping the stitched image onto a surface defined by a point cloud generated by the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and apparatus for creating 3d image of vehicle surroundings

    CN104093600A