Method and computing unit for determining at least one normal vector on a surface
The method calculates surface normals using affine parameters from disparity values, addressing noise and edge accuracy issues in stereo images, enabling efficient and accurate real-time processing.
Patent Information
- Application Number
- DE102023212862
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-18
- Publication Date
- 2025-06-18
AI Technical Summary
Existing methods for determining surface normals from stereo images are prone to noise and inaccuracies, particularly when using disparity maps, and struggle to maintain accuracy at edges and corners, which affects performance in real-time applications.
A method that calculates normal vectors using affine parameters derived from disparity values, employing a geometric approach without machine learning, and utilizing GPU acceleration through convolution kernels to determine normal vectors robustly and efficiently, with optional adaptive methods for edge handling.
The method achieves real-time, accurate determination of surface normals, maintaining precision at edges and corners, suitable for high-resolution applications, and is adaptable for various image processing tasks.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
The present invention relates to a method for determining at least one normal vector on a surface, and to a computing unit and a computer program for carrying it out.BACKGROUND OF THE INVENTIONSurface normal vectors may be used along with point clouds for a variety of tasks such as 3D surface reconstruction, scene segmentation, object detection, and others. A method for determining surface normals from depth maps is described, for example, in Moradi, S. et al., Fast, Accurate and Object Boundary-Aware Surface Normal Estimation from Depth Maps, arXiv:2209.08241.One possibility for the acquisition of a point cloud is the use of a plurality of cameras and the triangulation of the world coordinates from the disparity of the matching pixels.In order to calculate a 3D point cloud from a stereo image pair, the disparity map must first be determined. This can be done either by searching for pixel-by-pixel correspondences between the two images (cf. Hernandez-Juarez, D. et al., Embedded real-time stereo estimation via semi-global matching on the GPU, arXiv:1610.04121 ) or with a neural network (cf. Shen, Z. et al., PCW-Net: Pyramid Combination and Warping Cost Volume for Stereo Matching, arXiv: 2006.12797). In both cases, inaccuracies occur which lead to the disparity map being significantly contaminated with noise.The disparity map and the camera parameters may then be used to triangulate the 3D coordinates.A very fast method for determining surface normals is described in Y. Feng et al., D2NT: A High-Performing Depth-to-Normal Translator, arXiv:2304,12031, which is however very susceptible to noise. Also in S. Moradi et al., ibid., a robust estimator (with a limited range) is described, which also operates in real time, albeit slower.Disclosure of the InventionAccording to the invention, a method for determining at least one normal vector on a surface and a computing unit and a computer program for carrying it out are proposed with the features of the independent patent claims. Advantageous embodiments are the subject matter of the dependent claims and of the following description.Within the scope of the invention, a method for determining at least one normal vector on a surface from at least two disparity values is presented, which is very simple to implement and yet can operate in real time, in particular when using conventional GPU acceleration methods (graphics processing unit), such as OpenGL. The method is robust to noise and provides a novel compromise between accuracy and speed.The invention makes use of the measure to determine the normal vector from affine parameters, which results in a very efficient and robust GPU implementation that reduces the majority of the computations on convolution of the disparity map with two special convolution cores.In detail, two images are obtained, which are recorded with a stereo camera arrangement having two cameras arranged at a fixed baseline distance from one another. The obtaining may include capturing with the stereo camera arrangement itself; the obtaining may also include receiving image data captured elsewhere and / or at another time. Disparity values between corresponding pixels of the two images are determined. Thereafter, affine parameters of an affine transformation are determined or calculated depending on the disparity values, wherein the affine transformation transforms pixels of one of the two images into corresponding pixels in the other of the two images using an affine transformation matrix, respectively. The affine transformation matrix contains the affine parameters. The at least one normal vector is then determined depending on the affine parameters.In other words, the affine parameters are not predefined, but rather are calculated from the recorded images and the known geometric relationships. Once the affine parameters are known, the normal vector can be readily determined, as set forth below.The invention simplifies the determination of normal vectors for a stereo image pair. For this purpose, the normal vector is calculated as a function of an affine transformation which maps the images onto one another. The affine parameters are determined from the disparity values. In particular, the affine parameters may be determined using a least squares method based on the selected disparity in the vicinity of each pixel.The stereo image may be rectified to facilitate determining the disparity values. Rectification or equalization is understood to mean the elimination of misalignments in a pair of images. The equalization serves primarily to recalculate the image pair as if it were recorded by cameras whose sensors are located in the same plane, i.e. have a perfect standard stereo camera setup (cf. FIG. 2 ). The main advantage of equalized image pairs is that corresponding pixels share the same horizontal line, which simplifies the disparity search.The invention is based on a geometric consideration without any machine learning components. Therefore, no training data is required, the results are deterministic and depend only on the disparity. Nevertheless, the proposed methods are easily differentiable, so that the approach can be incorporated into a complex system as a component for which continuous training is possible.In advantageous refinements, an adaptive method for taking into account surface edges or edges is proposed. Thus, a large portion of undesirable smoothing of the normal vectors in the vicinity of edges can be eliminated. This can lead to loss of power, but still achieves real-time speed on an average consumer graphics processor at full HD resolution.The invention can be used particularly advantageously as an intermediate step in surface reconstruction or in segmentation tasks, which are used, for example, in a multiplicity of image processing applications. Particularly advantageous application cases are autonomous or semi-autonomous vehicles, such as robot lawn mowers or vacuum cleaners or driver assistance systems.A computing unit according to the invention, e.g. a control device of an autonomous or semi-autonomous vehicle, is configured, in particular by programming, to carry out a method according to the invention.The implementation of a method according to the invention in the form of a computer program or computer program product with program code for carrying out all method steps is also advantageous since this causes particularly low costs, in particular if an executing control device is also used for further tasks and is therefore present in any case. Finally, a machine-readable storage medium is provided with a computer program stored thereon, as described above. Suitable storage media or data carriers for providing the computer program are, in particular, magnetic, optical and electrical memories, such as hard disks, flash memories, EEPROMs, DVDs, among others. Download of a program via computer networks (Internet, intranet, etc.) is also possible. Such a download can be effected in a wired or wired or wireless manner (e.g. via a WLAN network, a 3G, 4G, 5G or 6G connection, etc.).Further advantages and embodiments of the invention will become apparent from the description and the accompanying drawing.The invention is schematically illustrated in the drawing on the basis of exemplary embodiments and is described below with reference to the drawing.Brief Description of the DrawingsFIG. 1 schematically shows a construction as can be used as the basis for the invention. FIG. 2 shows a geometric structure for recording a stereo image. FIG. 3 shows a block diagram of an embodiment of a method according to the invention. FIG. 4 shows two corresponding regions of a stereo image pair and their geometric relationships. FIG. 5 schematically illustrates how to variably determine pixels to be included.Embodiment(s) of the InventionAn embodiment of the invention will be described below with reference to Figs. 1 to 4. FIG. 1 schematically shows a structure as can be used as the basis for the invention, wherein a stereo camera arrangement 100 having two cameras 110, 111 is fastened to a vehicle 1 purely by way of example. FIG. 2 shows geometric relationships in the images 10, 11 recorded by the two cameras 110, 111, and FIG. 3 shows an embodiment of a method according to the invention in a block diagram. FIG. 4 shows two corresponding regions of the stereo image pair which are placed one above the other.The stereo camera arrangement 100 records two offset images 10, 11 of a scene or points X i containing a traffic sign 2 in the example shown in a step 210. In a preceding initialization step 200, the stereo camera arrangement 100 and an evaluating or processing unit 120 carrying out the method, which in particular has a GPU, can be initialized.The initializing step 200 may include initializing an OpenGL interface, allocating required resources to the GPU, and loading predetermined setting, particularly including pre-computed convolution kernels (convolution kernels) and camera calibration data, into the GPU.Each of these two offset images 10, 11 comprises a number of pixels, wherein at least for the plurality of pixels which are not directly or close to the edge of the image, a corresponding pixel should be present in the other image.A disparity map can be determined from the two images in the GPU according to known methods in a step 220. This contains only information about the correspondences of the points between the two images. The disparity map forms the input data of the embodiments of the invention described below.Based on the disparity map, in a next step 230 in the GPU, the world coordinates of each pixel are determined by triangulation. Tried and tested methods are also available in the prior art for this purpose.Subsequently, according to embodiments of the invention, the unknown affine transformation parameters are determined in a step 240 in the GPU.Based on the calculations in Barath, Daaniel et al., Levente (2015); Optimal surface normal from affine transformation. In: International Conference on Computer Vision Theory and Applications. SciTePress, Berlin, pp. 305-316. ISBN 978-989758091-8, an affine matrix for a transformation p2=A p1 of the pixels of one image into pixels of the other image can be derived according tof x, f y: horizontal and vertical focal lengths of the cameras 110, 111b: Baseline Spacing of Cameras 110, 111n=[n x, n y, n z]: surface normal vectorp = [x, y, z]: world coordinates of the pixelThe horizontal and vertical focal lengths can be calculated from the horizontal and vertical pixel densities k h, k v( number of pixels on the sensor for each unit length in space) and the physical focal length f according to f x= f k h and f y= f k v. It is therefore not necessary to know f x and f y because their ratio depends only on k h and k v.The normal vector n ultimately results therefrom as:In order to be able to apply this, the affine transformation between the two corresponding image sections must be known. Within the scope of the invention, a robust method for the efficient calculation of the two required affine parameters a 1, a 2( horizontal scaling and horizontal shearing) from the disparity map is presented.Referring to Figure 4, the invention proceeds from viewing two corresponding pixels c, c' and the pixels p, p' surrounding them in the two images. As shown in FIG. 4, a region transformed by the corresponding affine matrix substantially coincides with the corresponding region on the image of the other camera. If the disparity map is known, the following equation may be established based on FIG. 4 :At the same time, v'=A v. This leads toThe top row yields a linear equation for each environmental pixel. Pixels selected with N result in an overdeterminated linear equation system:With the known least squares solution for this system, the affine parameters are given as:According to embodiments of the invention, if the relative position of the selected surrounding pixels is always the same for all observed pixels (e.g., an nxn square with the observed pixel as the center), the entire matrix S=(V T V) -1 V T can be precomputed.In calculating the affine parameters, then only d needs to be constructed for each observed pixel and then multiplied by the pre-calculated matrix from the left.In such a case, there are N (=nxn - 1) pixels in the neighborhood of the central pixel (for which the normal vector is to be determined). Each pixel has a relative coordinate v i to the central pixel. For example, the diagonally upper left adjacent pixel v has i= [-1,1].Using auxiliary terms α, β and y, the following results:The two convolution kernels (convolution kernels) are then given by the two rows of the following matrix S:Each column corresponds to a nearby pixel.The two affine parameters are determined by convoluting the two rows of the matrix with d i- d c where d i is the horizontal disparity of the pixel and d c is the horizontal disparity of the central pixel. The convolution of the first row yields α 1- 1, the convolution of the second row yields α 2.In contrast to the pre-calculated constants α, β and y, s 1i and s 2i are both dependent on the relative position vector v i of the ithpixel.With one conversion, the computation of each parameter can be broken down into two steps, a discrete 2D convolution of the disparity map with a particular pre-computed kernel and a multiplication with a pre-computed constant and a subtraction.In the summations of Eq. (12) and (13), s 1i and s 2i are multiplied only by the disparity of the pixel at v i. Therefore, it is a 2D convolution with the kernel comprising s 1,1, s 1,2,..., s 1,N for the calculation of α 1, and s 2,1, s 2,2,..., s 2,N for the calculation of α 2.In both cases, the first step is a 2D convolution of the disparity map with two kernels. This is followed by the other operations with the precomputed terms δ1and δ2(defined in the above equation).The major part of the calculations is the calculation of the two 2D convolutions. In practice, this is efficiently achieved with an FFT (Fast Fourier Transform) on either CPU or GPU. However, experiments have shown that on modern GPUs and with reasonable kernel sizes, the method is already very fast without the additional performance gain of the FFT.One challenge for robust surface normal determinations is maintaining accuracy at corners and edges.One problem here is that the inclusion of points of an adjacent surface distorts the estimated normal vectors. This is generally manifested by smoothing of the normals at the edges, which is particularly problematic if the results are used for segmentation tasks.According to a further embodiment, the affine parameters are determined as a function of surrounding pixels, which are iteratively expanded starting from the observed pixel until an abort criterion is reached. The termination criterion may comprise a maximum number of iterations and / or a detected edge. For example, a kernel having a variable size is used. In this embodiment, the neighbor pixels involved are dynamically selected, rather than using a fixed shape for the kernel. The embodiment includes iteratively expanding the area of the enclosed points from the observed pixel.This approach is illustrated in Figure 5. In this case, a central pixel C is started and then moves outwards in a number K directions (K=8 in the illustration). Each touched pixel is included until an abort criterion is reached, which in the present example comprises either a step boundary (in the illustration, for example 3) or a suitable condition. This can be, in particular, the detection of an edge in the depth map, which was detected using any desired edge detection method. An edge may be given by a threshold value for the relative difference of depth to the central pixel, or by another heuristic condition.In the first pass, α, β and γ are calculated (δ1, δ2 are no longer needed because this version is not based on a convolution). In the second pass, the affine parameters are calculated using Eq. (10) and (11).Optionally, the disparity may be weighted based on the result of an edge detection filter, which may improve accuracy in certain cases.In a further step 250, knowing the affine parameters, the normal vectors are calculated according to Eq. (2).In a further step 260, the normal vectors are used in surface reconstruction or in segmentation tasks, which are used, for example, in a multiplicity of image processing applications. Particularly advantageous application cases are autonomous or semi-autonomous vehicles, such as robot lawn mowers or vacuum cleaners or driver assistance systems.The method may be implemented as a streaming pipeline in a computing unit. These steps may be described, for example, as follows: 1. optionally, initialize OpenGL, including allocating GPU resources and loading the pre-computed convolution kernels, camera calibration data, and other settings. 2. uploading the input disparity map. 3. triangulating the world coordinates of each pixel based on the disparity map. 4. determining the affine parameters. 5. calculating the normal vectors from calibration data, world position and affine parameters. 6... optionally post-processing of the normal vectors (e.g. MRF step or an adaptive averaging filter) 7... retrieving the normal map from the computing unit 8... releasing resourcesReferences included in the specificationThis list of documents cited by the applicant has been produced in an automated manner and is only included for the better information of the reader. The list is not part of the German patent application or utility model application. The DPMA does not take any adhesion for any faults or omissions.Cited Non-Patent LiteratureMoradi, S. et al., Fast, Accurate and Object Boundary-Aware Surface Normal Estimation from Depth Maps, arXiv:2209.08241
[0002] Hernandez-Juarez, D. et al., Embedded real-time stereo estimation via Semi-Global Matching on the GPU, arXiv:1610.04121
[0004] Shen, Z. et al., PCW-Net: Pyramid Combination and Warping Cost Volume for Stereo Matching, arXiv:2006.12797
[0004] Y. Feng et al., D2NT: A High-Performing Depth-to-Normal Translator, arXiv:2304,12031
[0006] Barath, Daniel et al., Levente (2015); Optimal surface normal from affine transformation. In: International Conference on Computer Vision Theory and Applications. SciTePress, Berlin, pp. 305-316. ISBN 978-989758091-8
[0028]
Claims
A method for determining at least one normal vector on a surface, comprising: obtaining two images (10, 11), wherein the two images (10, 11) are captured with a stereo camera arrangement (100) having two cameras (110, 111) arranged at a fixed baseline distance (b) from each other; determining disparity values between mutually corresponding pixels of the two images (10, 11); determining affine parameters of an affine transformation that respectively transforms pixels of one of the two images (10, 11) into corresponding pixels in the other of the two images (10, 11) using an affine transformation matrix, wherein the affine transformation matrix contains the affine parameters, depending on the disparity values of the pixels of the two images (10, 11); determining the at least one normal vector depending on the affine parameters.The method of claim 1, further comprising: determining world coordinates of the pixels of the two images (10, 11) depending on the disparity values between the corresponding pixels of the two images (10, 11); determining the at least one normal vector depending on the affine parameters and at least one world coordinate.The method according to claim 2, wherein determining the world coordinates of the pixels of the two images (10, 11) depending on the disparity values between the corresponding pixels of the two images (10, 11) comprises: triangulating the world coordinates of the pixels of the two images (10, 11).The method according to any of the preceding claims, wherein the affine transformation matrix is given by A = [ n x ( x - b ) + n y y + n z z n x x + n y y + n z z - b f x n y f y ( n x x + n y y + n z z ) 0 1] with f x, f y: horizontal and vertical focal lengths of the cameras (110, 111) b : baseline distance of the cameras (110, 111) n = [ n x, n y, n z]: surface normal vector p = [ x, y, z]: world coordinates of the pixelMethod according to any of the preceding claims, wherein the at least one normal vector is determined depending on the affine parameters according to: n = [ b k h z ( a 1 - 1) a 2 b k v z b ( - a 1 k h x - a 2 k v y + b k h + k h x ) ] where b: baseline distance of the cameras (110, 111) α 1, α 2: affine parameters p = [ x, y, z]: world coordinates of the pixel k h, k v: horizontal and vertical pixel densities of sensors of the cameras (110, 111)Method according to one of the preceding claims, wherein the affine parameters are determined as a function of a fixed number of surrounding pixels.The method of claim 5, wherein the affine parameters are determined according to a 1 = 1 + ∑ i = 1 N 1 αγ - β 2 ( γ v i x - β v i y ) } s 1 i ( d i - d c ) } d i a 2 = ∑ i = 1 N 1 αγ - β 2 ( -β v i x - α v i y ) } s 2 i ( d i - d c ) } d i where α = ∑ i = 1 N v i 1 2, β = ∑ i = 1 N v i 1 v i 2, γ = ∑ i = 1 N v i 2 v i= coordinates of the surrounding pixel i α 1, α 2: affine parameters N-1: number of surrounding pixels d i: horizontal disparity of surrounding pixel d c: horizontal disparity of central pixelMethod according to one of Claims 1 to 5, wherein the affine parameters are determined as a function of surrounding pixels which are iteratively extended starting from the observed pixel until an abort criterion is reached.The method of claim 8, wherein the abort criterion comprises a maximum number of iterations and / or a detected edge.The method according to any one of the preceding claims, wherein obtaining the two images (10, 11) comprises: capturing the two images (10, 11) with a stereo camera arrangement (100) comprising two cameras (110, 111) arranged at the fixed baseline distance (b) from each other;Method according to any of the preceding claims, further comprising at least one step selected from: determining surfaces in the two images (10, 11) depending on the at least one normal vector; segmenting surfaces (10, 11) in the two images depending on the at least one normal vector;Computing unit which is configured to carry out all method steps of a method according to one of the preceding claims.Vehicle (1) having a computing unit (120) according to Claim 12 and a stereo camera arrangement (100) having two cameras (110, 111) arranged at a fixed baseline distance (b) from one another.A computer program that causes a computing unit to perform all method steps of a method according to any one of claims 1 to 11 when executed on the computing unit.A machine readable storage medium having stored thereon a computer program according to claim 14.