Method for determining at least one scaling factor for a first depth map
A method for accurately scaling depth maps using pixel-specific ratios and frequency distribution analysis addresses the inconsistency issue in monocular depth maps, enhancing fusion quality and computer vision tasks.
Patent Information
- Application Number
- DE102024100899
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-12
- Publication Date
- 2025-07-17
AI Technical Summary
Existing depth maps created from camera images with monocular optics often have inconsistent scales due to the scaleless nature of monocular optics, leading to discontinuities when fused, and existing methods for scaling these maps are prone to imprecision and distortion.
A method involving pixel-specific ratio calculations and frequency distribution analysis is used to determine a global scaling factor for depth maps, utilizing a Gaussian mixture model to approximate the histogram of these ratios, ensuring accurate and precise scaling before fusion.
This approach ensures accurate scaling of depth maps, enabling reliable fusion and improving the performance of computer vision tasks such as local surface normals, road surface estimation, and low-height object detection.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Disclosure of the invention
[0001] The invention provides a novel method for scaling multiple depth maps for subsequent fusion of the depth maps. Depth maps represent depth information in images. By "depth information" we mean information about how far away objects or displayed object surfaces are in an image (e.g., a camera image). Depth information can be obtained in various ways, e.g., through photogrammetry (also called "structure from motion"), machine learning (also called "deep learning"), or stereo disparity methods. Depth information for an image (e.g., a camera image) is often obtained using different techniques, so that multiple depth maps exist for one image (e.g., a camera image). In order to fuse this data in a post-processing step, the scale of the depth maps must be adjusted. If this is not done, discontinuities in the resulting fused depth map will result.Therefore, the present invention shows how to correctly align multiple depth maps to enable central fusion.
[0002] Here, a new method is proposed how different depth maps can be scaled to each other in order to enable a fusion of depth maps and a joint data processing of the depth information from different depth maps.
[0003] Described here is a method for determining at least one scaling factor for a first depth map, comprising the following steps: a) Extracting first depth values d 1,i for a plurality of pixels i of the first depth map; b) Determining a comparison depth value d 2,i for each first depth value d 1,i ; c) Calculating ratios Sr,i=d1,id2,i of the first depth value d 1,i with the comparison depth value d 2,ifor each pixel i; d) Determine a frequency distribution of the ratio S r,i over all pixels i; and e) determining the at least one scaling factor based on the frequency distribution.
[0004] It is particularly preferred if the comparison depth value d 2,i in step b) was extracted from a second depth map for a pixel corresponding to the respective pixel i in the first depth map.
[0005] The described method is particularly suitable for processing a first depth map and a second depth map together, both of which provide depth information for an image from a single camera. Preferably, the first depth map and the second depth map are created based on the same camera image using different algorithms and, if appropriate, preferably with the use of different additional data.
[0006] Particularly preferably, a depth is specified for each pixel in the depth maps. The depth describes a distance that the object visible in the respective pixel in the vicinity of the camera or the object described by the pixel has from the camera. In other words: the depth describes the distance of a scene point to the camera. A pixel is a projected 2D point in the image (a pixel). A scene point is a 3D point in the world with a distance to the camera. The distance of the scene point to the camera can, for example, refer to a direct distance to the camera. This means that the distance along a viewing direction from the camera to the pixel is measured. It is also conceivable and included here, but not preferred, that the depth is defined in each case relative to a camera plane in which the camera is located and which has a defined orientation to the camera orientation.Preferably, such a camera plane is aligned normal to an axis of the camera.
[0007] Different algorithms and methods for creating depth maps from camera images can be applied to different depth maps. Since depth information can usually only be estimated relatively (i.e., compared to other depth information) using conventional algorithms, there is often a difference in scaling between such depth maps. The method described here describes an algorithm that can be used to find a scaling factor that describes this scaling and thus enables a combination of depth maps.
[0008] The method can also be used in principle if depth information or depth maps are to be evaluated that were created based on camera images that originate from different cameras or that were recorded by a single camera with a time delay.
[0009] Additionally, it is then necessary to perform pixel mapping of the different camera images. Otherwise, the comparison of depth information at the pixel level is not feasible.
[0010] An algorithm that creates a depth map provides depth values for each pixel of a single image. Such an algorithm may also include additional images, for example, those taken before or after the same camera, to generate the depth information of the depth map.
[0011] The creation of the first depth map and the creation of the second depth map may differ in that different additional images were used to create the respective depth map.
[0012] In principle, deviations in the scaling of depth maps created from camera images can occur primarily because cameras typically capture images using monocular optics. Information captured using monocular optics is initially scaleless, and an external reference is required to obtain depth information. The external reference can, if necessary, be generated using continued vehicle movement; as the vehicle continues to move, the image from the same camera shows visible objects in the camera image from a different perspective. Camera images captured previously or subsequently by the same camera can therefore be used to create depth maps. This usually also requires additional data describing the vehicle's movement.Such additional data can be received, for example, via a vehicle's bus system, and such data is available, for example, as navigation data and / or vehicle movement data from the drivetrain. If necessary, data from additional cameras mounted at a different location on the vehicle or from completely different sensors, such as LIDAR or RADAR sensors located in the camera's area, can also be used as additional data for creating depth maps.
[0013] To create depth maps based on camera images with monocular optics, filters generated using deep learning, or neural networks, which can optionally be used in the form of data filters, are often particularly effective. However, these filters have the fundamental problem that the training data suitable for training such filters can often only generate relative depth information. If attempts are made to obtain absolute depth information using such approaches, the algorithms are often very sensitive to the camera mounting position and would theoretically have to be completely retrained if the camera mounting position is changed.
[0014] In principle, there are many possible applications for the approach described here of finding a scaling factor for scaling between two depth maps, because cameras with monocular optics are widespread.
[0015] Here, it is proposed to determine a ratio to a comparison depth value for each individual pixel of a depth map, wherein the comparison depth value is preferably a second depth value for a corresponding pixel of a second depth map. This multitude of ratios are essentially pixel-specific scaling factors that are calculated in step c). This results in a type of scaling factor map that regularly shows different scaling factors distributed across the entire depth map. A frequency distribution of the ratios thus determined is then determined (step d). Based on this frequency distribution, a global scaling factor (overall scaling factor) is determined for the entire first depth map with respect to the comparison depth values (in particular with respect to the second depth map). This frequency distribution is evaluated according to step e) to determine the scaling factor.
[0016] Compared to other approaches that directly determine an overall scaling factor from the individual depth maps, this approach prevents certain inaccuracies in the scaling of the individual depth maps from compensating for each other when determining the scaling factor and thus distorting the scaling factor. The approach described here ensures that the scaling for each individual pixel of the depth map is always correct. The frequency distribution makes fluctuations in the scaling factor across the entire depth map visible. The frequency distribution is also often referred to as a histogram.
[0017] The method is particularly preferred if, for carrying out the method steps, groups and / or subsets of actual image pixels are considered as pixel i.
[0018] Then the depth values d are determined 1,iand the determination of the comparison depth values d 2,i and the conditions Sr,i=d1,id2,i then for each of these groups of image pixels. Such a group of image pixels can be defined, for example, as n times n image pixels, for example, as 5 x 5 image pixels. In embodiments, for such a group of image pixels, an average / common depth value d1 and an average / common comparison depth value d 2,i which are then processed according to the described procedure.
[0019] According to this embodiment of the described method, the computational effort required to implement the described method can be reduced. In further embodiments of the method, it may also be possible to consider only a subset of pixels, ideally with the same spatial distribution. If necessary, every Xth pixel (e.g., every tenth pixel) or groups of pixels can be considered as a single pixel and examined as a single pixel for the implementation of the method described here.
[0020] The method described here describes an approach, also known as "histogram-based," that can be used to find scaling factors for scaling between depth maps. For ideal data, the ratio extracted from a single pixel would be satisfactory. However, a variety of errors can occur during depth map creation, so a statistic across the entire image is the most robust approach.
[0021] It is preferred to use a so-called Gaussian mixture model to determine the scaling factor of the described method. In this way, the superposition of several Gaussian distributions is used to approximate the histogram. By determining the mean of the distributions, the presence of a single peak value as opposed to multiple strong peak values can be determined. The standard deviation provides information about the steepness of the curve and thus the reliability of the estimate.
[0022] It is also preferred if the at least one scaling parameter in step e) is determined as at least one parameter of a frequency distribution.
[0023] A common and suitable frequency distribution here is, for example, the Gaussian distribution described above, which is also referred to as the normal distribution. It is preferred if, in step e), at least one normal distribution is determined that approximates the frequency distribution determined in step d). A normal distribution has various parameters with which it is described, in particular the expected value or mean, and the variance.
[0024] It is further preferred if the at least one scaling parameter is determined as the mean value of the normal distribution.
[0025] If the scaling of the two depth maps or the depth map and the comparison depth values is uniform, the frequency distribution will correspond to a sharp normal distribution with a high and significant mean and a narrow standard deviation.
[0026] It is also preferred if, in step e), a confidence measure for the determined at least one scaling factor is determined from a standard deviation of the normal distribution.
[0027] By using the approach of determining the scaling factor using a normal distribution, the standard deviation becomes a parameter that is well suited as a confidence measure for the scaling factor and can even be used directly for this purpose.
[0028] Furthermore, it is preferred if in step e) it is checked whether the frequency distribution has a plurality of extreme values, wherein a plurality of extreme values in step e) are included in an evaluation according to which the accuracy of the depth information of the first depth map and / or a second depth map is evaluated.
[0029] It is further preferred if, in step e), in the case of a frequency distribution with a plurality of extreme values, the frequency distribution is reproduced with a plurality of normal distributions.
[0030] Furthermore, it is preferred if the at least one scaling factor determined in step e) is a relative scaling factor with which the scaling of the first depth values of at least one first depth map is comparable with the scaling of comparison depth values, wherein the method additionally comprises the determination of an absolute reference scale for depth values of the at least one depth map.
[0031] It is also preferred if the absolute reference scale in step e) is determined by converting the absolute reference scale of a comparison depth map with the at least one scaling factor.
[0032] Furthermore, it is preferred if the absolute reference scale is determined based on at least one of the following approaches: - Determining a calibrated base width of two cameras with which the first depth map was created using a stereo disparity method; - Using size references within the first depth map; - Analysis of the at least one first depth map with a neural network to determine an absolute reference scale.
[0033] The main application of the approach presented here is depth map fusion. Once the depth maps have been appropriately scaled, a depth map fusion algorithm can be applied.
[0034] The method described here enables the fusion of different depth maps obtained through photogrammetry. These depth maps are obtained from image pairs with different temporal separations, which means that the optical flow is calculated for images with different temporal separations from the current image.
[0035] In the context of photogrammetric methods, dynamic objects, low ego vehicle speeds, occlusions, and repetitive structures reduce the quality of the depth map. Therefore, the use of neural network filters is desirable, which could, in principle, enable the automatic minimization of such influences.
[0036] The fusion of these depth maps leads to a significant improvement in the performance of geometric computer vision algorithms, such as: local surface normals, road surface estimation, and low-altitude object detection.
[0037] Furthermore, the applications of relative scaling are not limited to the fusion of depth maps. The presented relative scaling technique can also be used for evaluation purposes when a depth map is available as a true reference ("ground truth"). The true reference ("ground truth") can be obtained, for example, from LiDAR or stereo data, and the evaluated depth map can be a depth map obtained using photogrammetry or depth from mono. Consequently, in addition to a trivial pixel-by-pixel depth value comparison, the scale difference is also evaluated.
[0038] Also to be described here is a device for data processing comprising a processor which is adapted / configured to carry out the described method.
[0039] The device can, in particular, be a control unit and / or a module within an overall system of a motor vehicle, which is configured to execute highly automated and possibly even autonomous driving functions. The depth maps processed using the method were preferably created based on sensor data (in particular based on camera data), which are processed with the overall system to provide parameters and / or control variables for providing the highly automated driving function.
[0040] Also claimed is a computer program product comprising instructions which, when the computer program product is executed by a computer, cause the computer to carry out the described method.
[0041] A computer-readable storage medium is also to be described, comprising instructions which, when executed by a computer, cause the computer to carry out the described method.
[0042] The invention and the technical context of the invention are explained in more detail below with reference to the figures. The figures show preferred embodiments to which the invention is not limited. It should be noted in particular that the figures, and in particular the proportions shown in the figures, are only schematic. They show: Fig. 1: schematic of an inconsistent fusion of depth maps; Fig. 2: schematic representation of a fusion of depth maps using the method described here; Fig. 3: a flowchart of the described method; and Fig. 4a to 4c: different frequency distributions that can be investigated using the described procedure.
[0043] The Fig. Figure 1 shows an example of offset or differently scaled depth maps 11 created with a camera 10, where different algorithms may have been used to create the depth map from camera raw data. The two offset depth maps 11 may also have been created with different raw data. Fig. 2 further shows the result of an inconsistent fused depth map 12. From such an inconsistent fused depth map 12, reliable information cannot usually be obtained.
[0044] The Fig. In comparison, Figure 2 shows an example according to which the two shifted or differently scaled depth maps 11 were merged into a scaled fused depth map 13 using the method described here. In the lower part of the Fig. 2 shows exemplary scaling scales 14 of the first depth map 1 and the second depth map 2. If at least one suitable scaling scale 14 has been determined between these scaling scales 14 using the described method, the two differently scaled depth maps 11 can be merged into a scaled fused depth map 13, from which reliable information can be obtained.
[0045] The Fig. Figure 3 shows a schematic flow diagram of the described method. At the top, a first depth map 1 can be seen from which first depth values 3 can be extracted. Further visible is a second depth map 2 from which second depth values 4 can be determined, which are to be used here as comparison depth values in the sense of the described method. This corresponds to steps a) and b) of the described method. Subsequently, a frequency distribution 5 of the ratios Sr,i=d1,id2,i of the first depth value 3 d 1,i with the comparison depth value d 2,i for each pixel i. The ratios Sr,i=d1,id2,i are calculated according to step c) and the frequency distribution 5 is determined according to step d). The scaling factor is preferably determined using statistical parameters of this frequency distribution 5. Approximating the frequency distribution 5 with a normal distribution 6 is particularly suitable, so that the scaling factor can be assumed to be the mean 7 of the normal distribution 6. A standard deviation 8 of the normal distribution 6 can, if necessary, be output as a confidence measure of the determined scaling factor.
[0046] The Fig. 4a to Fig. 4c now show different forms of frequency distributions 5, which can be investigated within the framework of the described procedure. Fig. Figure 4a shows a frequency distribution 5 with a clearly pronounced extreme value 9 of the frequency, which can be approximated quite well with a normal distribution 6 and allows a reliable determination of a scaling factor.
[0047] Fig. Figure 4b shows a situation with multiple peaks / extreme values 9 of similar height. This suggests that consistent fusion may not be possible. Such cases can occur when the depth reconstruction in an algorithm for adding depth information to a camera image did not work as expected. Such a configuration is a very difficult error to correct, since multiple equivalent solutions exist and the depth information can be scaled differently in different areas of the depth map 1,2. If necessary, such a situation can be used as a reason to discard a depth map 1,2 for further processing.
[0048] The Fig.Figure 4c shows another situation indicating impaired quality of at least one depth map 1,2. An extreme value 9 is very weak here. The described method could be used to determine a low confidence of a scaling factor based on a standard deviation. Such a situation may indicate that the relative scale of a camera 10 used to create the depth map 1,2 is not well defined. Such an effect can arise from rolling shutter effects, motion blur, or low accuracy of the depth reconstruction (possibly due to poor optical flow quality).
Claims
[1] Method for determining at least one scaling factor for a first depth map (1) comprising the following steps: a) Extracting first depth values (3) d 1,i for a plurality of pixels i of the first depth map (1); b) Determination of a comparison depth value (4) d 2,i for each first depth value (3) d 1,i ; c) Calculating ratios Sr,i=d1,id2,i of the first depth value (3) d 1,i with the comparison depth value d 2,i for each pixel i; d) Determine a frequency distribution (5) of the ratio S r,i over all pixels i; and e) determining the at least one scaling factor based on the frequency distribution (5). [2] Method according to claim 1, wherein the comparison depth value d 2,iin step b) was extracted from a second depth map (2) for a pixel corresponding to the respective pixel i in the first depth map (1). [3] Method according to one of the preceding claims, wherein for carrying out the method steps, groups and / or subsets of actual image pixels are considered as pixel i. [4] Method according to one of the preceding claims, wherein the at least one scaling parameter in step e) is determined as at least one parameter of at least one frequency distribution (5). [5] Method according to claim 4, wherein the at least one scaling parameter is determined based on at least one mean value (7) of the at least one normal distribution (6). [6] Method according to claim 4 and 5, wherein in step e) a confidence measure for the determined at least one scaling factor is determined from a standard deviation (8) of the at least one normal distribution (6). [7] Method according to one of the preceding claims, wherein in step e) it is checked whether the frequency distribution (5) has a plurality of extreme values (9), wherein a plurality of extreme values (9) in step e) are included in an evaluation according to which the accuracy of the depth information of the first depth map (1) and / or a second depth map (2) is evaluated. [8] Method according to claim 7, wherein in step e) in the case of a frequency distribution (5) with a plurality of extreme values (9) a simulation of the frequency distribution (5) with a plurality of normal distributions (6) takes place. [9] Method according to one of the preceding claims, wherein the at least one scaling factor determined in step e) is a relative scaling factor with which the scaling of the first depth values (3) of at least one first depth map (1) is comparable with the scaling of comparison depth values (4), wherein the method additionally comprises the determination of an absolute reference scale for depth values (3, 4) of the at least one depth map (1, 2). [10] Method according to claim 9, wherein the absolute reference scale is determined in step e) by converting the absolute reference scale of a comparison depth map (2) with the at least one scaling factor. [11] The method of claim 9 or 10, wherein the absolute reference scale is determined based on at least one of the following approaches: - determining a calibrated base width of two cameras (10) with which the first depth map (1) was created by a stereo disparity method; - Using size references within the first depth map (1); - analyzing the at least one first depth map (1) with a neural network to determine an absolute reference scale; [12] A data processing device comprising a processor adapted / configured to carry out the method according to any one of claims 1 to 11. [13] A computer program product comprising instructions which, when the computer program product is executed by a computer, cause the computer to carry out the method according to any one of claims 1 to 11. [14] A computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Augmented reality 3D reconstruction
EP4049245B1
Image distance measuring device and method
KR1020230096167A
Method for estimating depth, electronic device, and storage medium
US20230394692A1