Method for determining at least one scaling factor of first depth map
The method addresses the issue of inconsistent scaling in depth map fusion by determining a scaling factor through ratio analysis and Gaussian mixture modeling, ensuring accurate alignment and enhancing depth map fusion for applications like autonomous driving.
Patent Information
- Application Number
- CN202510014727.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-12
- Filing Date
- 2025-01-06
- Publication Date
- 2025-07-15
AI Technical Summary
The prior art is difficult to effectively align multiple depth maps to achieve correct fusion, resulting in discontinuity in the fusion depth map.
By extracting the depth values of multiple pixels of the first depth map, calculating the ratio of each pixel, determining the frequency distribution, and determining the scaling factor based on the frequency distribution, the scaling of the alignment of multiple depth maps is achieved.
The correct alignment and fusion of depth maps are achieved, the accuracy and reliability of the fusion depth maps are improved, and the performance of geometric computer vision algorithms is enhanced.
Smart Images

Figure CN120318088A_ABST
Abstract
Description
Technical Field
[0001] The present invention provides a new method for scaling multiple depth maps for subsequent depth map fusion. Background Art
[0002] A depth map describes the depth information in an image. Here, the "depth information" refers to information related to how far away an object in the image (e.g., a camera image) or the surface of the displayed object is. The depth information can be obtained in various ways, such as photogrammetry (also known as structure from motion), machine learning (also known as deep learning), or stereoscopic parallax methods. Generally, various techniques are used to obtain depth information related to an image (e.g., a camera image), such that there are multiple depth maps for the image (e.g., a camera image). To fuse this data in a post-processing step, the scales of the depth maps must be aligned. If not, it will result in discontinuities in the resulting fused depth map. Therefore, the present invention shows how to correctly align multiple depth maps for central fusion. Summary of the Invention
[0003] In this context, the object of the present invention is to propose a new method that shows how different depth maps scale relative to each other to achieve the fusion of depth maps and to perform joint data processing on the depth information from different depth maps.
[0004] A method for determining at least one scaling factor for a first depth map is described herein, the method having the following steps:
[0005] a) Extract a first depth value d for a plurality of pixels i of the first depth map 1,i ;
[0006] b) Determine a comparison depth value d for each first depth value d 1,i ; 2,i ;
[0007] c) Calculate the ratio between the first depth value d for each pixel i 1,i and the comparison depth value d 2,i ;
[0008] d) Determine the frequency distribution of the ratio S r,i over all pixels i; and
[0009] e) Determine at least one scaling factor based on the frequency distribution.
[0010] It is particularly preferred if, in step b), the comparison depth value d of the pixel corresponding to the corresponding pixel in the first depth map is extracted from a second depth map. 2,i ;
[0011] The method is particularly suitable for processing a first depth map and a second depth map together, both of which provide depth information for an image of a single camera. The first depth map and the second depth map are preferably created using different algorithms based on the same camera image and preferably with different additional data added.
[0012] For each image point in the depth map, a depth is particularly preferably specified. The depth describes the distance between an object near the camera and the object visible in the corresponding image point, or the distance between the object described by the image point and the camera. In other words: The depth describes the distance between a scene point and the camera. An image point is a projected 2D point (pixel) in the image. A scene point is a 3D point in the world at a certain distance from the camera. For example, the distance between a scene point and the camera can be related to the direct distance from the camera. Thus, this means measuring the distance along the viewing direction from the camera to the image point. In each case, the definition of the depth is related to the camera plane in which the camera is located, which has a defined alignment relative to the camera alignment, which is conceivable and is also included here, but is not advisable. Such a camera plane is preferably aligned normal to the axis of the camera.
[0013] Various algorithms and methods for creating depth maps for camera images can be applied to different depth maps. Since depth information can usually only be estimated relatively (i.e., compared with other depth information) using conventional algorithms, there will usually be different scalings between these depth maps. The method described herein is used to describe an algorithm using which a scaling factor describing such a scaling can be found, enabling the combination of depth maps.
[0014] If the aim is to evaluate depth information or depth maps created based on camera images from different cameras or camera images recorded by a single camera with a time offset, this method can also be used in principle.
[0015] In addition, pixel assignment also needs to be performed for various camera images. Otherwise, a comparison of depth information at the pixel level cannot be performed.
[0016] The algorithm for creating the depth map provides a depth value for each pixel of a single image. To generate the depth information of the depth map, such an algorithm may extract additional images that have been recorded previously or later using the same camera.
[0017] The creation of the first depth map and the creation of the second depth map may be different because different additional images are used to create the respective depth maps.
[0018] In principle, there may be deviations in the scaling of the depth map created from the camera image, mainly because cameras typically capture images using a monocular optical system. The information captured using a monocular optical system initially has no scale, and to obtain depth information, an external reference is required. The external reference may be generated by further movement of the vehicle because if the vehicle moves further, the images of the same camera in the camera image show visible objects from different angles. Therefore, a depth map can be created using camera images recorded by the same camera either beforehand or later. For this purpose, data describing the movement of the vehicle is usually also required. For example, these additional data can be received via the vehicle's bus system, and these data can be obtained from the drive system, such as the vehicle's navigation data and / or movement data. In appropriate cases, data from additional cameras attached to the vehicle at different positions or completely different sensors (such as LIDAR or RADAR sensors arranged in the camera area) can also be used as additional data for creating the depth map.
[0019] To create a depth map based on the camera image of a monocular optical system, filters or neural networks generated using deep learning can generally be used particularly effectively, and these filters or neural networks can optionally be used in the form of data filters. However, there is a fundamental problem with these filters, that is, the training data suitable for training such filters often only generates relative depth information. If this method is tried to obtain absolute depth information, the algorithm usually varies with respect to the camera mounting position, and if the camera mounting position changes, theoretically, the algorithm must be completely retrained.
[0020] In principle, since cameras with monocular optical systems are widely used, the method described in this article has many application possibilities in finding the scaling factor for the scaling between two depth maps.
[0021] This article proposes to determine the ratio of each pixel of the depth map to the comparison depth value, where the comparison depth value is preferably the second depth value of the corresponding pixel of the second depth map. This multiple ratio is basically the individual pixel scaling factor calculated in step c). This results in a scaling factor map that regularly shows different scaling factors distributed across the entire depth map.
[0022] Subsequently, the frequency distribution of the ratios determined in this way is decided (step d). Based on this frequency distribution, it is now possible to decide the global scaling factor (overall scaling factor) of the entire first depth map with respect to the comparison depth value (especially with respect to the second depth map). The frequency distribution is evaluated according to step e) to decide the scaling factor.
[0023] Compared with other methods that directly determine the overall scaling factor from a single depth map, this method avoids the distortion of the determination of the scaling factor due to the specific inaccuracies of the scale of a single depth map canceling each other out. Due to the method described herein, in any case, the scale of each pixel of the depth map is correct. The frequency distribution makes the fluctuations of the scaling factor visible across the entire depth map. The frequency distribution is also commonly referred to as a histogram in this article.
[0024] This method is particularly preferred if in each case a group and / or subset of actual image pixels is considered as pixel i for performing the method steps.
[0025] Then the depth value d of this group of image pixels is determined 1,i , and the comparison depth value d is determined 2,i , and their ratio is determined For example, such a group of image pixels can be defined as n by n image pixels, such as 5x5 image pixels. In an embodiment variant, first (before performing the method steps) an average / general depth value d1 and an average / general comparison depth value d can be determined for such a group of image pixels 2,i , and then these values are processed in each case according to the method.
[0026] According to this embodiment variant of the method, the computational overhead of performing the method can be reduced. In a further embodiment variant of the method, it is also possible to optionally only consider a subset of the pixels, preferably pixels with the same spatial distribution. In appropriate cases, every X pixels (for example, every 10 pixels) or a group of pixels in each case are considered as one pixel to perform the method described herein and are examined as one pixel.
[0027] The method described herein describes a method that can also be referred to as a "histogram-based" method and can be used to find the scaling factor for scaling between depth maps. In the case of ideal data, the ratio extracted from a single pixel is satisfactory. However, various errors may occur during the creation of depth maps, so performing statistics on the entire image is the most robust method.
[0028] Preferably, a so-called Gaussian mixture model is used to determine the scaling factor of the method. In this way, the superposition of multiple Gaussian distributions is used to approximate the histogram. By determining the mean of the distribution, the presence of a single peak relative to multiple strong peaks can be determined. The standard deviation provides information about the gradient of the curve and thus information about the reliability of the estimate.
[0029] It is also preferred if at least one scaling parameter is determined as at least one parameter of the frequency distribution in step e).
[0030] A conventional frequency distribution applicable here is, for example, the Gaussian distribution, also known as the normal distribution, which has been further described above. It is preferred if in step e) at least one normal distribution approximating the frequency distribution determined in step d) is determined. The normal distribution has various descriptive parameters, especially the expected value or mean and the variance.
[0031] It is further preferred if at least one scaling parameter is determined as the mean of the normal distribution.
[0032] If the scaling of both depth maps or the scaling of the depth map and the comparison depth value is consistent, then the frequency distribution will correspond to a sharp normal distribution with a high and significant mean and a narrow standard deviation.
[0033] It is also preferred if in step e) a confidence measure for at least one determined scaling factor is determined from the standard deviation of the normal distribution.
[0034] By determining the scaling factor on the normal distribution, a parameter with a standard deviation can be directly obtained, which is very suitable as a confidence measure for the scaling factor and can even be directly used for this purpose.
[0035] Furthermore, it is preferred if in step e) an examination is made to determine whether the frequency distribution has multiple extrema, where the multiple extrema in step e) are included in the evaluation, and based on this evaluation, the accuracy of the depth information of the first depth map and / or the second depth map is evaluated.
[0036] Furthermore, in step e), if the frequency distribution has multiple extrema, it is preferred to map the frequency distribution to multiple normal distributions.
[0037] Furthermore, it is preferred if at least one determined scaling factor in step e) is a relative scaling factor, and using this relative scaling factor, the scaling of the first depth value of at least one first depth map is comparable to the scaling of the comparison depth value, where the method further includes determining an absolute reference scale for the depth values of at least one depth map.
[0038] It is also preferred if the absolute reference scale in step e) is determined by converting the absolute reference scale of the comparison depth map with at least one scaling factor.
[0039] Furthermore, it is preferred if the absolute reference scale is determined according to at least one of the following methods:
[0040] - Determining the calibration base width of two cameras and creating a first depth map using the calibration base width by the stereoscopic parallax method;
[0041] - Using a dimension reference within the first depth map;
[0042] - Analyze at least one first depth map using a neural network to determine an absolute reference scale.
[0043] The main application possibility of the method proposed in this paper is depth map fusion. After the depth maps are appropriately scaled, depth map fusion algorithms can be applied.
[0044] The method described in this paper realizes the fusion of different depth maps obtained by photogrammetry. These depth maps are obtained from image pairs at different time intervals, which means that the optical flow is calculated based on images at different time intervals from the current image.
[0045] Regarding photogrammetry methods, dynamic objects, low autonomous control vehicle speeds, occlusions, and repetitive structures degrade the quality of depth maps. Therefore, it is advisable to use a neural network filter, which in principle can automatically minimize this effect.
[0046] The fusion of these depth maps greatly improves the performance of geometric computer vision algorithms, such as local surface normals, road surface estimation, and the recognition of objects at low heights.
[0047] Furthermore, the application of relative scaling is not limited to depth map fusion. When the depth map serves as a true reference ("ground truth"), the introduced relative scaling technique can also be used for evaluation purposes. The true reference ("ground truth") can be obtained from, for example, LIDAR or stereo data, and the depth map to be evaluated can be a depth map obtained by photogrammetry or depth from a monocular. Therefore, in addition to a simple pixel-by-pixel comparison of depth values, the scale difference is also evaluated.
[0048] Similarly, the aim of this paper is to describe a device for data processing, which includes a processor adapted / configured to implement the method.
[0049] The device can in particular be a control device and / or module in an overall motor vehicle system, which is configured to perform highly automated and possibly even autonomous driving functions. Preferably, depth maps processed by this method are created based on sensor data (especially based on camera data) processed by the overall system to provide parameters and / or control variables for providing highly automated driving functions.
[0050] Similarly, the aim is to claim a computer program product including instructions that, when the computer program product is executed by a computer, cause the computer to perform the method.
[0051] Similarly, the aim is to describe a computer-readable storage medium including instructions that, when the storage medium is executed by a computer, cause the computer to perform the method. Description of the Drawings
[0052] The present invention and its technical field will be described in more detail below with reference to the accompanying drawings. The drawings illustrate preferred exemplary embodiments that do not limit the present invention. It should be specifically noted that the dimensional ratios in the drawings, especially in the figures, are purely schematic. In the figures:
[0053] Figure 1 : schematically shows the inconsistent fusion of depth maps;
[0054] Figure 2 : schematically shows the depth map fusion performed by the method described herein;
[0055] Figure 3 : shows a flowchart of the method; and
[0056] Figures 4a to 4c : shows various frequency distributions that can be examined within the scope of the method. Detailed Description of the Invention
[0057] Figure 1 Shows examples of depth maps 11 that are shifted or scaled differently relative to each other. These depth maps 11 are created using cameras 10, and different algorithms may be used to create the depth maps from the original camera data. If appropriate, two depth maps 11 that are shifted relative to each other may also be generated from different original data. Figure 2 Further shows the result of the inconsistent fused depth map 12. Reliable information is generally not obtainable from such an inconsistent fused depth map 12.
[0058] In contrast, Figure 2 Shows an example according to which two depth maps 11 that are shifted or scaled differently relative to each other are fused to form a scaled fused depth map 13 by the method described herein. For example, in Figure 2 the lower region of the Figure 1 first depth Figure 2 and second depth
[0059] Figure 3 are plotted with a scaling factor 14. If at least one suitable scaling factor 14 is determined among these scaling factors 14 using the method, two differently scaled depth maps 11 can be fused to form a scaled fused depth map 13 from which reliable information can be obtained. Figure 1 is shown schematically at the top. From this, a first depth value 3 can be extracted. A second depth Figure 2 can also be identified, from which a second depth value 4 can be determined. This second depth value is intended to be used as a comparison depth value in the sense of the method. This corresponds to steps a) and b) of the method. Subsequently, the first depth value 3d for each pixel i is shown1,i with the comparison depth value d 2,i the ratio between the frequency distribution 5. Calculate the ratio according to step c) Determine the frequency distribution 5 according to step d). Preferably, the scaling factor is determined based on the statistical parameters of the frequency distribution 5. It is particularly suitable to approximate the frequency distribution 5 as a normal distribution 6, so the scaling factor can be assumed to be the mean value 7 of the normal distribution 6. The standard deviation 8 of the normal distribution 6 can be output as a confidence measure of the determined scaling factor.
[0060] Figures 4a to 4c Now different forms of the frequency distribution 5 are shown, which can be checked within the scope of the method. Figure 4a The frequency distribution 5 is shown, having an extreme value 9 of the frequency that can be well approximated by a normal distribution 6 and the scaling factor can be reliably determined.
[0061] Figure 4b The situation where multiple peaks / extremes 9 of the frequency are highly similar is shown. This indicates that continuous fusion may be impossible. This may occur when the algorithm depth reconstruction for adding depth information in the camera image does not work as expected. This configuration is a very difficult error to correct because there are multiple equivalent solutions, or the depth Figure 1 depth Figure 2 The depth information of different regions can also be scaled differently. If necessary, this situation can be used as a reason to reject the depth Figure 1 depth Figure 2 from further processing.
[0062] Figure 4c Another situation is shown, indicating that at least one of the depth Figure 1 depth Figure 2 has impaired quality. In this case, the significance of the extreme value 9 is very low. Here, the low confidence of the scaling factor determined based on the standard deviation can be used by the method. This situation may indicate that the relative scale of the camera 10 used to create the depth Figure 1 depth Figure 2 is not well defined. This effect may be caused by the rolling shutter effect, motion blur, or low depth reconstruction accuracy (possibly due to poor optical flow quality).
Claims
1. A method for determining at least one scaling factor of a first depth map (1), the method having the following steps: a) Extract the first depth value (3) d of multiple pixels i of the first depth map (1) 1,i ; b) Determine each first depth value (3) d 1,i of the comparison depth value (4) d 2,i ; c) Calculate the ratio between the first depth value (3)d of each pixel i 1,i and the comparison depth value (4)d 2,i d) Determine the ratio S r,i The frequency distribution (5) over all pixels i; and e) Determine at least one scaling factor based on a frequency distribution (5).
2. The method according to claim 1, wherein in step b), a comparison depth value d of a pixel corresponding to the corresponding pixel i in the first depth map (1) is extracted from the second depth map (2). 2,i .
3. The method according to claim 1, wherein A group and / or subset of actual image pixels are each considered as pixel i for performing the method steps.
4. The method according to claim 1, wherein in step e), at least one scaling parameter is determined as at least one parameter of at least one frequency distribution (5).
5. The method according to claim 4, wherein at least one scaling parameter is determined based on at least one mean value (7) of at least one normal distribution (6).
6. The method according to claim 5, wherein, In step e), a confidence measure of the determined at least one scaling factor is determined from the standard deviation (8) of at least one normal distribution (6).
7. The method according to claim 1, wherein in step e), a check is made to determine whether the frequency distribution (5) has multiple extrema (9), and wherein the multiple extrema (9) in step e) are included in an evaluation, and based on the evaluation, the accuracy of the depth information of the first depth map (1) and / or the second depth map (2) is evaluated.
8. The method according to claim 7, wherein In step e), in the case where the frequency distribution (5) has multiple extrema (9), the frequency distribution (5) is simulated with multiple normal distributions (6).
9. The method according to claim 1, wherein at least one scaling factor determined in step e) is a relative scaling factor, by which the scaling of the first depth value (3) of at least one first depth map (1) is comparable to the scaling of a comparison depth value (4), and wherein the method further includes determining an absolute reference scale for the depth values (3, 4) of at least one depth map (1, 2).
10. The method according to claim 9, wherein the absolute reference scale in step e) is determined by converting the absolute reference scale of a comparison depth map (2) with at least one scaling factor.
11. The method according to claim 9, wherein the absolute reference scale is determined according to at least one of the following methods: - Determine the calibration base width of two cameras (10), use the calibration base width and create a first depth map (1) by stereoscopic parallax method; - Use a dimension reference within the first depth map (1); - Use a neural network to analyze at least one first depth map (1) to determine the absolute reference scale.
12. A device for data processing, including a processor adapted / configured to execute the method according to claim 1.
13. A computer program product, including instructions which, when the computer program product is executed by a computer, cause the computer to execute the method according to claim 1.
14. A computer-readable storage medium, including instructions which, when the instructions are executed by a computer, cause the computer to execute the method according to claim 1.