High-precision depth map acquisition method and device

Through a method based on the deep-supervised unbounded neural radiation field model, combined with multi-view image information and neural network model, the problem of insufficient accuracy in complex scenarios of traditional three-dimensional reconstruction methods is solved, and high-precision three-dimensional reconstruction and depth map generation are realized, providing an important data set for automatic driving systems.

CN120147416APending Publication Date: 2025-06-13ZHONGKE HUIYAN (TIANJIN) ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510185563.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional three-dimensional reconstruction methods are difficult to meet the needs of high accuracy and accuracy in complex scenarios, and it is difficult to obtain the depth truth data of large outdoor scenes.

Method used

Using a method based on the deep-supervised unbounded neural radiation field (DS-MipNeRF360) model, combining multi-view image information and neural network model, high-precision three-dimensional features are rendered through monocular 2D feature images, and 2D feature maps and depth maps of any specified view angle are synthesized.

Benefits of technology

High-precision three-dimensional reconstruction of large-scale outdoor scenes without boundary is realized, the generated depth map accuracy reaches ideal indicators, solves the problem of difficulty in obtaining depth truth data, and provides important data sets for training and evaluation of autonomous driving systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147416A_ABST
    Figure CN120147416A_ABST
Patent Text Reader

Abstract

The invention discloses a high-precision depth image acquisition method and device, which are used for synthesizing a new visual angle image with high fidelity and authenticity and a corresponding high-precision depth image based on a small number of real scene images. The method comprises the following steps: acquiring an original video file to be processed, and inputting the original video file into a COLMAP model for preprocessing to obtain a processed image and pose information; and inputting the processed image and pose into a depth supervision unbounded neural radiation field network for training, and obtaining a new view and a high-precision depth map corresponding to the new view.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of assisted driving, and particularly to a method and device for obtaining high-precision depth maps. Background Art

[0002] With the continuous development of autonomous driving technology, the demand for environmental perception and scene understanding is increasing day by day. As an important environmental perception data, depth maps play a key role in autonomous driving systems. However, traditional 3D reconstruction methods are limited by accuracy and precision, and it is difficult to meet the requirements in complex scenarios. Therefore, the present invention proposes a method based on the DeepSDF-MipNeRF360 model, which is an improvement on the MipNeRF360[1] model, and its effect is verified to be superior to the MipNeRF360 model through a large number of experiments. This method combines multi-view image information and neural network models, and renders high-precision and realistic 3D features through monocular 2D feature images with limited views, realizing high-precision 3D reconstruction of unbounded outdoor large scenes, and synthesizing 2D feature maps of any specified view and depth maps under that view. The reconstruction of 3D features can better reflect the position, shape, and size of objects in the environment, which helps autonomous vehicles make more accurate decisions, avoid collisions and accidents, and improve overall driving safety.

[0003] In addition, the development of autonomous driving technology requires a large amount of data to train and verify algorithm models, including image data and annotation data. However, traditional manual image acquisition is costly and requires a large amount of time and human resources. And the acquisition of depth ground truth data for large outdoor scenes is a challenge in the field of autonomous driving. Currently, depth ground truth data is mainly obtained through lidar scanning, stereo cameras, and depth cameras. Lidar scanning calculates depth information by obtaining 3D point cloud data; stereo cameras take pictures of the scene from multiple angles and calculate depth using parallax information; depth cameras generate depth images based on technologies such as time-of-flight and structured light, but depth cameras only work well for small indoor scenes and are not suitable for large outdoor scenes.

[0004] In view of this, the present invention is proposed. Summary of the Invention

[0005] The main purpose of the present invention is to disclose a method and device for obtaining high-precision depth maps, which are used to synthesize new view images with high fidelity and authenticity, as well as corresponding high-precision depth images, based on a small number of real scene images.

[0006] To achieve the above object, according to one aspect of the present invention, a method for obtaining high-precision depth maps is provided, and the following technical solutions are adopted:

[0007] The high-precision depth map acquisition method includes: obtaining the original video file to be processed, inputting the original video file into the COLMAP model for preprocessing to obtain the processed image; inputting the processed image into the depth-supervised unbounded neural radiance field network for training to obtain new views and their corresponding high-precision depth maps.

[0008] Further, before obtaining the original video file to be processed, the high-precision depth map acquisition method further includes: constructing a depth-supervised unbounded neural radiance field network: adding a depth supervision signal based on the synthesizable scene model MipNeRF360 to obtain the DS-MipNeRF360 model; performing high-precision three-dimensional reconstruction and data generation based on the DS-MipNeRF360 model to obtain high-precision depth maps.

[0009] Further, the inputting the original video file into the COLMAP model for preprocessing includes: extracting frames from the video file through the ffmpeg tool, and calculating the pose information of the monocular image from the extracted image frames through the COLMAP model.

[0010] Further, the inputting the processed image into the depth-supervised unbounded neural radiance field network for training to obtain new views and their corresponding high-precision depth maps includes: inputting the extracted image frames and their corresponding pose information into the DS-MipNeRF360 model; calculating the model depth accuracy of the input image frames and their corresponding pose information to obtain new perspective views, new perspective depth maps, and new perspective point cloud files.

[0011] Further, the calculating the model depth accuracy of the input image frames and their corresponding pose information further includes: model reliability verification: based on the monocular camera, the transformation from pixel coordinates to the world coordinate system: (1) the relationship from the camera coordinate system to the pixel coordinate system:

[0012] (1)

[0013] Obtaining the transformation from pixels to the camera coordinate system:

[0014] (2)

[0015] where, according to the COLMAP principle and the camera calibration principle:

[0016] (3)

[0017] Let be represented by the model output depth and then substitute it into Equation (2) to obtain the coordinates in the camera coordinate system

[0018] The model depth is obtained according to similar triangles and the Pythagorean theorem and The relationship is as follows:

[0019] (4)

[0020] We get:

[0021] (5)

[0022] Substitute equation (5) into equation (2) to obtain the coordinates in the camera coordinate system

[0023] (2) Then, from the relationship between the camera coordinate system and the world coordinate system:

[0024] (6)

[0025] We obtain the transformation from the camera coordinate system to the world coordinate system:

[0026] (7)

[0027] where is the rotation matrix for the transformation from the camera coordinate system to the world coordinate system,

[0028] are respectively the representations of the axis of the camera coordinates in the world coordinate system; is the translation vector for the transformation from the camera coordinate system to the world coordinate system, which is the coordinate representation of the center of the camera coordinate system in the world coordinate system;

[0029] Calculate the distance between two points in space, i.e., the length, according to the Euclidean distance formula:

[0030] (8)

[0031] Calculate the depth accuracy from the length predicted by the model and the true length of the object:

[0032] Use the root mean square error (RMSE) and mean absolute percentage error (MAPE) metrics to calculate the accuracy of the depth;

[0033] RMSE:

[0034] (9)

[0035] where is the predicted value, is the true value. When represents a perfect model, the smaller the RMSE, the smaller the error;

[0036] MAPE:

[0037] (10)

[0038] wherein, is the predicted value, is the true value. When represents a perfect model, the smaller the MAPE, the smaller the error.

[0039] According to another aspect of the present invention, a high-precision depth map acquisition device is provided, and the following technical solutions are adopted:

[0040] A high-precision depth map acquisition device includes: a preprocessing module, configured to obtain an original video file to be processed, input the original video file into the COLMAP model for preprocessing, and then obtain the processed image and pose information; a training module, configured to input the processed image and pose into a depth-supervised unbounded neural radiance field network for training to obtain a new view and its corresponding high-precision depth map.

[0041] Further, the high-precision depth map acquisition device further includes: a construction module, configured to construct a depth-supervised unbounded neural radiance field network: adding a depth supervision signal based on the synthesizable scene model MipNeRF360 to obtain a DS-MipNeRF360 model; performing high-precision three-dimensional reconstruction and data generation based on the DS-MipNeRF360 model to obtain a high-precision depth map.

[0042] Further, the preprocessing module includes: a calculation module, configured to extract frames from the video file through the ffmpeg tool, and calculate the pose information of the monocular image from the extracted image frames through the COLMAP model.

[0043] Further, the training module includes: an input module, configured to input the extracted image frames and their corresponding pose information into the DS-MipNeRF360 model; a depth calculation module, configured to calculate the model depth accuracy of the input image frames and their corresponding pose information to obtain a new view view, a new view depth map, and a new view point cloud file.

[0044] The method proposed by the present invention can synthesize new perspective images with high fidelity and authenticity, as well as corresponding high-precision depth images, based on a small number of real-scene images. After inspection, the depth accuracy reaches the ideal index, so the synthesized new perspective depth map can be used as a depth ground truth image dataset. These datasets are very important for training deep learning models, conducting algorithm tests and evaluations, and are the basic resources for autonomous driving research and development. For example, the depth dataset can be used to supervise the training of perception algorithms, help train neural network models to improve the accuracy and robustness of perception; it can also be used to enhance object detection and recognition, enabling the autonomous driving system to more accurately distinguish and identify target objects such as vehicles, pedestrians, and traffic signs in front; in addition, the high-precision depth ground truth image can be used for multi-sensor data fusion, serving as a reference benchmark for sensors such as lidar, depth cameras, and stereo cameras to help fuse different sensor data. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0046] Figure 1 Flowchart of a method for obtaining a high-precision depth map according to an embodiment of the present invention;

[0047] Figure 2 Model framework diagram according to an embodiment of the present invention;

[0048] Figure 3 Another flowchart of a method for obtaining a high-precision depth map according to an embodiment of the present invention;

[0049] Figure 4 Monocular depth conversion diagram according to an embodiment of the present invention;

[0050] Figure 5 Structural diagram of a device for obtaining a high-precision depth map according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The following will detail the embodiments of the present invention with reference to the accompanying drawings. However, the present invention can be implemented in many different ways defined and covered by the claims.

[0052] Figure 1 Flowchart of a method for obtaining a high-precision depth map according to an embodiment of the present invention.

[0053] The method for obtaining a high-precision depth map includes:

[0054] S101: Obtain the original video file to be processed, input the original video file into the COLMAP model for preprocessing, and then the processed images and pose information can be obtained;

[0055] S103: Input the processed images and poses into the deep supervised unbounded neural radiance field network for training to obtain new views and their corresponding high-precision depth maps.

[0056] Specifically, obtain the original video file to be processed, put the original video into the COLMAP model for preprocessing, and then the processed images can be obtained; among them, the COLMAP model is a sparse reconstruction model, which can preprocess the image data with fixed features. In this way, the data obtained after preprocessing is put into the deep supervised unbounded neural radiance field network for training, and finally high-quality new views and their corresponding high-precision depth maps are obtained. The high-precision depth maps can be used as the depth ground truth of the RGB maps. This method can solve the problem of difficult manual image acquisition in the current autonomous driving field, alleviate the problem of difficult acquisition of depth ground truth in real scenes, and perform high-precision 3D reconstruction of autonomous driving scenes.

[0057] Figure 2 This is the model framework diagram described in the embodiment of the present invention.

[0058] See Figure 2 As shown, this method mainly includes two aspects of innovation:

[0059] First, based on the MipNeRF360 model, a depth supervision signal is added to propose the DS-MipNeRF360 model. Depth supervision is a training strategy that mainly introduces additional supervision signals in the middle layer of the network. This can promote the network to learn more useful features faster, thereby accelerating the model training process, enabling the network to reach higher performance in a shorter time. In addition, depth supervision can make the model more robust to noisy data and thus make reasonable predictions. Generally speaking, depth supervision can promote feature learning, accelerate training convergence, and improve model robustness. Such a training method has a significant effect on improving the performance of the model.

[0060] Second, high-precision 3D reconstruction and data generation are performed based on the DS-MipNeRF360 model. First, the present invention realizes realistic 3D reconstruction of unbounded outdoor large scenes, which helps autonomous vehicles perceive the driving environment, that is, estimate the positions, shapes, and sizes of objects in the environment, so as to make more accurate driving decisions, avoid collisions and accidents, and improve overall driving safety. In addition, the present invention can generate high-precision depth data, thus circumventing the problem that it is difficult to obtain the true depth values of current outdoor large scenes. The obtained depth dataset is very important for training deep learning models, conducting algorithm tests and evaluations, and is a basic resource for autonomous driving research and development. Through a large number of experiments and according to the calculation of the model depth accuracy in the following text, the reliability and authenticity of the depth dataset generated by this model are confirmed, indicating that the model proposed by the present invention can indeed produce the dataset.

[0061] Figure 3 It is a flowchart of another method for obtaining a high-precision depth map according to the embodiment of the present invention.

[0062] See Figure 3 As shown, in the invention process, the following steps are included:

[0063] Step 30: Data acquisition;

[0064] Step 31: Data preprocessing;

[0065] Step 32: DS-Mipnerf360 model training;

[0066] Step 33: Depth accuracy estimation.

[0067] More specifically, in terms of data acquisition, the shooting device: a monocular camera. Experimental tools: calibration board, tape measure. Shooting method: Slowly and uniformly surround the target scene to shoot a video. Then preprocess the data.

[0068] (1) Extract frames from the video file through the ffmpeg tool, and the extracted image frames have no artifacts or distortion phenomena.

[0069] (2) Calculate the pose information of the monocular image from the extracted image frames through the COLMAP model.

[0070] Then train the model, including: Model input: the extracted image frames and their corresponding pose information;

[0071] Model output: new view view, new view depth map, new view point cloud file;

[0072] Model type: supervised learning model, generative model.

[0073] Based on the monocular camera, the conversion from pixel coordinates to the world coordinate system:

[0074] (1) Relationship from camera coordinate system to pixel coordinate system:

[0075] (1)

[0076] Obtain the transformation from pixel to camera coordinate system:

[0077] (2)

[0078] Among them, according to the COLMAP principle and the camera calibration principle:

[0079] (3)

[0080] Since the depth output by the unbounded anti-aliasing neural radiance field model is Figure 4 in the length referred to, rather than the length in formula (2) (i.e., Figure 4 in the length), so needs to be expressed using the model output depth , and then substitute it into formula (2) to obtain the coordinates in the camera coordinate system

[0081] Figure 4 is the monocular depth conversion map described in the embodiment of the present invention; see Figure 4 the shown monocular depth conversion

[0082] According to similar triangles and the Pythagorean theorem, the relationship between the model depth and is as follows:

[0083] (4)

[0084] Obtain:

[0085] (5)

[0086] Substitute formula (5) into formula (2) to obtain the coordinates in the camera coordinate system

[0087] (2) Then, the relationship between the camera coordinate system and the world coordinate system:

[0088] (6)

[0089] Obtain the conversion from the camera coordinate system to the world coordinate system:

[0090] (7)

[0091] Among them, is the rotation matrix for the transformation from the camera coordinate system to the world coordinate system,

[0092] are respectively the representations of the axes of the camera coordinates in the world coordinate system; is the translation vector for the transformation from the camera coordinate system to the world coordinate system, and is the coordinate representation of the center of the camera coordinate system in the world coordinate system.

[0093] Calculate the distance between two points in space according to the Euclidean distance formula, that is, the length:

[0094] (8)

[0095] Use the Root Mean Square Error (RMSE) and Mean Absolute Percentage Error (MAPE) metrics to calculate the accuracy of the depth.

[0096] RMSE:

[0097] (9)

[0098] Among them, is the predicted value, is the true value. When represents a perfect model, the smaller the RMSE, the smaller the error;

[0099] MAPE:

[0100] (10)

[0101] Among them, is the predicted value, is the true value. When represents a perfect model, the smaller the MAPE, the smaller the error;

[0102] Statistically analyze the mean and variance of the depth accuracy of the model synthesizing all new views in each scenario to analyze the performance of the DS-MipNeRF360 model in each scenario. Visualize the difference between the length estimated by the model and the true physical length with a line chart to more intuitively understand the depth accuracy and the performance of the model.

[0103] Figure 5 This is the structural diagram of a high-precision depth map acquisition device described in the embodiments of the present invention.

[0104] See Figure 5As shown in the figure, the high-precision depth map acquisition device includes: a preprocessing module 50, which is used to obtain the original video file to be processed, input the original video file into the COLMAP model for preprocessing, and then the processed image can be obtained; a training module 52, which is used to input the processed image into the depth-supervised unbounded neural radiance field network for training to obtain a new view and its corresponding high-precision depth map.

[0105] Preferably, the high-precision depth map acquisition device further includes: a construction module (not shown in the figure), which is used to construct a depth-supervised unbounded neural radiance field network: adding a depth supervision signal based on the synthesizable scene model MipNeRF360 to obtain the DS-MipNeRF360 model; performing high-precision three-dimensional reconstruction and data generation based on the DS-MipNeRF360 model to obtain a high-precision depth map.

[0106] Preferably, the preprocessing module 50 includes: a calculation module (not shown in the figure), which is used to extract frames from the video file through the ffmpeg tool, and calculate the pose information of the monocular image from the extracted image frames through the COLMAP model.

[0107] Preferably, the training module 52 includes: an input module (not shown in the figure), which is used to input the extracted image frames and their corresponding pose information into the DS-MipNeRF360 model; a depth calculation module, which is used to calculate the model depth accuracy of the input image frames and their corresponding pose information to obtain a new perspective view, a new perspective depth map, and a new perspective point cloud file.

[0108] The method proposed by the present invention can synthesize new perspective images with high fidelity and authenticity, as well as corresponding high-precision depth images, based on a small number of real-scene images. After testing, the depth accuracy reaches the ideal index, so the synthesized new perspective depth map can be used as a depth ground truth image dataset. These datasets are very important for training deep learning models, conducting algorithm tests and evaluations, and are the basic resources for autonomous driving research and development. For example, the depth dataset can be used to supervise the training of perception algorithms, help train neural network models to improve the accuracy and robustness of perception; it can also be used to enhance object detection and recognition, enabling the autonomous driving system to more accurately distinguish and identify target objects such as vehicles, pedestrians, and traffic signs in front; in addition, the high-precision depth ground truth image can be used for multi-sensor data fusion, serving as a reference benchmark for sensors such as lidar, depth cameras, and stereo cameras to help fuse different sensor data.

[0109] Only some exemplary embodiments of this embodiment are described by way of illustration. Without doubt, for those of ordinary skill in the art, the described embodiments can be modified in various different ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. A method for obtaining a high-precision depth map, characterized in that: include: Get the original video file to be processed, input the original video file into the COLMAP model for preprocessing, and then get the processed image and posture information; The processed images and poses are input into the deep supervised unbounded neural radiance field network for training to obtain new views and their corresponding high-precision depth maps.

2. The high-precision depth map acquisition method according to claim 1, characterized in that: Before obtaining the original video file to be processed, the high-precision depth map acquisition method further includes: Construct a deep supervised unbounded neural radiance field network: Based on the synthesizable scene model MipNeRF360, add deep supervision signals to obtain the DS-MipNeRF360 model; Based on the DS-MipNeRF360 model, high-precision 3D reconstruction and data generation are performed to obtain a high-precision depth map.

3. The high-precision depth map acquisition method according to claim 2, characterized in that: The inputting of the original video file into the COLMAP model for preprocessing comprises: The ffmpeg tool is used to extract frames from the video file, and the extracted image frames are calculated through the COLMAP model to obtain the pose information of the monocular image.

4. The high-precision depth map acquisition method according to claim 3, characterized in that: The processing of the image and the pose is input into the deep supervised unbounded neural radiance field network for training to obtain a new view and its corresponding high-precision depth map, including: The extracted image frames and their corresponding pose information are input into the DS-MipNeRF360 model; The model depth accuracy is calculated for the input image frame and its corresponding pose information to obtain a new perspective view, a new perspective depth map, and a new perspective point cloud file.

5. The high-precision depth map acquisition method according to claim 4, characterized in that: The calculating of the model depth accuracy of the input image frame and its corresponding pose information also includes: Model reliability verification: Transformation of pixel coordinates to world coordinate system based on monocular camera: (1) Relationship from camera coordinate system to pixel coordinate system: (1) Get the transformation from pixel to camera coordinate system: (2) Among them, according to the COLMAP principle and camera calibration principle: (3) Will Output depth with model Then substitute it into equation (2) to get the coordinates in the camera coordinate system Determine the model depth based on similar triangles and the Pythagorean theorem and The relationship is as follows: (4) have to: (5) Substituting equation (5) into equation (2) yields the coordinates in the camera coordinate system: (2) Then, based on the relationship between the camera coordinate system and the world coordinate system: (6) Get the transformation from camera coordinate system to world coordinate system: (7) in, is the rotation matrix from the camera coordinate system to the world coordinate system. They are the camera coordinates The representation of axes in the world coordinate system; is the translation vector from the camera coordinate system to the world coordinate system, and is the coordinate representation of the center of the camera coordinate system in the world coordinate system; the distance between two points in space, i.e., the length, is calculated according to the Euclidean distance formula: (8) The depth accuracy is calculated by the length of the object predicted by the model and the actual length of the object: the root mean square error (RMSE) and mean absolute error (MAPE) indicators are used to calculate the depth accuracy; RMSE: (9) in, is the predicted value, is the true value. Represents a perfect model. The smaller the RMSE, the smaller the error. MAPE: (10) in, is the predicted value, is the true value. Represents a perfect model. The smaller the MAPE, the smaller the error.

6. A high-precision depth map acquisition device, characterized in that: include: The preprocessing module is used to obtain the original video file to be processed, and input the original video file into the COLMAP model for preprocessing to obtain the processed image and posture information; The training module is used to input the processed images and poses into the deep supervised unbounded neural radiance field network for training to obtain new views and their corresponding high-precision depth maps.

7. The high-precision depth map acquisition device according to claim 6, characterized in that: Also includes: A building module is used to construct a deep supervised unbounded neural radiance field network: based on the synthesizable scene model MipNeRF360, a deep supervision signal is added to obtain the DS-MipNeRF360 model; based on the DS-MipNeRF360 model, high-precision 3D reconstruction and data generation are performed to obtain a high-precision depth map.

8. The high-precision depth map acquisition device according to claim 7, characterized in that: The preprocessing module comprises: The calculation module is used to extract frames from the video file through the ffmpeg tool, and calculate the extracted image frames through the COLMAP model to obtain the pose information of the monocular image.

9. The high-precision depth map acquisition device according to claim 8, characterized in that: The training module includes: Input module, used to input the extracted image frames and their corresponding pose information into the DS-MipNeRF360 model; The depth calculation module is used to calculate the model depth accuracy of the input image frame and its corresponding pose information to obtain a new perspective view, a new perspective depth map, and a new perspective point cloud file.