Mobile RGB Camera Depth Map Generation via Multi-Frame Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The accuracy of the final depth map is low due to the limitations of the fusion algorithm in existing technologies.
Innovation Solution
An image processing method that determines a first estimated sparse depth map and pose information for a video frame sequence acquired by a mobile RGB camera, and then corrects these estimates using additional video frames and a target sparse depth map from a depth camera, ultimately generating a dense depth map.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a depth camera with high resolution is used to obtain a dense depth map, then the depth map quality is improved, but the cost and power consumption increase
Solution Approach 1:
The patent segments the depth map acquisition process into two stages: first acquiring a sparse depth map from a single video frame, then generating a dense depth map by integrating multiple video frames and sparse depth maps. This segmentation allows the system to achieve dense depth map quality without requiring a high-resolution depth camera, thereby reducing power consumption and cost.
Solution Approach 2:
The patent transitions from a single-frame sparse depth map (2D spatial information) to a multi-frame dense depth map by adding the time dimension. By integrating depth information across multiple video frames and utilizing pose information, the system reconstructs dense 3D depth maps without requiring high-resolution depth cameras, thus resolving the contradiction between depth map quality and power consumption.
2Measurement precision
If a depth camera with high resolution is used to obtain a dense depth map, then the depth map quality is improved, but the device cost increases
Solution Approach 1:
The patent segments the depth map acquisition process into two stages: first acquiring a sparse depth map from a single video frame, then generating a dense depth map by integrating multiple video frames and sparse depth maps. This segmentation allows the system to achieve dense depth map quality without requiring a high-resolution depth camera, thereby reducing power consumption and cost.
Solution Approach 2:
The patent uses RGB video frames as a copy or alternative source of spatial information to supplement the sparse depth map. By integrating multiple RGB frames with pose information, the system reconstructs dense depth maps, effectively using cheaper RGB camera data to replace the need for expensive high-resolution depth cameras.
3Quantity of substance
If fusion algorithm is used to obtain dense depth map from sparse depth maps, then the depth map density is improved, but the accuracy decreases
Solution Approach 1:
The patent performs preliminary correction of pose information and sparse depth maps before fusion. By pre-processing the input data to correct errors in pose estimation and sparse depth map acquisition, the system ensures that the subsequent fusion process maintains high accuracy while achieving dense depth map output.
Solution Approach 2:
The patent incorporates feedback mechanisms in the fusion process by using corrected pose information and multiple video frames to iteratively refine the dense depth map. The system continuously adjusts and optimizes the depth map generation by feeding back correction information from multiple sources, thereby maintaining high accuracy while achieving density.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
This application relates to the technical field of terminals, and provides an image processing method, an electronic device, a storage medium, and a program product. The method includes: determining a first estimated sparse depth map and first estimated pose information corresponding to a first video frame in a video frame sequence; determining a first corrected sparse depth map and first corrected pose information corresponding to the first video frame based on the first video frame, the first estimated sparse depth map, the first estimated pose information, a second video frame, a second estimated sparse depth map and second estimated pose information corresponding to the second video frame, and a target sparse depth map that is synchronized with the first video frame and acquired by a depth camera; and determining a dense depth map corresponding to the first video frame based on the first video frame, the first corrected sparse depth map, the first corrected pose information, the second video frame, and a second corrected sparse depth map and second corrected pose information of the second video frame. In this way, a depth image having accurate and dense depth information can be obtained.