3D Reconstruction Using IMU and Camera Data for Dense Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing SLAM technology in augmented and virtual reality environments produces only sparse maps, which are insufficient for real-time camera position and pose recognition, leading to challenges in achieving virtual and real occlusion.
Innovation Solution
A method performed by an electronic device that obtains camera position and pose, sparse maps, and high-density maps using frame images and IMU data, involving global optimization, neural network processing, and depth map generation to enhance map density and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing SLAM technology is used to obtain camera position and pose, then the system can quickly adapt to new scenes without pre-trained models, but the map density remains sparse and insufficient for real-time camera position and pose recognition
Solution Approach 1:
The patent combines multiple data sources including frame images from cameras, IMU data (acceleration and angular velocity), depth information, and feature point data into a unified processing framework. This merging of heterogeneous data streams enables the system to generate high-density maps with sufficient detail for accurate real-time camera position and pose recognition, resolving the limitation of sparse maps in traditional SLAM.
Solution Approach 2:
The patent transitions from 2D image data to 3D spatial representation by integrating depth information and generating three-dimensional maps. This dimensional transformation allows the system to capture spatial relationships and scene structure more comprehensively, enabling accurate camera pose estimation and achieving virtual-reality occlusion effects that were impossible with sparse 2D feature points alone.
2Productivity
If a sparse map is generated from feature points, then the system can operate in real-time without pre-trained models, but virtual-reality occlusion cannot be achieved
Solution Approach 1:
The patent performs preliminary processing of IMU data through integration to obtain position and attitude information before combining it with image data. Depth information is also pre-computed and integrated into the mapping process. These preliminary actions enable the system to maintain real-time processing capability while generating sufficiently dense maps for virtual-reality occlusion, as the heavy computational work is prepared in advance rather than during final rendering.
Solution Approach 2:
The patent introduces depth information as an intermediary element that bridges the gap between sparse feature points and dense volumetric maps. This intermediate representation allows the system to maintain real-time performance while achieving the density required for virtual-reality occlusion, as depth data can be efficiently integrated into the mapping pipeline without requiring full 3D reconstruction of every scene element.
3Measurement precision
If global optimization is performed on all frames, then the accuracy of camera position and pose is improved, but the computational complexity and processing time increase significantly
Solution Approach 1:
The patent segments the video stream into key frames and non-key frames, applying global optimization selectively to key frames that contain significant scene changes or new feature points. Non-key frames use incremental updates based on IMU data and previous map information. This segmentation reduces computational complexity while maintaining accuracy, as optimization is performed only when necessary rather than on every frame.
Solution Approach 2:
The patent implements periodic global optimization at key frame intervals rather than continuously on every frame. Between key frames, the system uses incremental updates and IMU data for pose estimation. This periodic approach maintains camera position and pose accuracy over time while significantly reducing computational complexity compared to continuous optimization of all frames.
Data Source
AI summary
A method performed by an electronic device, the electronic device, and a storage medium are provided. The method includes obtaining a frame image of a video from a camera and inertia data of an inertial measurement unit (IMU) corresponding to the frame image and obtaining a camera position and pose of the camera, a sparse map, and a high-density map corresponding to the frame image, based on the frame image and the inertia data of the IMU.


