Video-Based 3D Reconstruction for Remote AR Without Depth Sensors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing augmented reality (AR) systems in remote video sessions are limited by the inability of remote users to place AR objects outside the current frame without accurate depth and camera pose data, which is often unavailable due to bandwidth constraints or device limitations.
Innovation Solution
A method for reconstructing a 3D model from a video stream using Structure from Motion techniques and machine learning to extrapolate depth information, allowing remote users to place AR objects within an expanding 3D model that can be synchronized with the video feed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If depth and camera pose data are transmitted in remote video sessions, then AR object placement accuracy is improved, but bandwidth consumption increases
Solution Approach 1:
The patent extracts only the essential visual information needed for AR object placement from the video feed, rather than transmitting complete depth and camera pose data. By using computer vision algorithms to detect and track key features directly from video frames, the system achieves adequate AR placement accuracy while significantly reducing bandwidth consumption compared to transmitting full depth maps and pose information.
Solution Approach 2:
The patent creates a simplified virtual representation of the physical environment by generating 3D models from video frames rather than transmitting actual depth sensor data. This virtual copy allows remote users to place AR objects accurately enough for practical purposes while using minimal bandwidth, as only compressed video frames and essential feature data need to be transmitted.
2Measurement precision
If depth sensors such as LiDAR are used to capture depth information, then AR object placement precision is improved, but device complexity increases
Solution Approach 1:
The patent replaces physical depth sensing hardware (LiDAR, depth cameras) with computational methods using standard video cameras and computer vision algorithms. By substituting mechanical/optical depth measurement systems with software-based structure from motion and visual odometry techniques, the system achieves comparable AR placement precision while working with simpler, more widely available device components.
Solution Approach 2:
The patent enables standard video cameras to perform depth estimation and 3D reconstruction tasks traditionally requiring dedicated depth sensors. Through self-service computational photography techniques, the system uses the video camera's inherent capabilities combined with algorithmic processing to generate depth information, eliminating the need for additional specialized hardware.
3Adaptability or versatility
If complete depth and camera pose data are available, then remote users can place AR objects anywhere in the scene, but data transmission requirements increase
Solution Approach 1:
The patent implements partial action by providing just enough depth and spatial information for effective AR object placement without transmitting complete and exhaustive scene data. The system transmits selective feature points, simplified 3D models, or estimated depth maps that enable sufficient placement flexibility for most AR applications while keeping data transmission requirements manageable through selective information provision.
Data Source
AI summary
Embodiments include systems and methods for creation of a 3D mesh from a video stream or a sequence of frames. A sparse point cloud is first created from the video stream, which is then densified per frame by comparison with spatially proximate frames. A 3D mesh is then created from the densified depth maps, and the mesh is textured by projecting the images from the video stream or sequence of frames onto the mesh. Metric scale of the depth maps may be estimated where direct measurements are not able to be measured or calculated using a machine learning depth estimation network.


