Head-Mounted Display Depth Prediction for Distant XR Objects
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing depth-based XR applications are constrained by the maximum measurable depth range of their depth sensors, limiting user interaction with distant objects and hindering immersive outdoor experiences.
Innovation Solution
A head-mounted display system that includes a geospatial information module, image projector, domain transfer module, ranking module, and depth optimizer, which uses machine learning and image processing to generate precise depth information for distant objects by obtaining location and pose, performing image processing, and projecting 3D meshes to virtual planes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If depth sensors are used to measure depth in XR applications, then depth information can be obtained for nearby objects, but the maximum measurable depth range is limited
Solution Approach 1:
The patent introduces street view images and 3D mesh data as intermediary elements between the depth sensor and distant objects. Instead of directly measuring depth of distant objects with limited-range sensors, the system uses pre-captured street view images from databases and corresponding 3D mesh models as mediators to infer depth information for objects beyond the sensor's direct measurement capability.
Solution Approach 2:
The system creates a copy of the real-world environment using pre-captured street view images and 3D mesh models stored in databases. By matching features between the current camera view and these pre-captured images, the system can retrieve depth information from the 3D mesh copy, effectively extending the measurable depth range without requiring direct sensor measurement of distant objects.
2Adaptability or versatility
If more modules are added to extend depth range, then depth information for distant objects can be obtained, but device complexity increases
Solution Approach 1:
The patent implements a unified processing pipeline where the image processor performs multiple functions: it processes real-time camera images, compares them against street view images from databases, performs feature matching, and retrieves depth information from 3D mesh models. This multi-functional approach extends depth range without requiring separate dedicated hardware modules for each function, thereby managing system complexity.
Solution Approach 2:
The system uses pre-captured street view images and 3D mesh models that are already stored in databases, eliminating the need for additional active sensing hardware or complex real-time processing modules. The existing camera and image processor leverage these pre-prepared resources to extend depth capability, allowing the system to serve itself by utilizing readily available data rather than requiring additional specialized components.
Data Source
AI summary
A head-mounted display and a method for depth prediction are provided. The method includes: obtaining location information and a pose of the head-mounted display; obtaining a first street view from a database according to the location information and the pose; performing image processing on a first image captured by the head-mounted display; determining whether the processed first image matches the first street view; generating depth information for an image segment according to the first street view in response to the processed first image matching the first street view; and outputting the depth information.


