Video See-Through Depth Mapping Without Full 6 DOF Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

XR devices face challenges in supporting Video See-Through (VST) display effects due to reliance on 6 DOF data, leading to low fluency and poor user experience when 6 DOF data is incomplete or unavailable.

Innovation Solution

Implementing a method that uses preset stereoscopic shapes to provide fixed depth information, allowing VST to be realized even when full depth information is not obtained, by constructing a preset shape like a sphere and determining depth information based on this shape.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If VST technology relies on 6 DOF data for display, then the accuracy of spatial positioning is improved, but the system fails when 6 DOF data is incomplete or unavailable

Engineering Contradiction:
Improvespatial positioning accuracyVSAvoidsystem availability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent introduces a preset stereoscopic shape as an intermediary model to bridge the gap between incomplete depth information and VST display requirements. This virtual model serves as a mediator that can be constructed even when full 6 DOF data is unavailable, allowing the system to maintain functionality by relying on the predefined geometric structure rather than complete sensor data

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary construction of a stereoscopic shape model in advance, storing depth information of preset shapes before they are needed for VST display. This preparation allows the system to quickly generate depth maps without requiring real-time 6 DOF data, ensuring continuous operation even when sensor data is incomplete

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If the system uses preset stereoscopic shapes to provide fixed depth information, then VST can be realized with incomplete depth information, but the manufacturing complexity increases

Engineering Contradiction:
ImproveVST scenario compatibilityVSAvoidsystem structure complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent changes the parameter representation from requiring complete 6 DOF sensor data to using predefined geometric models with fixed depth parameters. By transforming the problem from data-dependent to model-dependent, the system achieves greater adaptability across different VST scenarios while the complexity is confined to the software layer rather than hardware

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If complete 6 DOF data is required for VST display, then the depth information accuracy is improved, but the image display fluency deteriorates when data is unavailable

Engineering Contradiction:
Improvedepth information accuracyVSAvoidimage display fluency
Core Design Contradiction:
Measurement precisionVSSpeed

Solution Approach 1:

The patent uses lightweight, computationally inexpensive preset stereoscopic shapes as temporary depth models instead of relying on expensive and time-consuming real-time 6 DOF data acquisition. These simple geometric models can be quickly constructed and updated, providing adequate depth information for VST display without compromising fluency, even though they are less accurate than complete sensor data

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Data Source

PatentUS20260067444A1Method and device for video see-through, storage medium, and program product
Publication Date: 2026.03.05 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20260067444A1 patent drawing
  • US20260067444A1 patent drawing
  • US20260067444A1 patent drawing

AI summary

Embodiments of the present disclosure provide a method and a device for Video See-Through (VST), a storage medium, and a program product. The method comprises: obtaining a first posture in a first Degree of Freedom mode; determining first depth information of a preset stereoscopic shape; and determining a first VST result according to the first depth information and the first posture.