VST XR Planar Frame Transformation for Low-Latency Head Pose Shifts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Optical see-through (OST) XR systems face challenges such as limited fields of view, indoor-only usage, and complex optical pipelines, while video see-through (VST) XR systems suffer from processing latencies due to high-resolution image frames, leading to incorrect or mistimed image frame displays.
Innovation Solution
Implementing dynamically-adaptive planar transformations in VST XR devices that project and transform image frames onto selected planes without requiring dense depth maps, using head motion prediction and adaptive plane selection to reduce latency and computational load.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If high-resolution image frames are processed in VST XR systems, then image quality is improved, but processing latency increases
Solution Approach 1:
The patent segments the image processing pipeline into multiple stages: capturing multiple lower-resolution image frames at different time points, predicting head pose changes, selectively rendering only certain frames based on predicted motion, and transforming these frames. This segmentation allows the system to maintain image quality while reducing overall processing latency by avoiding full processing of all high-resolution frames.
Solution Approach 2:
The patent applies partial action by selectively rendering only certain image frames based on predicted head pose changes. Instead of processing all captured frames at full resolution, the system identifies frames that need rendering based on motion prediction, thereby reducing processing latency while maintaining perceived image quality for the user.
2Measurement precision
If dense depth maps are used for image frame transformation, then transformation accuracy is improved, but computational load increases
Solution Approach 1:
The patent applies local quality by using plane-specific parameters rather than dense depth maps for all pixels. Each plane is characterized by local geometric parameters (normal vectors, distances) that are sufficient for transformation accuracy in that specific region. This approach maintains transformation accuracy where needed while significantly reducing computational load by avoiding full dense depth map generation and processing.
Solution Approach 2:
The patent changes the parameter representation from dense per-pixel depth values to sparse plane parameters (normal vectors, distances). This parameter transformation maintains the necessary geometric information for accurate image frame transformation while dramatically reducing the data volume and computational requirements for processing.
3Stability of the object's composition
If multiple image frames are captured and processed, then view continuity is improved, but processing time increases
Solution Approach 1:
The patent applies preliminary action by capturing multiple image frames in advance before they are needed for display. The system pre-captures frames at different time points, predicts which frames will be needed based on head pose prediction, and prepares them for rendering beforehand. This allows the system to maintain view continuity without processing all captured frames, thereby reducing processing time.
Solution Approach 2:
The patent applies dynamics by making the rendering process adaptive based on predicted head motion. The system dynamically determines which captured frames need to be rendered based on predicted head pose changes, rather than processing all frames statically. This dynamic selection maintains view continuity while significantly reducing the number of frames that require full processing.
Data Source
AI summary
A method includes obtaining multiple image frames captured using one or more imaging sensors of a video see-through (VST) extended reality (XR) device while a user's head is at a first head pose and depth data associated with the image frames. The method also includes predicting a second head pose of the user's head when rendered images will be displayed. The method further includes projecting at least one of the image frames onto one or more first planes to generate at least one projected image frame. The method also includes transforming the at least one projected image frame from the one or more first planes to one or more second planes corresponding to the second head pose to generate at least one transformed image frame. The method further includes rendering the at least one transformed image frame for presentation on one or more displays of the VST XR device.


