Split Transport Warping for Low-Latency AR Pose Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing warping techniques in cloud-based rendering for lightweight devices suffer from significant delay due to network jitter and latency, leading to outdated motion vectors and reduced accuracy in predicting virtual object positions and orientations, especially over wireless networks like 5G.
Innovation Solution
Utilize different transport channels with varying Quality of Service (QoS) characteristics, employing Ultra-Reliable Low Latency Communications (URLLC) for pose information and a lower latency high bandwidth channel for video streams to ensure up-to-date pose data at the client, reducing latency by 30-60 ms and improving pose prediction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If cloud-based rendering with single transport channel is used, then video transmission bandwidth is sufficient, but latency increases by 30-60 ms causing outdated motion vectors
Solution Approach 1:
The patent segments the transport channel into two separate channels: a first transport channel for pose information and a second transport channel for video streams. This segmentation allows each channel to be optimized for its specific data type, with the first channel prioritizing low latency for pose updates and the second channel providing sufficient bandwidth for video transmission, thereby resolving the contradiction between latency and reliability
Solution Approach 2:
The patent applies local quality by assigning different QoS characteristics to different data types. Pose information transmitted over the first channel receives ultra-low latency treatment (30-60 ms reduction) while video streams over the second channel receive standard bandwidth optimization. This localized quality differentiation ensures motion vectors remain accurate without compromising overall system performance
2Measurement precision
If ultra-low latency channel is used for pose information, then pose prediction accuracy improves, but network resource consumption increases
Solution Approach 1:
The patent segments network resource allocation by creating dedicated transport channels for different data types. The first channel is optimized for pose information with ultra-low latency characteristics, while the second channel handles video streams with standard bandwidth allocation. This segmentation prevents overprovisioning of radio resources while ensuring pose prediction accuracy improves by 30-60 ms
Solution Approach 2:
The patent changes the latency parameter specifically for pose information transmission by routing it through the first ultra-low latency channel. This parameter change improves pose prediction accuracy without requiring all network traffic to consume elevated resources, thereby avoiding excessive radio resource consumption while maintaining improved accuracy
3Measurement precision
If motion vectors are derived from rendered view frames, then accuracy is sufficient, but processing time increases causing delay
Solution Approach 1:
The patent applies preliminary action by transmitting pose information (including motion vectors and object state data) ahead of the actual video frames through the first ultra-low latency channel. This allows the client device to pre-compute motion vectors and perform warping operations before the corresponding video frames arrive, eliminating processing delays while maintaining accuracy
Solution Approach 2:
The patent introduces pose information as an intermediary that carries essential motion and object state data between the server and client. This intermediary contains pre-extracted motion vectors and object parameters that eliminate the need for time-consuming real-time extraction from video frames, thereby reducing processing time while preserving accuracy
Data Source
AI summary
A network device (110) generates (810) video data representing a viewing frustrum of a three-dimensional scene. A plurality of virtual objects are within the viewing frustrum. The network device (110) transmits (820) pose information and the video data to a computing device (120) over a first transport channel (130a) and a second transport channel (130b), respectively. The first transport channel (130a) has lower latency characteristics than the second transport channel (130b) and the pose information comprises a pose of a virtual object within the viewing frustrum. The computing device (120) receives (860), from the network device (110), the pose information and video data over the first transport channel (130a) and the second transport channel (130b), respectively. The computing device (120) predicts (870) a newer pose of the virtual object from the pose information and generates (880) a two-dimensional image using the predicted pose and the video data as inputs to a warping function.


