Volumetric XR Session Grouping for Lower-Latency RGB-D Streaming
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing XR volumetric conversation systems waste bandwidth and computational resources due to the lack of consideration for the viewer's position, leading to unnecessary data transmission and increased latency, especially in computationally constrained devices.
Innovation Solution
Directly send compressed RGB-D data from capture devices to a cloud-based media service, where location-based synchronization and spatial grouping are performed to optimize data handling, reducing latency and computational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If all capture devices send volumetric video data without considering viewer position, then complete 3D scene coverage is achieved, but bandwidth consumption and computational resources are wasted
Solution Approach 1:
The system transmits different quantities of capture device data to different receivers based on their specific viewing positions and requirements. Each receiver receives optimized data from capture devices that are most relevant to its viewing angle, rather than all receivers getting identical complete scene data. This localizes the data transmission quality to match local viewing needs, reducing overall bandwidth consumption while maintaining scene coverage completeness.
2Reliability
If all capture devices send volumetric video data without considering viewer position, then complete 3D scene coverage is achieved, but processing time and latency increase
Solution Approach 1:
The system optimizes processing latency by having each receiver process only the subset of capture device data that is most relevant to its viewing position. This selective processing based on local viewing requirements reduces the computational burden and processing time compared to processing all capture device data, while still achieving complete scene coverage from the appropriate perspectives.
3Reliability
If excessive cameras are used to capture volumetric data, then comprehensive scene capture is achieved, but system complexity and noise increase
Solution Approach 1:
The system dynamically determines which capture devices to use based on the receiver's viewing position and requirements. Rather than statically configuring all cameras to always transmit data, the system adaptively selects and activates only the necessary capture devices for each viewing scenario. This dynamic approach maintains comprehensive scene capture capability while reducing system complexity by avoiding the permanent activation and configuration of excessive cameras.
4Measurement precision
If high resolution 3D objects are transmitted, then visual quality is improved, but computational complexity and latency at receiver end increase
Solution Approach 1:
The system transmits capture device data at appropriate resolution and quality levels tailored to each receiver's specific viewing position, device capabilities, and requirements. Rather than uniformly transmitting high resolution data to all receivers, the system locally optimizes the quality and resolution of transmitted data to match each receiver's needs, thereby maintaining visual quality where necessary while reducing computational complexity at receiver ends that have more constrained requirements.
Data Source
AI summary
There is disclosed an apparatus and a method for spatial computing service session description for volumetric XR conversation. In accordance an embodiment the method comprises receiving volumetric video from a plurality of sources, the volumetric video data comprising color images and depth images representing at least a part of a scene; determining which of the plurality of sources belong to a same group; selecting a group identifier for the group; associating the volumetric video with the group identifier indicative of from which group the volumetric video data is received; associating the volumetric video with a source identifier indicative of from which source of the group the volumetric video data is received; associating the color image of the volumetric video with a color data identifier; and associating the depth image of the volumetric video with a depth data identifier.


