Hybrid Multi-View Sensor Coding with Virtual Pose Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current multi-view immersive video formats assume a depth map is associated with a physical or virtual camera, which limits the flexibility and efficiency of encoding and decoding when using hybrid sensor configurations, particularly in close-range or indoor setups.
Innovation Solution
Incorporate extrinsic parameters, such as sensor poses, into metadata to enable the generation of additional depth components by warping existing depth components, reducing the need to transmit multiple depth maps and allowing decoding to generate depth components at the client side.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple depth maps are transmitted for different sensor positions, then rendering accuracy is improved, but data transmission volume increases
Solution Approach 1:
The patent extracts only the essential depth map and extrinsic parameters from multiple sensor positions, transmitting only the minimum necessary data while enabling reconstruction of other views through warping operations at the decoding end
Solution Approach 2:
Instead of transmitting multiple depth maps, the patent transmits one depth map and uses extrinsic parameters to generate synthetic copies of depth information for other sensor positions through virtual camera warping, achieving multiple views from a single source
2Stability of the object's composition
If depth maps are associated with specific camera positions, then rendering consistency is improved, but system flexibility deteriorates
Solution Approach 1:
The patent makes the system dynamic by allowing depth maps to be associated with virtual camera positions rather than fixed physical cameras. The extrinsic parameters enable flexible repositioning and warping of depth maps to match different sensor configurations, combining consistency with adaptability
Solution Approach 2:
The patent changes the association parameters from fixed camera positions to virtual camera positions with extrinsic parameters. This parameter transformation enables the same depth map to be consistently rendered for different sensor configurations by adjusting the virtual camera parameters
Data Source
AI summary
A method for transmitting multi-view image frame data. The method comprises obtaining multi-view components representative of a scene generated from a plurality of sensors, wherein each multi-view component corresponds to a sensor and wherein at least one of the multi-view components includes a depth component and at least one of the multi-view components does not include a depth component. A virtual sensor pose is obtained for each sensor in a virtual scene, wherein the virtual scene is a virtual representation of the scene and wherein the virtual sensor pose is a virtual representation of the pose of the sensor in the scene when generating the corresponding multi-view component. Sensor parameter metadata is generated for the multi-view components, wherein the sensor parameter metadata contains extrinsic parameters for the multi-view components and the extrinsic parameters contain at least the virtual sensor pose of a sensor for each of the corresponding multi-view components. The extrinsic parameters enable the generation of additional depth components by warping the depth components based on their corresponding virtual sensor pose and a target position in the virtual scene. The multi-view components and the sensor parameter metadata is thus transmitted.


