Spatial Object Encoding for Adaptive Video Delivery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding technologies face challenges in delivering high-quality video content to devices with heterogeneous display and computational capabilities, particularly in scalable video and 360-degree virtual reality applications, where efficient bit-stream scalability and adaptive quality delivery are needed.
Innovation Solution
The proposed solution involves encoding media content into multiple spatial objects with independent encoding and decoding parameters, along with metadata that characterizes their relationships, allowing for incremental quality delivery and composition. This includes encoding a base quality layer and an incremental quality layer, where the incremental layer is derived by reconstructing and up-converting the base layer, and using different video coding standards for each object, enabling flexible codec selection and adaptive rendering.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If video content is encoded into multiple spatial objects with independent encoding parameters, then adaptability to different device capabilities is improved, but device complexity increases
Solution Approach 1:
The video content is segmented into multiple spatial objects (base quality layer and incremental quality layer) that can be independently encoded and decoded. Each spatial object is encoded separately with its own parameters, allowing selective rendering based on device capabilities while maintaining manageable complexity through modular processing
Solution Approach 2:
The system dynamically adapts the rendering process by selectively combining spatial objects based on device capabilities. The compositing operation is dynamically adjusted - devices can render only the base layer or combine base layer with incremental layer, allowing the system to optimize performance adaptively without requiring all devices to handle all layers
2Manufacturing precision
If multiple quality layers are transmitted and stored, then video quality is improved, but loss of time in processing and delivery increases
Solution Approach 1:
The video content is pre-encoded into multiple quality layers (base and incremental) during the encoding phase. This preliminary action allows the encoded layers to be stored and transmitted separately, so that during playback devices can quickly select and combine only the necessary layers without performing complex real-time encoding operations
Solution Approach 2:
The incremental quality layer is extracted as a separate component from the base quality layer. This extraction allows the incremental layer to be transmitted and processed independently, reducing the processing burden on devices that only need base quality while enabling high-quality reconstruction when both layers are available
3Adaptability or versatility
If different video coding standards are used for different spatial objects, then adaptability to heterogeneous devices is improved, but device complexity increases
Solution Approach 1:
The system uses a universal compositing framework that can handle multiple coding standards (H.264/AVC, H.265/HEVC, JPEG) within a single architecture. The base quality layer and incremental quality layer can be encoded with different standards, and the compositing operation universally processes them together, allowing devices to handle heterogeneous codecs without requiring separate processing paths for each standard
Data Source
AI summary
A media content delivery apparatus that encodes media content as multiple spatial objects is provided. The media content delivery apparatus encodes a first spatial object according to a first set of parameters. The media content delivery apparatus also encodes a second spatial object according to a second set of parameters. The first and second spatial objects are encoded independently. The media content delivery apparatus also generates a metadata based on the first set of parameters, the second set of parameters, and a relationship between the first and second spatial objects. The media content delivery apparatus then transmits or stores the encoded first spatial object, the encoded second spatial object, and the generated metadata.


