Scalable Codec Processing with AI-Generated Enhanced Layers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing scalable codecs, such as SVC and SHVC, face challenges in efficiently processing and transmitting scalable video streams due to complex decoding requirements and memory constraints, particularly when generating and using base and enhanced layers with different view ports and scalabilities.
Innovation Solution
An electronic device generates a base layer image based on image characteristics like resolution, color gamut, frame rate, bit depth, or bit rate, and uses artificial intelligence networks to compress and transmit this image along with configuration information, allowing receivers to generate enhanced layers efficiently using AI networks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional scalable codecs (SVC, SHVC) are used to process scalable video streams with base and enhanced layers, then video compression and scalability are achieved, but decoding complexity and memory requirements increase significantly
Solution Approach 1:
The patent extracts the enhanced layer processing from traditional scalable codec frameworks and implements it through separate AI neural networks. The base layer is decoded traditionally, then AI networks are applied to generate enhanced layers, separating complex processing into modular components that reduce overall decoding complexity while maintaining scalability.
Solution Approach 2:
The patent replaces traditional mechanical codec processing mechanisms with AI-based neural networks for enhanced layer generation. Instead of using complex inter-layer prediction and refinement mechanisms in SVC/SHVC, the system uses AI models to transform base layer output into enhanced layers, simplifying the processing architecture.
2Manufacturing precision
If traditional scalable codecs transmit both base layer and enhanced layer data, then complete video quality is maintained, but data transmission size increases
Solution Approach 1:
The patent extracts only the necessary base layer information for transmission and uses AI neural networks at the receiving end to generate enhanced layers. This eliminates the need to transmit redundant enhanced layer data while maintaining video quality through AI-based synthesis from the compressed base layer.
Solution Approach 2:
The patent creates virtual copies of enhanced layer data through AI neural network processing rather than transmitting actual enhanced layer bitstreams. The AI models generate synthetic enhanced layer representations from base layer input, reducing transmission quantity while preserving quality through intelligent reconstruction.
3Adaptability or versatility
If receivers decode both base layer and enhanced layer in traditional scalable codecs, then full scalability is achieved, but processing time and computational resources increase
Solution Approach 1:
The patent implements partial processing where receivers can selectively process only the base layer for basic viewing requirements, or apply AI neural networks to generate enhanced layers only when needed. This partial action approach allows receivers to adjust processing level based on capabilities and requirements, reducing processing time while maintaining scalability options.
Solution Approach 2:
The patent performs preliminary base layer decoding and AI model preparation before actual video processing. The AI neural networks are pre-configured and the base layer is decoded in advance, enabling faster enhanced layer generation when required and reducing overall processing time through preparatory actions.
Data Source
AI summary
An electronic device is provided. The electronic device includes a communication circuit, memory, comprising one or more storage media, storing instructions, and one or more processors communicatively coupled to the communication circuit and the memory, wherein the instructions, when executed by the one or more processors individually or collectively, cause the electronic device to determine information on a scalability applicable to an original image, generate a base layer (BL) image for the original image on the basis of at least one of the resolution, color gamut, gamma, bit depth, or bitrate of the original image, generate at least one artificial intelligence (AI) network on the basis of the original information, the information on the scalability about the original image, and the BL image, compress the BL image and control the communication circuit to transmit a bitstream including the compressed BL image and configuration information on the at least one AI network to an external electronic device.


