Temporal Scalability in Video Encoding with Scene Change Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video encoding systems that are not scalable from a temporal point of view fail to maintain scaling functionalities and compression/quality performance when scene changes occur, as they do not adapt to the hierarchical relation between pictures in a video sequence, leading to inefficient coding and potential loss of scalability.
Innovation Solution
A method and system for encoding time-scalable videos that dynamically adapt the temporal prediction structure to react to scene changes by converting successive Key Pictures to Intra-pictures and adjusting the Intra Period, maintaining uniform scalability and improving coding efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If motion-compensated prediction is used for encoding video pictures, then coding efficiency is improved, but temporal scalability is lost when scene changes occur
Solution Approach 1:
The patent applies dynamics by making the prediction structure adaptive rather than fixed. The encoder dynamically switches between hierarchical temporal prediction (for normal scenes) and non-hierarchical prediction (for scene changes), allowing the system to optimize coding efficiency while maintaining temporal scalability when needed. This is achieved through scene change detection that triggers structural modifications to the prediction hierarchy.
Solution Approach 2:
The patent changes the structural parameters of the prediction system based on scene content. When scene changes are detected, the system modifies the prediction structure by disabling hierarchical constraints and adjusting reference picture selection, thereby adapting the coding approach to maintain both efficiency and scalability under varying conditions.
2Manufacturing precision
If Intra-pictures are used frequently to maintain quality after scene changes, then quality is improved, but computational complexity increases
Solution Approach 1:
The patent applies local quality by applying different encoding strategies to different temporal regions. Instead of uniformly using Intra-pictures throughout, the system uses hierarchical prediction for stable scenes and switches to non-hierarchical prediction only when scene changes are detected, thereby maintaining quality where needed while reducing unnecessary computational complexity in stable regions.
Solution Approach 2:
The system dynamically adjusts the Intra-picture frequency based on scene change detection. Rather than using Intra-pictures at fixed intervals or always, the encoder activates non-hierarchical prediction selectively during scene transitions, optimizing the balance between quality maintenance and computational load.
3Adaptability or versatility
If hierarchical temporal prediction structure is maintained, then temporal scalability is preserved, but coding efficiency decreases during scene changes
Solution Approach 1:
The patent makes the prediction structure dynamic by allowing it to switch between hierarchical and non-hierarchical modes. During scene changes, the system temporarily adopts non-hierarchical prediction to improve coding efficiency, then returns to hierarchical structure when stability is restored, thus balancing scalability and efficiency based on real-time scene conditions.
Solution Approach 2:
The system changes the structural parameters of the prediction hierarchy in response to scene changes. When scene changes are detected, the encoder modifies the prediction structure by relaxing hierarchical constraints and adjusting reference picture relationships, thereby improving coding efficiency during transitions while preserving temporal scalability through selective application of the changes.
4Productivity
If scene change detection and adaptive coding mode variation are implemented, then coding efficiency is improved, but device complexity increases
Solution Approach 1:
The patent applies preliminary action by detecting scene changes in advance and proactively switching the prediction structure before encoding subsequent pictures. This allows the system to prepare the appropriate coding mode ahead of time, improving efficiency by avoiding suboptimal encoding of scene transition frames while keeping the added complexity manageable through pre-computed detection mechanisms.
Solution Approach 2:
The system uses feedback from scene change detection to dynamically adjust the prediction structure. The detector monitors scene stability and provides feedback that triggers structural changes in the prediction system, creating a closed-loop control mechanism that optimizes coding efficiency while managing complexity through intelligent, condition-based decision-making.
Data Source
Figure 1~2b
Figure 3~4
Figure 5~6b
AI summary
An encoder device (SE) allows generating, starting from a sequence (IS) of digital video pictures, a time-scalable encoded bitstream (SB) obtained by applying to the pictures a hierarchical prediction (P) wherein the pictures are organised in Groups Of Picture or GOPs including: - base time layer (L0) pictures, designated Key Picture and suitable for encoding a Inter or Intra, with and without motion-compensated prediction respectively, - higher time layer (L1, L2, L3) pictures, adapted to be selectively eliminated to effect time scalability of the encoded scalable bitstream (SB). Carried out for such purpose is the detection (IB) of the possible presence of scene changes (SC) in the sequence (IS) of digital video pictures, and, in the presence of a scene change (SC), the first Key Picture after the scene change (SC) is in any case encoded (120) as Intra (I0).