Recursive Scene Segmentation for HDR Video Metadata Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In cloud-based video coding architectures, existing techniques for processing HDR video often require frame-by-frame updates of reshaping metadata, leading to significant overhead, especially at low bit rates, due to the need for scene-based processing across multiple computing nodes.

Innovation Solution

The proposed solution involves segmenting video frames into scenes and sub-scenes, using a recursive scene splitting algorithm to minimize the need for reshaping function updates, thereby reducing metadata overhead by applying scene-based forward and backward reshaping functions to maintain temporal consistency across nodes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If frame-by-frame updates of reshaping metadata are used, then video quality is maintained, but metadata overhead increases significantly

Engineering Contradiction:
Improvevideo qualityVSAvoidmetadata overhead
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The video sequence is divided into multiple scenes, and within each scene, frames are grouped into segments that share common reshaping characteristics. This segmentation allows reshaping metadata to be updated at the segment level rather than frame-by-frame, reducing metadata overhead while maintaining video quality through localized updates only when scene characteristics change.

Inventive Principle:
Principle #1Segmentation

2Productivity

If scene-based processing is applied across multiple computing nodes, then processing efficiency is improved, but metadata overhead increases

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidmetadata overhead
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent segments video sequences into scenes that can be independently processed by different computing nodes. Each node processes a specific scene and generates reshaping metadata only for its assigned segment, eliminating the need for global frame-by-frame metadata updates across all nodes and reducing overall metadata overhead while maintaining parallel processing efficiency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by generating reshaping metadata with local validity scope limited to specific scenes or segments rather than globally across the entire video sequence. Each computing node generates metadata locally for its assigned scene, reducing redundant metadata transmission and processing across the distributed system while maintaining processing efficiency.

Inventive Principle:
Principle #3Local quality

3Quantity of substance

If recursive scene splitting is implemented, then metadata overhead is reduced, but computational complexity increases

Engineering Contradiction:
Improvemetadata overheadVSAvoidcomputational complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by performing recursive scene splitting and segment identification during the initial video analysis phase. The video sequence is pre-segmented into scenes and segments with assigned validity ranges for reshaping metadata before actual encoding begins. This preliminary segmentation reduces the computational burden during the main encoding process and minimizes metadata overhead by establishing segment boundaries in advance.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12190587B2Recursive segment to scene segmentation for cloud-based coding of HDR video
Publication Date: 2025.01.07 DOLBY LABORATORIES LICENSING CORP
  • US12190587B2 patent drawing
  • US12190587B2 patent drawing
  • US12190587B2 patent drawing

AI summary

In a cloud-based system for encoding high dynamic range (HDR) video, each node receives a video segment and bumper frames. Each segment is subdivided into primary scenes and secondary scenes to derive scene-based forward reshaping functions that minimize the amount of reshaping-related metadata when coding the video segment, while maintaining temporal continuity among scenes processed by multiple nodes. Methods to generate scene-based forward and backward reshaping functions to optimize video coding and improve the coding efficiency of reshaping-related metadata are also examined.