Scene Change Detection Using Sum of Variance and Encoding Cost

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video encoding and transcoding techniques are inefficient in accurately detecting scene features such as scene changes, fade-ins, and fade-outs, leading to suboptimal encoding and transcoding processes.

Innovation Solution

The use of sum of variance (SVAR) and estimated picture encoding cost (PCOST) metrics to dynamically adjust quantization parameters and identify scene features, enabling more efficient encoding and transcoding by tailoring quantization settings to frame complexities and detecting scene changes, fade-ins, and fade-outs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional sound level change detection is used to identify scene changes, then the detection process is simple, but the detection accuracy is poor and scene features are not accurately identified

Engineering Contradiction:
Improvescene change detection accuracyVSAvoiddetection process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent changes the detection parameter from sound level to visual parameters including sum of variance (SVAR) and estimated picture encoding cost (PCOST). By using multiple parameters (SVAR for spatial complexity, PCOST for encoding complexity) instead of a single sound parameter, the system achieves more accurate scene change detection while maintaining computational efficiency through parameter transformation of existing encoding data.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If accurate scene feature detection is implemented using multiple metrics, then encoding efficiency is improved, but computational complexity increases

Engineering Contradiction:
Improvevideo encoding efficiencyVSAvoidcomputational complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system uses self-service by leveraging data already generated during the normal video encoding process. The SVAR and PCOST metrics are derived from intermediate encoding data without requiring separate detection processes. This allows accurate scene feature detection to serve the encoding process itself, improving efficiency while avoiding additional computational overhead through resource reuse.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent makes the encoding metrics serve multiple functions: they are used both for the actual video encoding and simultaneously for scene change detection and quantization parameter adjustment. This multi-functionality allows the system to achieve accurate scene feature detection and optimized encoding efficiency using the same computational resources, eliminating the need for separate detection systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If quantization parameters are dynamically adjusted based on frame complexity, then coding efficiency is improved, but processing time increases

Engineering Contradiction:
Improvecoding efficiencyVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary action by calculating SVAR and PCOST metrics during the early stages of encoding before final quantization parameter determination. By pre-computing these complexity metrics from intermediate encoding data, the system prepares the necessary information in advance, allowing rapid QP adjustment decisions without adding significant processing time to the overall encoding pipeline.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9426475B2Scene change detection using sum of variance and estimated picture encoding cost
Publication Date: 2016.08.23 VIXS SYSTEMS INC
  • US9426475B2 patent drawing
  • US9426475B2 patent drawing
  • US9426475B2 patent drawing

AI summary

A video processing device includes a complexity estimation module to determine a first sum of variances metric and a first estimated picture encoding cost metric for a first picture of a video stream. The video processing device further includes a scene analysis module to determine a first threshold based on a first statistical feature for sum of variance metrics of a set of one or more pictures preceding the first picture in the video stream and a second threshold based on a second statistical feature for estimated picture encoding cost metrics of the set of one or more pictures. The scene analysis module further is to identify a scene change as occurring at the first picture based on the first sum of variances metric, the first estimated picture encoding cost metric, the first threshold, and the second threshold.