Video Coding Reference Picture Processing for Seamless Quality Switching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In scalable video coding, the use of open GOPs at segment boundaries in dynamic adaptive streaming over HTTP can result in decoding issues, such as the inability to decode random access skipped pictures, leading to picture rate glitches during quality and resolution switching.

Innovation Solution

A method and apparatus for encoding and decoding video representations that involve processing decoded pictures from a first coded video representation to create reference pictures for a second coded video representation, which differs in chroma format, sample bit depth, color gamut, or spatial resolution, through resampling and sample value scaling, allowing seamless switching between different video qualities and resolutions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If open GOPs are used at segment boundaries in DASH, then quality and resolution switching is enabled, but random access skipped pictures cannot be decoded, leading to picture rate glitches

Engineering Contradiction:
Improvequality switching capabilityVSAvoiddecoding reliability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent applies preliminary action by processing decoded pictures from the first representation in advance to create suitable reference pictures for the second representation before switching occurs. This pre-processing ensures that when quality switching happens, the decoder has ready-to-use reference pictures that are compatible with the new representation, preventing decoding failures of RASL pictures and eliminating picture rate glitches while maintaining reliable playback

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If processed decoded pictures from first representation are used as reference for second representation, then seamless quality switching is achieved, but additional processing steps are required

Engineering Contradiction:
Improverepresentation switching capabilityVSAvoiddecoding process complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent uses an intermediary approach by introducing processed decoded pictures as a intermediate element between the first and second representations. These processed pictures serve as a bridge that enables seamless transitions between different quality representations. The intermediary processed pictures contain the necessary reference information adapted from the first representation, allowing the second representation to be decoded correctly without requiring complex direct conversion processes

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies parameter changes by modifying the decoded pictures from the first representation through resampling and sample value scaling to create reference pictures suitable for the second representation. These parameter transformations (resolution, sampling rate, color space) enable the reference pictures to match the characteristics of the target representation, facilitating smooth quality switching while managing complexity through standardized processing operations

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240298017A1Apparatus, a method and a computer program for video coding and decoding
Publication Date: 2024.09.05 NOKIA TECHNOLOGIES OY
  • US20240298017A1 patent drawing
  • US20240298017A1 patent drawing
  • US20240298017A1 patent drawing

AI summary

There is provided methods, apparatuses and computer program products for video coding and decoding. A first part of a first coded video representation is decoded, and information on decoding a second coded video representation is received and parsed. The coded second representation differs from the first coded video representation in chroma format, sample bit depth, color gamut and/or spatial resolution, and the information indicates if the second coded video representation may be decoded using processed decoded pictures of the first coded video representation as reference pictures. If the information indicates that the second coded video representation may be decoded using processed decoded pictures of the first coded video representation as a prediction reference, decoded picture(s) of the first part is/are processed into processed decoded picture(s) by resampling and/or sample value scaling; and decoding a second part of a second video representation using said processed decoded picture(s) as reference pictures.