Selective Learning-Based Model Application in Video Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video encoding and decoding technologies using neural networks face challenges in optimizing the operation of neural network filters, particularly in deciding when to apply learning-based models for encoding and decoding video frames, which affects the quality and efficiency of the compression and reconstruction processes.

Innovation Solution

The proposed solution involves a method and apparatus that selectively apply a learning-based model for video frame encoding and decoding, where the model is fine-tuned based on the activation status of previously-decoded blocks, and the decision to use the model is encoded into a bitstream, allowing for adaptive processing and improved visual quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If a learning-based model is applied for encoding and decoding video blocks, then the visual quality of reconstructed content is improved, but the computational overhead and processing time increase

Engineering Contradiction:
Improvevisual quality of reconstructed contentVSAvoidcomputational overhead
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies the learning-based model selectively to only certain video blocks that benefit most from it, rather than applying it to all blocks. The system determines which blocks require the learning-based model based on specific criteria, applying the model partially to achieve quality improvement while avoiding unnecessary computational overhead on blocks where the model would not provide significant benefit.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent implements different processing qualities for different regions of the video content. By analyzing local block characteristics and applying the learning-based model only where needed, the system creates variable quality processing - high quality where the model is applied and standard quality where it is not - thereby optimizing the overall quality-computation trade-off.

Inventive Principle:
Principle #3Local quality

2Productivity

If the learning-based model is selectively fine-tuned based on previously-decoded blocks, then the encoding efficiency is improved, but the decision-making complexity increases

Engineering Contradiction:
Improveencoding efficiencyVSAvoiddecision-making complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements a feedback mechanism where the system monitors the performance and activation status of the learning-based model on previously decoded blocks. This feedback information is used to dynamically adjust and fine-tune the model's application strategy, allowing the system to learn from past decisions and improve encoding efficiency over time while adapting to different content characteristics.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs preliminary analysis of video blocks to determine in advance whether the learning-based model should be applied. By pre-assessing block characteristics and predicting which blocks would benefit from the model before actual encoding occurs, the system streamlines the decision-making process and improves encoding efficiency by avoiding unnecessary model applications.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If the decision on usage of the learning-based model is encoded into a bitstream, then the adaptability of the decoding process is improved, but the bitstream size increases

Engineering Contradiction:
Improveadaptability of the decoding processVSAvoidbitstream size
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential decision information needed for the learning-based model usage and encodes only this critical data into the bitstream, rather than encoding complete model parameters or extensive metadata. This selective extraction approach maintains the adaptability needed for proper decoding while minimizing the additional bitstream overhead.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial encoding of the model usage decisions, encoding information only for blocks where the learning-based model is applied or where such information is necessary for reconstruction. This partial approach provides sufficient adaptability for the decoding process while avoiding the overhead of encoding complete information for all blocks.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12108050B2Method, an apparatus and a computer program product for video encoding and video decoding
Publication Date: 2024.10.01 NOKIA TECHNOLOGIES OY
  • US12108050B2 patent drawing
  • US12108050B2 patent drawing
  • US12108050B2 patent drawing

AI summary

The embodiments relate to a method for encoding and a decoding, and apparatuses for the same. The method for encoding comprises receiving a block of a video frame for encoding (1510); making a decision on whether or not a learning-based model is to be applied as a processing step for encoding the block (1520); applying the learning-based model for said input block according to the decision, where the learning-based model has been selectively fine-tuned according to information relating to activation of the learning-based model of previously-decoded blocks (1530); encoding a signal corresponding to the decision on usage of the learning-based model into a bitstream (1540); and encoding the block into a bitstream with an information whether the block is to be used for finetuning (1550).