Contextual Video Compression Using Reference-Image Feature Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional residual-based video coding methods fail to optimally utilize the redundancy between frames, resulting in suboptimal reconstruction quality and compression efficiency due to limited characterization of contextual information.

Innovation Solution

A context-based coding solution that extracts higher-dimensional contextual feature representations from reference images to guide adaptive encoding and decoding, utilizing machine learning models to enhance compression efficiency and reconstruction quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional residual-based video coding methods are used, then the coding process is simple, but the reconstruction quality and compression efficiency are suboptimal due to limited characterization of contextual information

Engineering Contradiction:
Improvereconstruction qualityVSAvoidcoding process complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent transitions from conventional residual-based coding to deep contextual video coding by extracting contextual feature representations from reference images. This adds a new dimension of contextual information characterization beyond simple pixel residuals, enabling better reconstruction quality through enhanced feature extraction and conditional decoding mechanisms

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent changes the coding parameters by using deep learning models to extract contextual features from reference images. Instead of traditional residual calculations, the system employs conditional encoding and decoding based on extracted contextual representations, fundamentally altering the coding approach to achieve superior reconstruction quality

Inventive Principle:
Principle #35Parameter changes

2Productivity

If conventional residual-based video coding methods are used, then the implementation is straightforward, but the compression efficiency is limited due to insufficient utilization of frame redundancy

Engineering Contradiction:
Improvecompression efficiencyVSAvoidcontextual information characterization
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent introduces contextual feature representations as an intermediary between reference images and the coding process. These extracted features serve as mediators that capture essential contextual information, enabling more efficient compression by better representing the relationship between frames while reducing information loss

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces conventional mechanical residual-based coding mechanisms with deep learning-based contextual feature extraction. This substitution enables the system to automatically learn and utilize frame redundancy patterns, significantly improving compression efficiency through intelligent feature representation rather than fixed algorithmic approaches

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12593044B2Deep contextual video image compression
Publication Date: 2026.03.31 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12593044B2 patent drawing
  • US12593044B2 patent drawing
  • US12593044B2 patent drawing

AI summary

According to implementations of the present disclosure, there is provided a context-based image coding solution. According to the solution, a reference image of a target image is obtained. A contextual feature representation is extracted from the reference image, the contextual feature representation characterizing contextual information associated with the target image. Conditional encoding or conditional decoding is performed on the target image based on the contextual feature representation. In this way, the enhancement of the performance is achieved in terms of the reconstruction quality and the compression efficiency.