Perceptual Representations for Motion Vector Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing motion processing methods for video compression face challenges in accurately determining motion vectors, especially in low-texture regions and during changes in contrast, brightness, blur, noise, and transitions, leading to inefficiencies and increased processing costs.

Innovation Solution

The use of perceptual representations to generate motion vectors, which are based on models of human perception, improves accuracy and precision by enhancing imaging in low-contrast areas and maintaining resilience to changes, thereby reducing processing requirements and costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If direct methods (pyramidal and block-based searches) are used to determine motion vectors, then motion vector accuracy and precision can be increased, but processing power and processing costs increase significantly

Engineering Contradiction:
Improvemotion vector accuracyVSAvoidprocessing power
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent introduces perceptual quality metrics as an intermediary evaluation mechanism between motion estimation and motion compensation. Instead of directly comparing pixel values, the system uses perceptual representations to assess motion vector accuracy, reducing the computational burden of direct methods while maintaining or improving precision through human vision-based evaluation criteria

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If fidelity metrics are used to identify matches between estimated movement and generated motion vectors, then false matches can be removed, but opportunistic best matches and motion vector outliers occur which reduce compression quality and efficiency

Engineering Contradiction:
Improvefalse match removalVSAvoidmotion vector accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent changes the evaluation parameter from traditional fidelity metrics (which compare raw pixel differences) to perceptual quality metrics (which evaluate based on human visual perception characteristics). This parameter transformation allows the system to correctly identify and remove false matches while avoiding opportunistic best matches, as perceptual metrics better reflect actual visual quality implications of motion vector errors

Inventive Principle:
Principle #35Parameter changes

3Power

If existing evaluation methods using fidelity metrics are used, then processing can be kept at lower costs, but motion vector accuracy deteriorates in low-texture regions and during transitions

Engineering Contradiction:
Improveprocessing costVSAvoidmotion vector accuracy
Core Design Contradiction:
PowerVSMeasurement precision

Solution Approach 1:

The patent applies local quality assessment by using perceptual representations that are specifically designed to handle different image regions differently. The perceptual metrics account for local characteristics such as texture content, edge presence, and transition types, allowing accurate motion vector evaluation in challenging regions like low-texture areas and transitions without requiring uniformly high processing power across the entire image

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9503756B2Encoding and decoding using perceptual representations
Publication Date: 2016.11.22 ARRIS ENTERPRISES INC
  • US9503756B2 patent drawing
  • US9503756B2 patent drawing
  • US9503756B2 patent drawing

AI summary

Encoding a video signal including pictures includes generating perceptual representations based on the pictures. Reference pictures are selected and motion vectors are generated based on the perceptual representations and the reference pictures. The motion vectors and pointers for the reference pictures are provided in an encoded video signal. Decoding may include receiving pointers for reference pictures and motion vectors based on perceptual representations of the reference pictures. The decoding of the pictures in the encoded video signal may include selecting reference pictures using the pointers and determining predicted pictures, based on the motion vectors and the selected reference pictures. The decoding may include generating reconstructed pictures from the predicted pictures and the residual pictures.