3D Reference Motion Vector for Panoramic Video Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video coding methods for panoramic 360° videos are inefficient due to the limitations of traditional motion compensation models, which are not suitable for spherical video projections, leading to increased signaling overhead and reduced motion compensation effectiveness.

Innovation Solution

The use of 3D reference motion vectors is introduced to predict pixel positions in panoramic video encoding and decoding, allowing for more efficient motion compensation by projecting pixel positions onto a viewing sphere and de-projecting them back onto a 2D plane, enabling the use of a single motion vector for multiple pixels and accounting for geometrical distortions caused by sphere-to-plane projection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If traditional 2D motion compensation models are used for panoramic video, then the encoding process is simple, but motion compensation effectiveness is reduced due to geometrical distortions from sphere-to-plane projection

Engineering Contradiction:
Improveencoding simplicityVSAvoidmotion compensation effectiveness
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent transitions from 2D motion compensation to 3D motion compensation by mapping pixel positions onto a virtual sphere in three-dimensional space. The 3D reference motion vector operates in spherical coordinates (azimuth and elevation angles), allowing the system to account for the curved geometry of panoramic video while maintaining computational efficiency through dimensionality change rather than complexity increase.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If higher order compensation models with additional motion vectors are used, then complex transformations are covered, but signaling overhead increases reducing overall effectiveness

Engineering Contradiction:
Improvemotion model capabilityVSAvoidsignaling overhead
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent changes the parameter space from 2D pixel coordinates to 3D spherical coordinates (azimuth angle φ and elevation angle θ). This parameter transformation allows a single 3D reference motion vector to capture complex transformations including rotation, zoom, and perspective effects that would otherwise require multiple 2D motion vectors, thereby reducing signaling overhead while maintaining adaptability.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If a single motion vector is used for all pixels in a block, then signaling is efficient, but it cannot account for geometrical distortions in panoramic projection

Engineering Contradiction:
Improveencoding efficiencyVSAvoidmotion compensation precision
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

By introducing the third dimension (spherical space) for motion vector representation, the patent enables a single 3D reference motion vector to inherently account for geometrical distortions. The azimuth and elevation angle changes naturally represent the curved projection geometry, allowing one vector to achieve the precision that would otherwise require multiple vectors in 2D space.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentEP3610647B1Apparatuses and methods for encoding and decoding a panoramic video signal
Publication Date: 2021.12.08 HUAWEI TECH CO LTD
  • EP3610647B1 patent drawingFigure 1
  • EP3610647B1 patent drawingFigure 2
  • EP3610647B1 patent drawingFigure 3

AI summary

The invention relates to an apparatus (101) for encoding a video signal, wherein the video signal is a two-dimensional projection of a panoramic video signal and comprises a plurality of successive frames, including a reference frame and a current frame, wherein each frame of the plurality of successive frames comprises a plurality of video coding blocks and wherein each video coding block comprises a plurality of pixels. The encoding apparatus (101) comprises: an inter prediction unit (105) to generate a predicted video coding block for the current video coding block of the current frame on the basis of the corresponding video coding block of the reference frame, wherein the inter prediction unit (105) is configured to determine a 3D reference motion vector on the basis of a first 3D position defined by a projection of a first pixel of a current video coding block of the current frame onto a viewing sphere and on the basis of a second 3D position defined by a projection of a second pixel of the corresponding video coding block of the reference frame onto the viewing sphere and to predict at least one pixel of the current video coding block on the basis of the 3D reference motion vector; and an encoding unit (103) configured to encode the current video coding block of the current frame on the basis of the predicted video coding block. Moreover, the invention relates to a corresponding decoding apparatus (151), encoding method and decoding method.