Spherical Motion Estimation for 360-Degree Video Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video coding systems for 360° video struggle to effectively detect redundancies in two-dimensional representations of three-dimensional image content, leading to inefficient bandwidth usage due to distortions caused by the conversion of three-dimensional space into two-dimensional data.

Innovation Solution

The implementation of a video coding system that uses spherical-domain projections to predictively code 360° video data by transforming input and reference pictures into spherical representations, allowing for the detection of redundancies and efficient coding through motion vector estimation and differential coding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If video coding systems use traditional two-dimensional representation for 360° video, then the coding process is simple, but bandwidth efficiency deteriorates due to undetected redundancies caused by distortions

Engineering Contradiction:
Improvebandwidth efficiencyVSAvoidcoding system complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The patent transforms the video coding approach from traditional two-dimensional representation to spherical-domain representation. By mapping 360° video content onto a spherical coordinate system, the coding system can properly account for the three-dimensional nature of the content, enabling accurate redundancy detection and significantly improving bandwidth efficiency while maintaining manageable complexity through systematic transformation processes

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If video coding systems transform to spherical-domain projections, then redundancy detection improves, but processing complexity increases

Engineering Contradiction:
Improveredundancy detection accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies spherical projection transformation as a preliminary step before motion estimation and redundancy detection. By pre-transforming the 360° video content into spherical coordinates, the system establishes a proper geometric framework that enables accurate redundancy detection without requiring complex adaptive processing during the main coding stages, thus managing overall processing complexity

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If traditional motion estimation is used on distorted two-dimensional data, then processing is fast, but coding accuracy deteriorates

Engineering Contradiction:
Improvemotion estimation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent changes the coordinate system parameters from Cartesian (x,y) to spherical coordinates (θ,φ) for motion estimation. This parameter transformation allows motion vectors to be calculated in a coordinate system that naturally represents the geometry of 360° video content, significantly improving motion estimation accuracy. The systematic nature of the spherical transformation maintains processing efficiency comparable to traditional methods

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11818394B2Sphere projected motion estimation/compensation and mode decision
Publication Date: 2023.11.14 APPLE INC
  • US11818394B2 patent drawing
  • US11818394B2 patent drawing
  • US11818394B2 patent drawing

AI summary

Techniques are disclosed for coding video data predictively based on predictions made from spherical-domain projections of input pictures to be coded and reference pictures that are prediction candidates. Spherical projection of an input picture and the candidate reference pictures may be generated. Thereafter, a search may be conducted for a match between the spherical-domain representation of a pixel block to be coded and a spherical-domain representation of the reference picture. On a match, an offset may be determined between the spherical-domain representation of the pixel block to a matching portion of the of the reference picture in the spherical-domain representation. The spherical-domain offset may be transformed to a motion vector in a source-domain representation of the input picture, and the pixel block may be coded predictively with reference to a source-domain representation of the matching portion of the reference picture.