Joint Depth and Texture Motion Vector Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multiview video coding standards, such as MVC, do not efficiently code depth map videos and texture videos together, limiting bandwidth efficiency and preventing motion prediction between depth map and texture images.

Innovation Solution

Implement joint coding of depth map video and texture video, where motion vectors are predicted between the two, with scenarios including SVC-compliant configurations where depth map video is coded as a base layer and texture video as an enhancement layer, and vice versa, as well as inter-layer and inter-view prediction methods to optimize motion prediction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If separate coding of depth map video and texture video is used, then coding simplicity is maintained, but bandwidth efficiency deteriorates and motion prediction between depth and texture is prevented

Engineering Contradiction:
Improvebandwidth efficiencyVSAvoidcoding complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The patent merges the coding processes of depth map video and texture video into a unified joint coding framework. This allows motion vectors from depth map coding to be reused for texture video prediction, eliminating redundant motion estimation computations and improving bandwidth efficiency through shared coding resources and structures

Inventive Principle:
Principle #5Merging (Combining)

2Productivity

If motion vectors are shared between depth map and texture video, then coding efficiency is improved, but prediction accuracy may deteriorate due to motion differences between depth and texture

Engineering Contradiction:
Improvecoding efficiencyVSAvoidprediction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies local quality by allowing different prediction strategies for different regions. Motion vectors are shared where depth and texture motions align, while region-specific adjustments are applied where local motion differences exist, maintaining both coding efficiency and prediction accuracy through spatially adaptive prediction

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent modifies motion vector parameters by applying offsets and adjustments to the shared motion vectors based on local motion characteristics. This allows the base motion vectors from depth coding to be adapted for texture prediction, preserving coding efficiency while improving prediction accuracy through parameter optimization

Inventive Principle:
Principle #35Parameter changes

3Loss of energy

If joint coding of depth and texture is implemented, then bandwidth efficiency is enhanced, but device complexity increases due to inter-layer and inter-view prediction mechanisms

Engineering Contradiction:
Improvebandwidth efficiencyVSAvoidprediction mechanism complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by pre-computing and storing motion vectors from depth map coding that can be reused for texture video prediction. This preliminary computation avoids redundant motion estimation in texture coding, reducing overall computational complexity while maintaining bandwidth efficiency through pre-prepared prediction data

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10715779B2Sharing of motion vector in 3D video coding
Publication Date: 2020.07.14 NOKIA TECHNOLOGIES OY
  • US10715779B2 patent drawing
  • US10715779B2 patent drawing
  • US10715779B2 patent drawing

AI summary

Joint coding of depth map video and texture video is provided, where a motion vector for a texture video is predicted from a respective motion vector of a depth map video or vice versa. For scalable video coding, depth map video is coded as a base layer and texture video is coded as an enhancement layer(s). Inter-layer motion prediction predicts motion in texture video from motion in depth map video. With more than one view in a bitstream (for multiview coding), depth map videos are considered monochromatic camera views and are predicted from each other. If joint multiview video model coding tools are allowed, inter-view motion skip is used to predict motion vectors of texture images from depth map images. Furthermore, scalable multiview coding is utilized, where inter-view prediction is applied between views in the same dependency layer, and inter-layer (motion) prediction is applied between layers in the same view.