Light Field Image Encoding Using Epipolar Plane Images

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image and video codecs are inadequate for encoding and decoding light field based images, failing to capture the directional distribution of light rays and resulting in loss of depth information, which limits advanced imaging functionalities such as real-time focus and perspective manipulation.

Innovation Solution

A method for predicting blocks of pixels in light field data using epipolar plane images, determining optimal unidirectional prediction modes based on energy levels from neighboring pixels, and encoding residual errors to enhance image rendering and manipulation capabilities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If conventional image and video codecs are used to encode light field based images, then the encoding process is simple and compatible with existing standards, but the directional distribution of light rays and depth information are lost

Engineering Contradiction:
Improvedepth informationVSAvoidencoding complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the light field data into multiple views arranged in a matrix structure, where each view represents an image captured from a different viewpoint. This segmentation allows the system to process and encode light field data by handling individual views separately while maintaining the relationships between them, thereby preserving depth information through the multi-view geometry rather than using conventional single-image encoding methods

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from conventional two-dimensional image encoding to four-dimensional light field encoding by adding spatial and angular dimensions. The light field is represented as a function of position and direction, creating a 4D data structure that conventional 2D codecs cannot process. This dimensional transformation enables preservation of depth and directional information that would otherwise be lost

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If light field data is captured using plenoptic devices or camera arrays, then directional distribution and depth information are preserved, but the data volume and processing complexity increase significantly

Engineering Contradiction:
Improvedirectional distribution informationVSAvoiddata volume
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent merges multiple camera views into a unified light field representation by arranging them in a matrix structure and establishing geometric relationships between adjacent views. This merging process allows the system to process and encode light field data more efficiently by leveraging the correlations between neighboring views, thereby reducing the effective data volume that needs to be processed while preserving directional distribution information

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transforms the representation parameters of light field data from raw multi-view images to a compressed format that encodes depth maps, surface normals, and other geometric parameters. By changing the parameter space from pixel values to geometric descriptors, the system reduces data volume while maintaining the essential directional and depth information needed for rendering and manipulation

Inventive Principle:
Principle #35Parameter changes

3Productivity

If standard image codecs are used, then encoding and decoding are fast and efficient, but advanced imaging functionalities like real-time focus and perspective manipulation cannot be achieved

Engineering Contradiction:
Improveencoding speedVSAvoidimaging functionality
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent performs preliminary encoding of light field data into a format that preserves geometric relationships and depth information, enabling subsequent post-processing operations like refocusing and perspective changes. By preparing the data structure in advance with proper organization of multiple views and geometric parameters, the system enables fast and efficient advanced imaging functionalities without requiring re-encoding operations

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a universal encoding framework that can handle both conventional imaging tasks and advanced light field operations through the same data structure. The matrix of views representation and associated geometric parameters serve multiple purposes: they enable depth estimation, refocusing, perspective transformation, and other advanced functionalities, making the encoding system versatile rather than specialized for a single application

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10785502B2Method and apparatus for encoding and decoding a light field based image, and corresponding computer program product
Publication Date: 2020.09.22 INTERDIGITAL VC HOLDINGS INC
  • US10785502B2 patent drawing
  • US10785502B2 patent drawing
  • US10785502B2 patent drawing

AI summary

The present disclosure generally relates to a method for predicting at least one block of pixels of a view (170) belonging to a matrix of views (17) obtained from light-field data belong with a scene. According to present disclosure, the method is implemented by a processor and comprises for at least one pixel to predict of said block of pixels: —from said matrix of views (17), obtaining (51) at least one epipolar plane image (EPI) belong with said pixel to predict, —among a set of unidirectional prediction modes, determining (52) at least one optimal unidirectional prediction mode from a set of previous reconstructed pixels neighboring said pixel to predict in said at least one epipolar plane image, —extrapolating (53) a prediction value of said pixel to predict by using said at least one optimal unidirectional prediction mode.