Light Field Image Encoding via Graph Signal Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing light field image compression techniques, such as JPEG and video coding standards like AVC and HEVC, introduce redundancy and artifacts due to demosaicing and scaling, which damage angular information and limit parallax in light field images.

Innovation Solution

The method employs graph signal processing (GSP) techniques to generate a compact light field image representation that avoids redundancy by spatially displacing pixels to create a transformed multi-color image with dummy pixels, allowing for efficient encoding and decoding without demosaicing, and uses graph-based compression methods like Graph Fourier Transform (GFT) or Graph based Lifting Transform (GLT) to encode sub-views.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If conventional image compression techniques (JPEG, AVC, HEVC) are used to compress light field images, then compression is achieved, but redundancy is introduced and artifacts are generated that damage angular information and limit parallax

Engineering Contradiction:
Improvedata redundancyVSAvoidangular information
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The light field image is segmented into multiple sub-views based on angular information. Each sub-view represents a different angular perspective of the scene. By segmenting the 4D light field data into multiple 2D sub-views, the patent preserves angular information structure while enabling targeted compression strategies for each view, avoiding the information loss that occurs when treating the light field as a single conventional image.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the 4D light field data into multiple 2D sub-views, effectively projecting higher-dimensional information into lower-dimensional representations that can be processed by conventional compression techniques. This dimensional transformation allows the angular information to be preserved in the spatial arrangement of sub-views while reducing the overall data complexity for compression.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of manufacture

If demosaicing and scaling are applied during compression, then color patterned raw lenselet images can be processed, but artifacts are introduced that damage angular information

Engineering Contradiction:
Improveprocessing capabilityVSAvoidangular information
Core Design Contradiction:
Ease of manufactureVSLoss of information

Solution Approach 1:

The patent performs preliminary organization of color patterned raw lenselet images into sub-views before applying compression. By pre-arranging the data in a structured sub-view format, the patent enables direct compression of the organized data without requiring subsequent demosaicing and scaling operations that would introduce artifacts and damage angular information.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts angular information from color patterned raw lenselet images by organizing pixels into distinct sub-views based on their angular coordinates. This extraction process separates the angular dimension from the spatial dimension, allowing compression to be applied to each sub-view independently without the need for demosaicing that would mix angular information and introduce artifacts.

Inventive Principle:
Principle #2Taking out (Extraction)

3Quantity of substance

If light field images are compressed as single 2D images, then compression is achieved, but the 4D nature and angular information are lost

Engineering Contradiction:
Improvedata sizeVSAvoidangular information
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent segments the 4D light field image into multiple 2D sub-views, each representing a different angular perspective. This segmentation preserves the angular information structure by maintaining separate representations for different viewing angles, unlike treating the light field as a single 2D image which would collapse all angular information into one view.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent handles the 4D light field data by transforming it into multiple 2D sub-views, effectively using the arrangement and organization of these sub-views to represent the fourth dimension (angular space). This approach reduces data size for compression while preserving angular information through the structured arrangement of sub-views rather than losing it in a single 2D projection.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentEP3607745B1Methods and apparatuses for encoding and decoding digital light field images
Publication Date: 2024.11.20 SISVEL TECH
  • EP3607745B1 patent drawingFigure 1(a)~1(b)
  • EP3607745B1 patent drawingFigure 2(a)~2(c)
  • EP3607745B1 patent drawingFigure 3(a)~3(b)

AI summary

A method for encoding a raw lenselet image (f), comprising a receiving phase (705), wherein at least a portion of a raw lenselet image (f) is received, said image (f) comprising a plurality of macro-pixels (650, 1320, 1340), each macro-pixel (650, 1320, 1340) comprising pixels corresponding to a specific view angle for the same point of a scene and an output phase, wherein a bitstream (fd^) comprising at least a portion of an encoded lenselet image (f) is outputted. The method comprises an image transform phase (710), wherein the pixels of said raw lenselet image (f) are spatially displaced in a transformed multi-color image (660) having a larger number of columns and rows with respect to the received raw lenselet image, wherein dummy pixels (610, 620) having undefined value are inserted into said raw lenselet image (f) and wherein said displacement is performed so as to put the estimated center location of each macro-pixel (650, 1320, 1340) onto integer pixel locations. Moreover, the method comprises a sub-view generation phase, wherein a sequence of sub-views (fd ') is generated, said sub-views (1310, 1510, 1620) comprising pixels of the same angular coordinates extracted from different macro-pixels (650,1320,1340) of said transformed raw lenselet image (f). Finally, the method comprises a graph coding phase, wherein a bitstream (fd^) is generated by encoding a graph representation of at least one of the sub-views (1310, 1510, 1620) of said sequence (fd' ) according to a predefined graph signal processing (GSP) technique. The method comprises outputting said graph-coded bitstream (fd^) for its transmission and/or storage.