Light Field Image Encoding via Graph Signal Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing light field image compression techniques, such as JPEG and video coding standards like AVC and HEVC, introduce redundancy and artifacts due to demosaicing and scaling, which damage angular information and limit parallax in light field images.
Innovation Solution
The method employs graph signal processing (GSP) techniques to generate a compact light field image representation that avoids redundancy by spatially displacing pixels to create a transformed multi-color image with dummy pixels, allowing for efficient encoding and decoding without demosaicing, and uses graph-based compression methods like Graph Fourier Transform (GFT) or Graph based Lifting Transform (GLT) to encode sub-views.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional image compression techniques (JPEG, AVC, HEVC) are used to compress light field images, then compression is achieved, but redundancy is introduced and artifacts are generated that damage angular information and limit parallax
Solution Approach 1:
The light field image is segmented into multiple sub-views based on angular information. Each sub-view represents a different angular perspective of the scene. By segmenting the 4D light field data into multiple 2D sub-views, the patent preserves angular information structure while enabling targeted compression strategies for each view, avoiding the information loss that occurs when treating the light field as a single conventional image.
Solution Approach 2:
The patent transforms the 4D light field data into multiple 2D sub-views, effectively projecting higher-dimensional information into lower-dimensional representations that can be processed by conventional compression techniques. This dimensional transformation allows the angular information to be preserved in the spatial arrangement of sub-views while reducing the overall data complexity for compression.
2Ease of manufacture
If demosaicing and scaling are applied during compression, then color patterned raw lenselet images can be processed, but artifacts are introduced that damage angular information
Solution Approach 1:
The patent performs preliminary organization of color patterned raw lenselet images into sub-views before applying compression. By pre-arranging the data in a structured sub-view format, the patent enables direct compression of the organized data without requiring subsequent demosaicing and scaling operations that would introduce artifacts and damage angular information.
Solution Approach 2:
The patent extracts angular information from color patterned raw lenselet images by organizing pixels into distinct sub-views based on their angular coordinates. This extraction process separates the angular dimension from the spatial dimension, allowing compression to be applied to each sub-view independently without the need for demosaicing that would mix angular information and introduce artifacts.
3Quantity of substance
If light field images are compressed as single 2D images, then compression is achieved, but the 4D nature and angular information are lost
Solution Approach 1:
The patent segments the 4D light field image into multiple 2D sub-views, each representing a different angular perspective. This segmentation preserves the angular information structure by maintaining separate representations for different viewing angles, unlike treating the light field as a single 2D image which would collapse all angular information into one view.
Solution Approach 2:
The patent handles the 4D light field data by transforming it into multiple 2D sub-views, effectively using the arrangement and organization of these sub-views to represent the fourth dimension (angular space). This approach reduces data size for compression while preserving angular information through the structured arrangement of sub-views rather than losing it in a single 2D projection.
Data Source
Figure 1(a)~1(b)
Figure 2(a)~2(c)
Figure 3(a)~3(b)
AI summary
A method for encoding a raw lenselet image (f), comprising a receiving phase (705), wherein at least a portion of a raw lenselet image (f) is received, said image (f) comprising a plurality of macro-pixels (650, 1320, 1340), each macro-pixel (650, 1320, 1340) comprising pixels corresponding to a specific view angle for the same point of a scene and an output phase, wherein a bitstream (fd^) comprising at least a portion of an encoded lenselet image (f) is outputted. The method comprises an image transform phase (710), wherein the pixels of said raw lenselet image (f) are spatially displaced in a transformed multi-color image (660) having a larger number of columns and rows with respect to the received raw lenselet image, wherein dummy pixels (610, 620) having undefined value are inserted into said raw lenselet image (f) and wherein said displacement is performed so as to put the estimated center location of each macro-pixel (650, 1320, 1340) onto integer pixel locations. Moreover, the method comprises a sub-view generation phase, wherein a sequence of sub-views (fd ') is generated, said sub-views (1310, 1510, 1620) comprising pixels of the same angular coordinates extracted from different macro-pixels (650,1320,1340) of said transformed raw lenselet image (f). Finally, the method comprises a graph coding phase, wherein a bitstream (fd^) is generated by encoding a graph representation of at least one of the sub-views (1310, 1510, 1620) of said sequence (fd' ) according to a predefined graph signal processing (GSP) technique. The method comprises outputting said graph-coded bitstream (fd^) for its transmission and/or storage.