Deep Learning Light Field Compression for Multi-Layer Tensor Displays
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing light field compression methods for multi-layer 3D displays, such as tensor displays, are inefficient and result in poor pixel estimates due to their highly redundant nature, with current approaches like sequence-based compression and least-squares algorithms being sub-optimal.
Innovation Solution
A machine learning model, specifically a deep learning deformable SAI feature embedding network, is trained to optimize pixel representations for each layer of a 3D display, utilizing deformable convolution and residual learning to minimize loss functions, enabling efficient compression and encoding aligned with the properties of multi-layer displays.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If sequence-based compression of Synthetic Aperture Image (SAI) is employed, then compression is achieved, but encoding efficiency is poor requiring 169 pieces of information for a 17×17 SAI
Solution Approach 1:
The patent segments the light field data into multiple layers corresponding to different depth planes in the tensor display. Instead of compressing the entire SAI as a single sequence, the data is divided into L layers, each representing a specific depth plane. This segmentation allows independent processing and optimization of each layer, improving encoding efficiency while maintaining compression.
Solution Approach 2:
The patent transforms the traditional 2D SAI representation into a 3D tensor representation by adding the depth dimension. The light field data is reorganized from a 2D synthetic aperture image into a 3D tensor with dimensions corresponding to spatial coordinates and depth layers, enabling more efficient compression by exploiting the additional dimensional structure.
2Quantity of substance
If least-squares algorithm is used to compress LF data for tensor display, then compression is achieved, but pixel estimates for multi-layer displays are poor
Solution Approach 1:
The patent performs preliminary action by pre-computing and storing depth information and layer assignments for each pixel during the encoding phase. This preliminary processing allows the decoder to accurately reconstruct pixel values for each layer without relying solely on least-squares estimation, thereby improving pixel estimation accuracy while maintaining compression efficiency.
Solution Approach 2:
The patent introduces an intermediary depth map and layer assignment structure that mediates between the compressed light field data and the final multi-layer display reconstruction. This intermediary representation provides additional structural information that guides the reconstruction process, improving pixel estimation accuracy beyond what least-squares alone can achieve.
3Quantity of substance
If traditional compression schemes are used for light field data, then storage and communication are enabled, but computing and network resources are excessive
Solution Approach 1:
The patent changes the representation parameters by transforming light field data from traditional SAI format into tensor format with explicit depth layering. This parameter change enables more efficient compression by exploiting the structured nature of tensor representations, reducing the computational resources needed for both compression and decompression operations.
Solution Approach 2:
The patent creates a universal tensor-based representation framework that serves multiple functions: compression, depth information extraction, and layer-specific processing. This multi-functional approach eliminates the need for separate processing pipelines, reducing overall computing resources and network bandwidth requirements.
Data Source
AI summary
Systems, methods and apparatuses are described herein for training a machine learning model to accept as input synthetic aperture image (SAI) training data for a three-dimensional (3D) display, the 3D display comprising a plurality of layers. The machine learning model may be trained to output respective pixel representations of the SAI training data for each of the plurality of layers of the 3D display. The provided systems, methods and apparatuses may access image data, input the image data to the trained machine learning model, and determine, using the trained machine learning model, respective pixel representations of the input image data for each of the plurality of layers of the 3D display. The provided systems, methods and apparatuses may encode the respective pixel representations of the input image data, and transmit, for display at the 3D display, the encoded respective pixel representations of the input image data.


