Multi-dimensional Object Representation Compression for AR
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current compression techniques are insufficient for efficiently sending large numbers of images to client devices for local rendering, particularly on mobile devices, which results in resource-intensive transmission and processing challenges due to large file sizes and memory limitations, impacting real-time user experiences.
Innovation Solution
The proposed solution involves encoding object images and segmentation masks into video files, using a spiral configuration and partitioning them into patches, with keyframes for quick frame retrieval, to reduce file size and improve rendering speed on client devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If multiple object images are sent to client device for photorealistic rendering, then rendering quality is improved, but transmission resource consumption increases
Solution Approach 1:
The patent segments the object representation into multiple views or frames that can be compressed and transmitted efficiently. Instead of sending complete high-resolution images, the system divides the object into viewable segments that can be reconstructed on the client device, reducing transmission data while maintaining rendering quality.
Solution Approach 2:
The patent transitions from transmitting complete 2D images to transmitting compressed 3D model data or view synthesis parameters. By changing the dimensionality of the transmitted data from full image pixels to geometric or parametric representations, the system reduces transmission resources while enabling photorealistic rendering through view synthesis algorithms.
2Manufacturing precision
If large number of images are transmitted to client device, then photorealistic rendering is improved, but file size increases
Solution Approach 1:
Instead of transmitting multiple complete image copies, the patent transmits a compressed representation or template that can be used to generate or copy views on the client device. The system sends essential geometric and textural data that enables local rendering of multiple views, dramatically reducing file size while maintaining photorealistic quality.
Solution Approach 2:
The patent performs preliminary processing and compression of object data on the server side before transmission. By pre-computing and encoding the object representation in a compressed format suitable for client-side rendering, the system reduces the file size transmitted while ensuring photorealistic rendering capability is preserved through pre-prepared rendering assets.
3Manufacturing precision
If multiple object images are sent for local rendering, then rendering capability is improved, but processing time increases
Solution Approach 1:
The patent performs preliminary processing of object data into optimized formats during server-side encoding. By pre-computing geometric structures, material properties, and view synthesis parameters before transmission, the system enables faster client-side rendering while maintaining high rendering capability. The preprocessing shifts computational burden from client to server, reducing client processing time.
4Measurement precision
If complete object images are transmitted, then rendering accuracy is improved, but memory consumption on mobile devices increases
Solution Approach 1:
The patent segments the complete object representation into essential geometric structures and textural details that can be stored in compressed formats. By dividing the object model into hierarchical levels of detail and transmitting only necessary components for accurate rendering, the system maintains rendering accuracy while reducing memory consumption on mobile devices.
Solution Approach 2:
The patent changes the representation parameters from storing complete pixel data to storing compressed geometric and material parameters. By transforming the object representation from image-space to object-space parameters, the system achieves the same rendering accuracy with significantly reduced memory footprint, enabling mobile device compatibility.
Data Source
AI summary
Objects can be rendered in three dimensions and viewed and manipulated in an augmented reality environment. A number of object images, a number of segmentation masks, and an object mesh structure are used by a client device to render the object in three dimensions. The object images and segmentation masks can be sequenced into frames. The object images and segmentation masks can be partitioned into patches and sequenced, or ordered, within each patch, and a keyframe can be assigned in each patch. Then, the object images and segmentation masks can be encoded into video files and sent to a client device. The client device can quickly retrieve a requested object image and segmentation mask based at least in part on identifying the keyframe in the same patch as the object image and segmentation mask.


