Volumetric Neural Network Compression for Stable Virtual Viewpoints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing volumetric representation neural networks face challenges in maintaining spatiotemporal consistency during compression, leading to reduced immersion and potential dizziness when replaying videos, and require efficient encoding/decoding methods to reduce transmission costs.
Innovation Solution
A method involving generating multi-view image sets, volumetric representation neural networks, and decomposing features into one-dimensional vectors or two-dimensional planes, followed by encoding these features using tensor decomposition and video codecs, while maintaining spatiotemporal consistency through feature relearning and prediction based on feature types.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If volumetric representation neural networks are transmitted instead of encoded images, then high-quality virtual viewpoint generation is achieved, but transmission cost increases significantly
Solution Approach 1:
The patent creates compressed representations (copies) of the volumetric neural network data through tensor decomposition and video coding techniques. Instead of transmitting the full high-precision neural network, a compressed version is transmitted that can be reconstructed at the receiving end, significantly reducing transmission cost while maintaining virtual viewpoint generation capability
Solution Approach 2:
The volumetric neural network data is segmented into multiple components through tensor decomposition (e.g., splitting into spatial, temporal, and feature dimensions). This segmentation allows selective transmission of essential components and enables efficient compression by treating different segments with appropriate coding strategies
2Productivity
If existing compression methods (quantization, pruning, tensor decomposition) are applied to volumetric representation neural networks, then transmission efficiency is improved, but spatiotemporal consistency is lost
Solution Approach 1:
The patent performs preliminary actions by grouping multi-view images into multi-view image sets before compression and maintaining their associations throughout the encoding process. This preliminary organization ensures that spatiotemporal relationships are preserved in the compressed representation, preventing consistency loss during decompression
Solution Approach 2:
The patent incorporates feedback mechanisms in the encoding process that monitor and preserve spatiotemporal consistency. By using reference frames and predictive coding that considers temporal and spatial relationships, the system can adjust compression parameters to maintain consistency while achieving efficient compression
3Device complexity
If multi-view images are compressed without considering spatiotemporal changes, then compression complexity is reduced, but video replay quality deteriorates causing dizziness and reduced immersion
Solution Approach 1:
The patent applies dynamic compression strategies that adapt to spatiotemporal changes in the multi-view image sequences. By analyzing motion and temporal correlations between frames, the system dynamically adjusts compression parameters to maintain quality in regions with significant changes while applying higher compression where appropriate, balancing complexity and quality
Data Source
AI summary
The present disclosure relates to a method and apparatus for volumetric representation neural network coding/decoding. A method for encoding a volumetric representation neural network according to one aspect of the present disclosure may include: generating one or more multi-view image sets by grouping a plurality of multi-view images; generating one or more volumetric representation neural networks expressing three-dimensional characteristics of the one or more multi-view image sets; generating features capable of reconstructing the one or more volumetric representation neural networks from the one or more volumetric representation neural networks; and encoding the features to generate a bitstream.


