Image Signal Generation Using Prediction Quality Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current virtual reality applications face challenges in providing efficient and immersive experiences, particularly when based on real-world captures, due to high computational and communication resource requirements, and suboptimal user experiences with reduced quality and restricted freedom.
Innovation Solution
An apparatus and method for generating an image signal that includes a receiver for source images from different view poses, a combined image generator that creates images from multiple source images, an evaluator for determining prediction quality measures, a determiner for identifying segments with high prediction errors, and an image signal generator that combines these elements to reduce data rate while maintaining image quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple source images from different view poses are captured to provide immersive virtual reality experience, then image quality and viewing freedom are improved, but data rate and communication resource requirements increase
Solution Approach 1:
The image signal is segmented into multiple layers including a base layer and enhancement layers. Each layer serves a specific function: the base layer provides essential scene information from a reference view pose, while enhancement layers provide additional view pose information. This segmentation allows progressive transmission where the base layer ensures basic functionality with lower data rate, and enhancement layers progressively improve viewing freedom and image quality as bandwidth allows.
Solution Approach 2:
Different parts of the image signal are transmitted with different qualities and data rates. The base layer transmits critical scene information at higher quality, while enhancement layers transmit supplementary view information at variable qualities. This local quality approach ensures that essential viewing experience is maintained with lower data rate, while enhanced viewing freedom is provided where bandwidth permits.
2Adaptability or versatility
If multiple source images from different view poses are captured to allow dynamic viewing position changes, then user experience is improved, but computational resource requirements increase
Solution Approach 1:
Image data from multiple view poses is captured and processed in advance during the encoding phase. The encoder performs view synthesis and generates multiple layers of image data representing different view poses before transmission. This preliminary action shifts the computational burden to the encoding stage, allowing the decoder to operate with reduced computational resources while still providing dynamic viewing position changes.
Solution Approach 2:
The patent introduces an intermediary representation format that bridges the gap between multiple source images and the final rendered view. This intermediary format uses a base layer with reference view data and enhancement layers with view synthesis data, allowing the system to manage computational complexity by processing information in staged layers rather than requiring all source images to be simultaneously processed.
3Adaptability or versatility
If a predetermined virtual model of the scene is used for virtual reality application, then rendering flexibility is improved, but image quality and realism are reduced
Solution Approach 1:
The patent merges the advantages of both virtual models and real captured images by combining synthesized view data from a virtual model with actual captured image data from multiple view poses. The base layer may use virtual model synthesis to provide rendering flexibility, while enhancement layers incorporate real captured image data to maintain image quality and realism. This merging allows the system to achieve both rendering flexibility and high image quality.
Data Source
AI summary
Generating an image signal comprises a receiver (401) receiving source images representing a scene. A combined image generator (403) generates combined images from the source images. Each combined image is derived from only parts of at least two images of the source images. An evaluator (405) determines prediction quality measures for elements of the source images where the prediction quality measure for an element of a first source image is indicative of a difference between pixel values in the first source image and predicted pixel values for pixels in the element. The predicted pixel values are pixel values resulting from prediction of pixels from the combined images. A determiner (407) determines segments of the source images comprising elements for which the prediction quality measure is indicative of a difference above a threshold. An image signal generator (409) generates an image signal comprising image data representing the combined images and the segments of the source images.


