Virtual View Rendering Text Positioning Depth Warping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for virtual view synthesis in stereoscopic and multi-view displays often result in erroneous depth maps, leading to improperly positioned text characters in rendered images due to errors in depth data, which degrades user comfort, especially for salient objects like subtitles.
Innovation Solution
A method and system for virtual view rendering that detects text in input videos and warps it with a constant depth distribution independent of the background, using depth data post-processing to prevent character shifts and maintain accurate text positioning.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If depth base rendering is used to generate virtual views, then viewing angle extension and depth adjustment are improved, but text positioning accuracy deteriorates due to erroneous depth maps
Solution Approach 1:
The patent segments the image into text regions and non-text regions, applying different depth processing strategies to each. Text regions are identified through text detection algorithms and processed separately with constant depth values, while non-text regions use the original depth map, thereby resolving the contradiction between viewing angle extension and text positioning accuracy
Solution Approach 2:
The patent applies local quality by using constant depth values specifically for text regions while maintaining variable depth values for non-text regions. This localized differentiation ensures that text positioning remains accurate while still allowing depth-based rendering effects in other areas of the image
2Ease of manufacture
If text regions are processed as whole boxes, then text detection simplicity is improved, but text-background separation deteriorates causing mixing
Solution Approach 1:
The patent extracts text regions from the overall image processing flow by detecting text boundaries and creating separate masks for text and non-text areas. This extraction allows independent processing of text regions with constant depth values, preventing text-background mixing while maintaining detection simplicity
Solution Approach 2:
The patent introduces text detection masks and region segmentation as intermediary elements between the original image and the final rendered output. These intermediaries enable precise control over which regions receive constant depth treatment versus variable depth treatment, achieving clean text-background separation
Data Source
Figure 1~2
Figure 3
AI summary
Present invention discloses method and system for virtual view rendering, in which input video (101) with text is rendered in accordance with depth data (106). Said method comprises steps of; determining format of the input video (101) in terms of view numbers (102); detecting text on input video (101) (103) by performing text detection algorithm in accordance with determined format of the input video (101); generating rendered view (107) by performing depth data post processing (105). Said system comprises means for determining the format of an input video (101) in terms of view number; means for detecting at least one text on said video in accordance with determined format; means for depth data post processing, which warps detected text with constant depth value.