Stereoscopic Caption Z-Plane Rendering for Occlusion-Free Stability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Captions in stereoscopic video can cause occlusion or eye fatigue due to improper placement in the Z-plane, and sudden changes in Z-plane position lead to viewer discomfort.
Innovation Solution
A system adjusts disparity values for caption rendering using a high-quality per-frame optical flow/disparity map and a bidirectional window to smooth out noisy measurements, ensuring captions are displayed at a just noticeable difference rate, maintaining stability and avoiding occlusion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If captions are placed too far from the viewer in the Z-plane, then occlusion by objects is reduced, but captions may be occluded by objects in front of the captions
Solution Approach 1:
The system dynamically adjusts the Z-plane position of captions frame-by-frame based on real-time disparity analysis. By computing the minimum disparity value across all objects in each frame and applying temporal smoothing, the caption position adapts to changing scene depth, ensuring captions remain visible without causing occlusion.
Solution Approach 2:
The system uses feedback from disparity map analysis to continuously optimize caption positioning. The rendered minimum disparity value from previous frames feeds into the current frame's caption position calculation, creating a closed-loop system that maintains optimal caption visibility while avoiding occlusion by dynamically responding to scene content.
2Reliability
If captions are placed too close to the viewer in the Z-plane, then captions are always visible, but eye fatigue results where the viewer has to focus on objects that are farther away
Solution Approach 1:
The system changes the disparity parameter of captions to match the scene's depth characteristics. By setting caption disparity to the minimum disparity value found in the scene plus an offset, captions are positioned at an appropriate depth relative to other objects, eliminating the need for viewers to refocus between near captions and distant objects, thus preventing eye fatigue.
3Adaptability or versatility
If the Z-plane position of captions changes drastically over successive frames, then captions can adapt to scene content, but eye fatigue occurs where the viewer is hunting for captions in different Z-plane positions
Solution Approach 1:
The system applies temporal smoothing to the disparity values before rendering captions. By averaging the minimum disparity values from multiple frames and applying a smoothing factor, the system cushions against abrupt changes in caption position, ensuring smooth transitions that prevent viewer discomfort while maintaining adaptation to scene content.
4Manufacturing precision
If disparity values are used directly for caption rendering, then caption positioning reflects scene depth, but noisy measurements cause unstable caption positions
Solution Approach 1:
The system applies temporal smoothing to cushion against noisy disparity measurements. By averaging disparity values across multiple frames with a smoothing factor, the system reduces measurement noise while preserving the underlying depth structure, resulting in stable and accurate caption positioning.
Solution Approach 2:
The system performs preliminary analysis to find the minimum disparity value across all objects in the scene before rendering captions. This preliminary depth characterization ensures that caption positioning is based on accurate scene understanding, placing captions at an appropriate depth relative to other objects.
Data Source
AI summary
In some embodiments, a method determines a disparity value from a plurality of disparity values in a current frame of a stereoscopic video. The disparity value is based on a difference of a value for a pixel between a first video and a second video of the stereoscopic video. A location is determined in a current frame that include the disparity value. The method analyzes first frames prior to the current frame to adjust disparity values in the first frames to generate one or more adjusted first disparity values. Also, the method analyzes second frames after the current frame to adjust disparity values in the second frames to generate one or more adjusted second disparity values. The one or more adjusted first disparity values and the one or more adjusted second disparity values are output for use in displaying captions in the first video or the second video.


