Timed Text Rendering in 360 Video Using Spherical Coordinates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies face challenges in rendering timed text and graphics within virtual reality (VR) videos, particularly in 360° video environments, where text can become distorted and difficult to display effectively due to the spherical format, which is not well-handled by existing video coding standards.
Innovation Solution
The solution involves a control mechanism to adjust the display of timed text within a 360° video, using a cue box that can be altered for depth and distortion correction, and ensuring text is visible relative to the user's viewport, either always visible regardless of direction or only visible when looking in a specific direction, utilizing networked systems and client devices to transmit and render the text correctly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If timed text is embedded directly into 360° video using traditional video coding standards, then the text can be displayed in the video, but the text becomes distorted and difficult to read due to the spherical format
Solution Approach 1:
The patent transitions from 2D text rendering to 3D spherical coordinate system for text placement. Text is positioned using spherical coordinates (azimuth, elevation, radius) rather than traditional 2D screen coordinates, allowing text to be placed at specific locations in the 360° environment while maintaining proper orientation and readability relative to the user's viewpoint
Solution Approach 2:
The patent changes the coordinate system parameters from 2D Cartesian (x, y) to 3D spherical (azimuth, elevation, radius). It also dynamically adjusts text orientation parameters (roll, pitch, yaw) based on the user's head orientation and the text's position in the spherical space, ensuring text remains readable regardless of user movement
2Device complexity
If timed text position is fixed relative to the video frame, then the text rendering is simple, but the text may not be visible when the user looks in a different direction
Solution Approach 1:
The patent makes text positioning dynamic by linking text location and orientation to the user's viewpoint. As the user moves their head, the text automatically adjusts its position and orientation in the spherical coordinate system to maintain visibility and readability. The text can be configured to follow the user's gaze or maintain a fixed relationship to the user's head orientation
Solution Approach 2:
The system continuously receives feedback from head-tracking sensors that monitor user orientation and position. This feedback is used to dynamically recalculate and update text positioning and orientation parameters in real-time, ensuring the text remains visible and readable as the user moves through the 360° environment
3Productivity
If timed text is rendered without depth disparity adjustment, then the rendering process is simpler, but the text appears distorted in stereoscopic VR displays
Solution Approach 1:
The patent applies different rendering treatments to different parts of the text based on its position in the spherical coordinate system. Text elements are individually adjusted for depth disparity, orientation, and scale according to their specific location and the user's viewpoint, ensuring each portion of the text is rendered with appropriate quality for stereoscopic display
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An electronic device, a server and a method for rendering timed text within an omnidirectional video are disclosed. The method includes receiving a signaling message including a flag indicating whether a position of the timed text within the omnidirectional video is dependent on a viewport of the omnidirectional video. The method also includes determining whether the position of the timed text within the omnidirectional video is dependent on the viewport based on the flag. The method further includes rendering the timed text within the omnidirectional video based on the determination.