360-Degree Video Subtitle Adaptation and Multi-View Redundancy Removal
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current VR systems face challenges in providing efficient data transmission for large VR content, network robustness, and flexibility, especially for mobile reception, and lack adapted subtitle features for 360-degree video scenarios.
Innovation Solution
An apparatus and method for transmitting and receiving video that includes an inter-view redundancy remover, packing unit, encoder, decoder, un-packer, selector, block processor, merger, block boundary processor, and renderer, which handle multi-view video packing, inter-view redundancy, and signaling information to support interactive 360-degree video experiences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If general TTML based subtitles or bitmap based subtitles are used, then subtitle functionality is provided, but they are not adapted for 360-degree video scenarios
Solution Approach 1:
The patent changes the fundamental parameters of subtitle representation from traditional 2D TTML or bitmap formats to a 3D spatial coordinate system that matches the equirectangular projection of 360-degree video. Subtitles are defined by 3D coordinates (x, y, z) and orientation parameters that allow them to be correctly positioned and oriented in the spherical video space, making them adaptable to any viewport while maintaining high quality.
2Productivity
If multi-view video data is transmitted without redundancy removal, then all viewing position information is preserved, but data transmission efficiency decreases
Solution Approach 1:
The patent uses the primary view as a reference copy and generates subsidiary views by copying and transforming this reference. Instead of transmitting independent full-resolution data for each viewing position, the system transmits the primary view and then transmits only the transformation parameters (rotation angles, translation vectors) needed to generate other views at the receiver端, dramatically improving transmission efficiency while preserving all viewing position information.
Solution Approach 2:
The system performs preliminary encoding of the primary view and pre-calculates the transformation relationships between different viewing positions. By preparing the reference data and transformation rules in advance, the receiver can efficiently generate any required view without needing to receive and process complete data for all possible viewing positions.
Data Source
AI summary
An apparatus for transmitting a video comprises an inter-view redundancy remover configured to remove redundant information of pictures for viewing positions, where redundant pixel information between the pictures in adjacent viewing positions is removed, a packing unit configured to pack the pictures and generate packing information, and an encoder configured to encode the pictures. The apparatus is further configured to perform multi-view packing the pictures into a packed picture, where each view for the picture includes different types of a texture and a depth map, and a residual of texture and a depth map are generated for a subsidiary view based on redundancy between each view. The apparatus is further configured to generate signaling information for the inter-view redundancy remover, the packing unit, or the multi-view packing.


