Subtitle Placement via Audio Source Position Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing subtitle display technologies in DVB-based broadcasting face challenges in effectively positioning subtitles relative to the speaker's location in video images, particularly when 3D audio is integrated, as they lack precise synchronization and positioning mechanisms for accurate subtitle placement.
Innovation Solution
A transmission apparatus and method that encodes video, subtitle, and audio streams, incorporating metadata to associate subtitle data with audio data and position information, allowing for precise control of subtitle placement on the receiving end based on sound source position, using a container stream format that includes video, subtitle, and audio streams with inserted metainformation for accurate synchronization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If subtitle information is sent using bitmap data in DVB-based broadcasting, then subtitle display is simple and straightforward, but subtitle positioning relative to speaker location cannot be achieved
Solution Approach 1:
The subtitle data is segmented into multiple regions (first subtitle region and second subtitle region) corresponding to different speakers. Each region can be independently positioned and controlled, allowing subtitles to be accurately placed relative to specific speaker locations while maintaining manageable data structure.
Solution Approach 2:
Different subtitle regions are assigned different positional characteristics based on their corresponding speakers' locations. The first subtitle region is positioned according to the first speaker's location, and the second subtitle region according to the second speaker's location, enabling precise local positioning rather than uniform placement.
2Manufacturing precision
If text-based subtitle information is used with font expansion to match receiving side resolution, then subtitle quality improves, but synchronization with audio and positioning becomes complex
Solution Approach 1:
Position information for each subtitle region is pre-calculated and embedded in the transmitted data based on speaker locations in the video. The receiving device simply needs to retrieve and apply this pre-prepared position information, avoiding complex real-time calculations and reducing synchronization complexity.
Solution Approach 2:
Position information acts as an intermediary that links subtitle regions to speaker locations. This intermediary data structure simplifies the relationship between subtitles and audio sources, making synchronization and positioning more manageable by providing explicit mapping information.
3Adaptability or versatility
If 3D audio with position information is integrated, then spatial audio experience improves, but subtitle alignment with speaker position becomes difficult
Solution Approach 1:
The patent merges the positioning systems of 3D audio and subtitles by using the same speaker position information for both audio rendering and subtitle placement. This unified approach ensures that subtitles and audio sources are aligned in the spatial domain, maintaining consistency across different media types.
Solution Approach 2:
The position information extracted from video data serves multiple functions: it is used both for 3D audio rendering (mapping audio to speaker positions) and for subtitle positioning (placing subtitle regions). This multi-functional use of the same data simplifies the system while maintaining alignment accuracy.
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
Effective display of subtitles is permitted on a receiving side. A video stream is generated that has video coded data. A subtitle stream is generated that has subtitle data corresponding to a speech of a speaker. An audio stream is generated that has object coded data including audio data whose sound source is a speech of a speaker and position information regarding the sound source. A container stream in a given format is sent that includes each of the streams. Metainformation is inserted, into a layer of the container stream, for associating subtitle data corresponding to each speech included in the subtitle stream and audio data corresponding to each speech included in the audio stream.