Subtitle Placement via Audio Source Position Metadata

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing subtitle display technologies in DVB-based broadcasting face challenges in effectively positioning subtitles relative to the speaker's location in video images, particularly when 3D audio is integrated, as they lack precise synchronization and positioning mechanisms for accurate subtitle placement.

Innovation Solution

A transmission apparatus and method that encodes video, subtitle, and audio streams, incorporating metadata to associate subtitle data with audio data and position information, allowing for precise control of subtitle placement on the receiving end based on sound source position, using a container stream format that includes video, subtitle, and audio streams with inserted metainformation for accurate synchronization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If subtitle information is sent using bitmap data in DVB-based broadcasting, then subtitle display is simple and straightforward, but subtitle positioning relative to speaker location cannot be achieved

Engineering Contradiction:
Improvesubtitle display simplicityVSAvoidsubtitle positioning accuracy
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The subtitle data is segmented into multiple regions (first subtitle region and second subtitle region) corresponding to different speakers. Each region can be independently positioned and controlled, allowing subtitles to be accurately placed relative to specific speaker locations while maintaining manageable data structure.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different subtitle regions are assigned different positional characteristics based on their corresponding speakers' locations. The first subtitle region is positioned according to the first speaker's location, and the second subtitle region according to the second speaker's location, enabling precise local positioning rather than uniform placement.

Inventive Principle:
Principle #3Local quality

2Manufacturing precision

If text-based subtitle information is used with font expansion to match receiving side resolution, then subtitle quality improves, but synchronization with audio and positioning becomes complex

Engineering Contradiction:
Improvesubtitle qualityVSAvoidsynchronization and positioning complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

Position information for each subtitle region is pre-calculated and embedded in the transmitted data based on speaker locations in the video. The receiving device simply needs to retrieve and apply this pre-prepared position information, avoiding complex real-time calculations and reducing synchronization complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Position information acts as an intermediary that links subtitle regions to speaker locations. This intermediary data structure simplifies the relationship between subtitles and audio sources, making synchronization and positioning more manageable by providing explicit mapping information.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If 3D audio with position information is integrated, then spatial audio experience improves, but subtitle alignment with speaker position becomes difficult

Engineering Contradiction:
Improvespatial audio capabilityVSAvoidsubtitle-speaker alignment accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent merges the positioning systems of 3D audio and subtitles by using the same speaker position information for both audio rendering and subtitle placement. This unified approach ensures that subtitles and audio sources are aligned in the spatial domain, maintaining consistency across different media types.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The position information extracted from video data serves multiple functions: it is used both for 3D audio rendering (mapping audio to speaker positions) and for subtitle positioning (placing subtitle regions). This multi-functional use of the same data simplifies the system while maintaining alignment accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3393130B1Transmission device, transmission method, receiving device and receiving method for associating subtitle data with corresponding audio data
Publication Date: 2020.04.29 SONY GROUP CORP
  • EP3393130B1 patent drawingFigure 1
  • EP3393130B1 patent drawingFigure 2
  • EP3393130B1 patent drawingFigure 3~4

AI summary

Effective display of subtitles is permitted on a receiving side. A video stream is generated that has video coded data. A subtitle stream is generated that has subtitle data corresponding to a speech of a speaker. An audio stream is generated that has object coded data including audio data whose sound source is a speech of a speaker and position information regarding the sound source. A container stream in a given format is sent that includes each of the streams. Metainformation is inserted, into a layer of the container stream, for associating subtitle data corresponding to each speech included in the subtitle stream and audio data corresponding to each speech included in the audio stream.