Dynamic Caption Generation Adapting to Speaker Voice State

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current caption generation methods for broadcasting services lack dynamic adaptation to the voice state of speakers, such as volume, tone, and emotion, which limits the engagement and visual appeal of captions in real-time broadcasts.

Innovation Solution

A method and apparatus that generate caption text and style information based on the voice state of speakers, including control over size, color, font, position, rotation, and special effects, by analyzing voice states and applying corresponding changes to the caption text and screen style in real-time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If static caption generation is used, then the caption display is simple and stable, but the viewer engagement and visual appeal are limited

Engineering Contradiction:
Improveviewer engagementVSAvoidcaption generation system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies dynamics by making the caption style adjustable based on voice state changes. The system transitions from static caption display to dynamic caption generation where caption characteristics (size, color, font, position, rotation, special effects) automatically adapt in real-time according to the speaker's voice state, thereby enhancing viewer engagement without requiring complex manual intervention

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent implements parameter changes by modifying multiple caption parameters simultaneously based on voice state analysis. When voice state changes are detected (volume, tone, emotion), the system changes corresponding caption parameters such as size, color, font, position, rotation, and special effects, creating an adaptive captioning system that responds to spoken content characteristics

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If real-time voice state analysis is implemented, then caption adaptability is improved, but processing time and computational resources increase

Engineering Contradiction:
Improvecaption adaptabilityVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system maintains continuous voice state analysis throughout the broadcast, enabling real-time caption adaptation without interruption. The useful action of analyzing voice state and generating adaptive captions continues seamlessly as the broadcast progresses, ensuring that captions always reflect the current speaker state without requiring periodic pauses or batch processing

Inventive Principle:
Principle #20Continuity of useful action

3Adaptability or versatility

If multiple caption style options are provided, then viewer engagement is enhanced, but the complexity of controlling and managing caption styles increases

Engineering Contradiction:
Improvevisual appealVSAvoidcaption control system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs self-service by automatically selecting and applying appropriate caption styles based on voice state analysis without requiring manual control or complex user input. The caption generation system independently analyzes voice characteristics and determines the most suitable caption style, reducing the burden on operators while maintaining high visual appeal and engagement

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11330342B2Method and apparatus for generating caption
Publication Date: 2022.05.10 NCSOFT CORP
  • US11330342B2 patent drawing
  • US11330342B2 patent drawing
  • US11330342B2 patent drawing

AI summary

A method and apparatus for generating a caption are provided. The method of generating a caption according to one embodiment comprises: generating caption text which corresponds to a voice of a speaker included in broadcast data; generating reference voice information using a part of the voice of the speaker included in the broadcast data; and generating caption style information for the caption text based on the voice of the speaker and the reference voice information.