Dynamic Caption Generation Adapting to Speaker Voice State
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current caption generation methods for broadcasting services lack dynamic adaptation to the voice state of speakers, such as volume, tone, and emotion, which limits the engagement and visual appeal of captions in real-time broadcasts.
Innovation Solution
A method and apparatus that generate caption text and style information based on the voice state of speakers, including control over size, color, font, position, rotation, and special effects, by analyzing voice states and applying corresponding changes to the caption text and screen style in real-time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If static caption generation is used, then the caption display is simple and stable, but the viewer engagement and visual appeal are limited
Solution Approach 1:
The patent applies dynamics by making the caption style adjustable based on voice state changes. The system transitions from static caption display to dynamic caption generation where caption characteristics (size, color, font, position, rotation, special effects) automatically adapt in real-time according to the speaker's voice state, thereby enhancing viewer engagement without requiring complex manual intervention
Solution Approach 2:
The patent implements parameter changes by modifying multiple caption parameters simultaneously based on voice state analysis. When voice state changes are detected (volume, tone, emotion), the system changes corresponding caption parameters such as size, color, font, position, rotation, and special effects, creating an adaptive captioning system that responds to spoken content characteristics
2Adaptability or versatility
If real-time voice state analysis is implemented, then caption adaptability is improved, but processing time and computational resources increase
Solution Approach 1:
The system maintains continuous voice state analysis throughout the broadcast, enabling real-time caption adaptation without interruption. The useful action of analyzing voice state and generating adaptive captions continues seamlessly as the broadcast progresses, ensuring that captions always reflect the current speaker state without requiring periodic pauses or batch processing
3Adaptability or versatility
If multiple caption style options are provided, then viewer engagement is enhanced, but the complexity of controlling and managing caption styles increases
Solution Approach 1:
The system performs self-service by automatically selecting and applying appropriate caption styles based on voice state analysis without requiring manual control or complex user input. The caption generation system independently analyzes voice characteristics and determines the most suitable caption style, reducing the burden on operators while maintaining high visual appeal and engagement
Data Source
AI summary
A method and apparatus for generating a caption are provided. The method of generating a caption according to one embodiment comprises: generating caption text which corresponds to a voice of a speaker included in broadcast data; generating reference voice information using a part of the voice of the speaker included in the broadcast data; and generating caption style information for the caption text based on the voice of the speaker and the reference voice information.


