Digital Broadcast Terminal Audio Signal Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current digital broadcast terminals lack the capability to edit and customize digital broadcast contents according to user preferences, specifically failing to insert background music and display text captions from separated audio signals.
Innovation Solution
A method and terminal that extract audio and video signals from digital broadcast content, separate voice signals, convert them to text, and synchronize with additional audio data for customized playback, allowing users to insert background music and display captions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If digital broadcast terminal receives and plays broadcast content as-is, then the terminal structure remains simple, but the terminal cannot customize content according to user preferences
Solution Approach 1:
The audio signal is separated into voice signal and non-voice signal components. The terminal divides the audio stream into distinct segments that can be independently processed - the voice portion for caption generation and the non-voice portion for background music replacement, enabling customization without complete audio reprocessing
Solution Approach 2:
The voice signal is extracted from the mixed audio signal using voice activity detection. This extraction allows the system to isolate and process only the necessary components (voice for captions) while leaving other components (music effects) intact, reducing processing complexity while achieving customization
2Adaptability or versatility
If the terminal separates and processes audio signals to insert background music, then user personalization is enabled, but the processing time and computational resources increase
Solution Approach 1:
Voice activity detection is performed continuously during normal audio reception to pre-identify voice segments. This preliminary detection allows the system to prepare caption generation in advance during natural pauses in speech, reducing the perceived processing time when captions are actually needed
Solution Approach 2:
The system automatically detects voice activity and triggers caption generation without user intervention. The terminal self-manages the complex processing tasks by monitoring audio characteristics and autonomously determining when to separate signals, insert background music, or generate captions, minimizing user waiting time
3Ease of operation
If voice signals are converted to text for caption display, then accessibility and user experience improve, but the terminal requires additional processing functions
Solution Approach 1:
The terminal applies different processing qualities to different audio segments based on their characteristics. Voice signals undergo intensive processing (separation, recognition, caption generation), while non-voice segments receive minimal processing or none at all, optimizing resource allocation and reducing overall system complexity requirements
Data Source
AI summary
A method of converting digital broadcast contents and digital broadcast terminal having a function of the same are disclosed, by which music in compliance with a user's taste can be inserted as a background music by separating an audio signal of digital broadcast contents and by which a text can be displayed as a caption on a screen of the terminal in a manner of converting a voice recognized from a separated audio signal to the text. In converting digital broadcast contents including a video signal and an audio signal received via a digital broadcast network in a digital broadcast terminal, the present invention includes a step (a) of extracting the audio signal and the video signal for a specific section of the digital broadcast contents and a step (b) of synthesizing the extracted video signal with prescribed audio data.


