Voice Extraction from Mixed Audio for Video Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for identifying video content through sound channels, such as watermarking, are ineffective when music and voice portions are mixed together, as they require pre-watermarking and are not efficient in separating voice from music.
Innovation Solution
The method involves converting audio signals from Descriptive Video Service (DVS) or Secondary Audio Program (SAP) to text using filtering, modulation, and nonlinear transformations, allowing for the separation of voice from music and subsequent identification of video content without altering the content prior to distribution or transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If watermarking is used to identify video content, then content identification is possible, but the process requires pre-watermarking and cannot separate voice from music when they are mixed
Solution Approach 1:
The patent extracts voice information from the mixed audio signal by analyzing specific frequency ranges where human speech typically occurs. The system separates voice from music by identifying frequency bands characteristic of speech signals and isolating those components for recognition, eliminating the need for pre-watermarking while maintaining identification accuracy
Solution Approach 2:
The patent introduces speech recognition algorithms as an intermediary between the mixed audio signal and content identification. This intermediary process converts audio signals to text and compares them against a database of known dialog, enabling accurate content identification without requiring prior watermarking or complex pre-processing
2Measurement precision
If narrow band pass filtering is applied to separate voice from music, then voice separation is achieved, but musical signals and frequencies are rejected
Solution Approach 1:
The patent extracts only the necessary voice frequency components from the audio signal using narrow band pass filtering. By selecting specific frequency ranges where speech occurs, the system isolates voice information while naturally excluding music frequencies, achieving separation without requiring removal of musical content
3Measurement precision
If frequency translation is used to translate narrow band spectrum to high frequency, then speech intelligibility is improved, but the original frequency spectrum is altered
Solution Approach 1:
The patent applies frequency translation to shift the narrow band voice spectrum to higher frequencies where it becomes more intelligible. This parameter change in the frequency domain enhances speech recognition capability by positioning voice frequencies in a range more suitable for human perception and automated recognition systems
Data Source
AI summary
A system for identification of video content in a video signal is provided via a sound track audio signal. The audio signal is processed with filtering, frequency translation, and or non linear transformations to extract voice signals from the sound track channel. The extracted voice signals are coupled to a speech recognition system to provide in text form, the words of the video content, which is later compared with a reference library of words or dialog from known video programs or movies. Other attributes of the video signal or transport stream may be combined with closed caption data or closed caption text for identification purposes. Example attributes include DVS/SAP information, time code information, histograms, and or rendered video or pictures.


