Voice Extraction from Mixed Audio for Video Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for identifying video content through sound channels, such as watermarking, are ineffective when music and voice portions are mixed together, as they require pre-watermarking and are not efficient in separating voice from music.

Innovation Solution

The method involves converting audio signals from Descriptive Video Service (DVS) or Secondary Audio Program (SAP) to text using filtering, modulation, and nonlinear transformations, allowing for the separation of voice from music and subsequent identification of video content without altering the content prior to distribution or transmission.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If watermarking is used to identify video content, then content identification is possible, but the process requires pre-watermarking and cannot separate voice from music when they are mixed

Engineering Contradiction:
Improvecontent identification accuracyVSAvoidpre-processing requirement
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts voice information from the mixed audio signal by analyzing specific frequency ranges where human speech typically occurs. The system separates voice from music by identifying frequency bands characteristic of speech signals and isolating those components for recognition, eliminating the need for pre-watermarking while maintaining identification accuracy

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces speech recognition algorithms as an intermediary between the mixed audio signal and content identification. This intermediary process converts audio signals to text and compares them against a database of known dialog, enabling accurate content identification without requiring prior watermarking or complex pre-processing

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If narrow band pass filtering is applied to separate voice from music, then voice separation is achieved, but musical signals and frequencies are rejected

Engineering Contradiction:
Improvevoice separation accuracyVSAvoidmusical signal rejection
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent extracts only the necessary voice frequency components from the audio signal using narrow band pass filtering. By selecting specific frequency ranges where speech occurs, the system isolates voice information while naturally excluding music frequencies, achieving separation without requiring removal of musical content

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If frequency translation is used to translate narrow band spectrum to high frequency, then speech intelligibility is improved, but the original frequency spectrum is altered

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidfrequency spectrum stability
Core Design Contradiction:
Measurement precisionVSStability of the object's composition

Solution Approach 1:

The patent applies frequency translation to shift the narrow band voice spectrum to higher frequencies where it becomes more intelligible. This parameter change in the frequency domain enhances speech recognition capability by positioning voice frequencies in a range more suitable for human perception and automated recognition systems

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8527268B2Method and apparatus for improving speech recognition and identifying video program material or content
Publication Date: 2013.09.03 ROVI TECHNOLOGIES CORP
  • US8527268B2 patent drawing
  • US8527268B2 patent drawing
  • US8527268B2 patent drawing

AI summary

A system for identification of video content in a video signal is provided via a sound track audio signal. The audio signal is processed with filtering, frequency translation, and or non linear transformations to extract voice signals from the sound track channel. The extracted voice signals are coupled to a speech recognition system to provide in text form, the words of the video content, which is later compared with a reference library of words or dialog from known video programs or movies. Other attributes of the video signal or transport stream may be combined with closed caption data or closed caption text for identification purposes. Example attributes include DVS/SAP information, time code information, histograms, and or rendered video or pictures.