Broadcast Audio Identification via Spectral Feature Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for automatically detecting and identifying broadcast programming, such as music or speech, are inefficient and prone to errors due to reliance on embedded cues, direct signal correlation, or limited computational capabilities, especially when dealing with time-compressed or altered content.
Innovation Solution
A system that registers known programming by digitally sampling and processing it into a database, then detects and identifies it by extracting feature sets and comparing them against stored codes, using a pattern vector generation and database search module to achieve accurate and robust identification without requiring embedded cues or impractical computational resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If direct signal correlation or embedded cues are used for program identification, then identification accuracy may be improved, but the system becomes vulnerable to time-compressed or altered content and requires impractical computational resources
Solution Approach 1:
The audio signal is divided into multiple short-time frames, and each frame is further segmented into frequency sub-bands. This segmentation allows the system to extract local spectral features that are more robust to time-compression than full-signal correlation, while reducing computational complexity through localized analysis.
Solution Approach 2:
The patent transforms the audio signal from the time domain to the frequency domain using Fourier transforms, and then extracts spectral features (energy, centroid, spread) that are less sensitive to playback speed variations. This parameter transformation enables the system to maintain identification accuracy under time-compression conditions.
2Measurement precision
If comprehensive signal analysis is performed to improve identification accuracy, then measurement precision increases, but computational complexity and processing time increase
Solution Approach 1:
The patent extracts only the most discriminative spectral features (energy, centroid, spread) from each frequency sub-band, rather than analyzing the entire signal spectrum. This feature extraction reduces computational complexity while maintaining identification accuracy by focusing on the most informative characteristics.
Solution Approach 2:
The system analyzes a subset of frequency sub-bands and selects only the most informative features for pattern matching, rather than processing all possible signal parameters. This partial analysis approach achieves sufficient identification accuracy with reduced computational burden.
3Device complexity
If traditional pattern recognition methods are used, then system complexity is reduced, but false positive and false negative rates increase
Solution Approach 1:
The patent extends traditional pattern recognition by incorporating temporal sequencing of spectral features across multiple frames and frequency sub-bands. This multi-dimensional feature space (frequency × time × feature type) provides more discriminative power for distinguishing similar programs, reducing false positives and negatives while maintaining reasonable system complexity.
Data Source
AI summary
This invention relates to the automatic detection and identification of broadcast programming, for example music, speech or video that is broadcast over radio, television, the Internet or other media. “Broadcast” means any readily available source of content, whether now known or hereafter devised, including streaming, peer to peer delivery or detection of network traffic. A known program is registered by deriving a numerical code for each of many short time segments during the program and storing the sequence of numerical codes and a reference to the identity of the program. Detection and identification of an input signal occurs by similarly extracting the numerical codes from it and comparing the sequence of detected numerical codes against the stored sequences. Testing criteria is applied that optimizes the rate of correct detections of the registered programming. Other optimizations in the comparison process are used to expedite the comparison process.


