Audio Fingerprint Indexing via Time-Variant Spectrogram Transforms
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio signal identification techniques fail to accurately identify noisy or distorted audio signals due to noise masking the underlying signal, leading to false negatives and inefficient processing requirements.
Innovation Solution
A system generates noise- and distortion-insensitive audio fingerprints by applying a time-variant transformation to the frequency spectrums of overlapping audio frames, selecting low-frequency components for indexing, and using phase differences to enhance matching accuracy and reduce computational resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional audio fingerprinting techniques are used to identify audio signals, then identification speed is improved through index-based searching, but noise and distortion in the signal cause false negatives and reduce identification accuracy
Solution Approach 1:
The patent transforms the audio signal from time domain to frequency domain using Fourier transform, then applies Mel-scale filtering to convert linear frequency to perceptual frequency. This parameter transformation allows the system to capture noise-insensitive features in the frequency domain that remain stable even when noise is present, thereby maintaining both fast index-based searching and high identification accuracy
Solution Approach 2:
The patent introduces an intermediate representation called audio fingerprint that extracts essential spectral features from the original audio signal. This fingerprint serves as a mediator between the raw audio signal and the identification process, capturing only the most discriminative features while filtering out noise. The fingerprint can be efficiently indexed and searched, enabling both fast retrieval and accurate matching even in noisy conditions
2Measurement precision
If the signal to noise ratio is very low (less than −6 dB), then noise completely masks the signal making identification difficult, but conventional techniques still attempt to match noisy portions resulting in false negatives
Solution Approach 1:
The patent converts the harmful effect of noise by transforming to the frequency domain where noise and signal occupy different spectral regions. The Mel-scale filtering and spectral feature extraction process selectively captures signal-dominated frequency regions while ignoring noise-dominated regions. This transforms the noise from a masking factor into a filterable component, allowing accurate signal identification even when noise completely masks the time-domain signal
Solution Approach 2:
The patent extracts only the most discriminative spectral features from the audio signal to create the fingerprint, rather than using the complete signal. By selecting only the most stable and informative frequency components (through Mel-scale filtering and spectral peak detection), the system achieves accurate identification with minimal signal data, effectively ignoring the noisy portions that would otherwise prevent identification
3Adaptability or versatility
If tempo shifting occurs when audio signal is played faster or slower than original speed, then spectral content shifts along time axis causing noise to increasingly mask the original signal
Solution Approach 1:
The patent uses dynamic time warping and elastic matching algorithms that can adapt to tempo variations. Instead of requiring exact temporal alignment, the system dynamically adjusts the matching process to accommodate speed changes. The fingerprint matching algorithm is designed to be elastic, allowing temporal stretching or compressing of the query fingerprint to match reference fingerprints at different playback speeds, thereby maintaining identification accuracy across tempo variations
4Speed
If index-based selection of reference fingerprints is used for matching, then searching speed is improved, but noise and distortion in the index cause failures to match against database indexes
Solution Approach 1:
The patent pre-processes the audio signal to create a compact fingerprint representation that captures essential spectral features before the matching process. This preliminary action creates a noise-resistant summary of the signal that can be efficiently indexed. The fingerprint is designed to be stable under various transformations, ensuring that index-based searching remains both fast and accurate even when the original signal contains noise or distortion
Data Source
AI summary
An audio identification system generates audio fingerprints and indexes associated with the audio fingerprints based on discrete and overlapping frames within a sample of an audio signal. The system applies a time-to-frequency domain transform to a time-sequence of frames, which may be filtered. The audio identification system then applies a time-variant transformation (e.g., a Discrete Cosine Transform) to the transformed frames and generates an audio fingerprint and index by selecting sets of coefficients of the time-variant transformation. The system selects coefficients that are less sensitive to possible noise and/or distortions in the underlying signal, such as low-frequency coefficients. The time-variant transformation provides sufficient sampling among the indexes by incorporating the phase information of the frames into the indexes. The system stores the audio fingerprint and other identifying information by index for efficient retrieval and matching of the retrieved fingerprints.


